Scheduling is one of the oldest problems in computer science. How do you decide which job runs on which resource, and when? For decades the answer was heuristics: First-Fit, Round-Robin, Least-Loaded, Shortest-Job-First. Rules written by engineers, tuned by intuition, and largely unchanged since the 1970s.
Then, starting around 2016, a handful of research groups began asking a different question: what if a scheduling agent could learn from the workload itself — discovering strategies that no human engineer would think to write down?
This is back when Deep Learning was a niche and still research driven sport. Now with the adoption and growth of AI, and everyone distracted with LLMs its time to revisit these technologies and considered where they are applied to commercial interests.
The Problem with Heuristics
Before getting into the algorithms, it’s worth being clear about why heuristics fall short.
AI workloads and AI model training demands have changed not just the amount of powerful compute we need but the patterns of demand as well. For most businesses that either built or used software based services, infrastructure demands were based on patterns of consumption. Patterns you could forecast but not prove as compute consumption scheduling was all based on proven workload demands.
For years the tech industry used heuristic scheduling approaches. The cloud providers were particularly effective here.
A heuristic like First-Fit is fast and simple: assign each job to the first resource that can take it. It works fine when jobs are interchangeable. But in modern infrastructure they rarely are. A Stable Diffusion request that can reuse a model already loaded in GPU memory takes 1 second. The same request landing on a pod that needs to load the model from scratch takes 11 seconds. The difference is invisible to First-Fit.
The deeper problem is that heuristics cannot reason across time. They see one job at a time. A scheduler that could look ahead — that could learn “if I place this job here, the next three jobs of this type will also land on a warm pod” — would consistently outperform any fixed rule.
That forward-looking, adaptive reasoning is exactly what reinforcement learning is designed for.
2016: The First Wave — Resource Management with Deep RL
The paper that opened the door was “Resource Management with Deep Reinforcement Learning” (Mao et al., MIT CSAIL, 2016), sometimes called DeepRM.
The problem: allocate CPU and memory resources across competing jobs in a cluster. The existing approach was rule-based bin-packing. DeepRM replaced it with a policy gradient agent that took a visual representation of the resource grid as input (literally a stack of colour images showing which resources were occupied) and learned to assign jobs to minimise average job slowdown.
The result was striking. DeepRM outperformed all heuristics — not by a small margin, but by learning qualitatively different strategies. It would sometimes hold back short jobs to wait for a better slot to open up, a behaviour no greedy heuristic would produce.
Algorithm used: REINFORCE (Williams, 1992) — one of the earliest policy gradient methods. Simple but noisy. The agent samples trajectories, observes the total reward, and nudges the policy in the direction of better outcomes.
Why it mattered: It proved that RL could learn useful scheduling policies from scratch with no human-written rules. The 2016 paper is the intellectual ancestor of almost every RL scheduling paper that followed.
2017: Adaptive Bitrate Streaming — Pensieve
The same MIT group (Mao et al.) next applied RL to a different resource allocation problem: adaptive bitrate video streaming in their 2017 paper “Real-World Performance of Adaptive Bitrate Algorithms”, and more directly in Pensieve (SIGCOMM 2017).
The problem: Netflix and YouTube need to decide, chunk by chunk, what video quality to serve over a fluctuating network connection. Too high and you get buffering. Too low and the picture looks bad.
Pensieve trained an A3C agent (Asynchronous Advantage Actor-Critic — Mnih et al., DeepMind, 2016) on real network traces to make these decisions. It outperformed all hand-tuned heuristics because it could implicitly model network patterns over time rather than reacting only to the current throughput measurement.
Algorithm used: A3C (DeepMind, 2016). Multiple parallel agents explore in different environments simultaneously, feeding gradients back to a shared central model. Much faster to train than single-agent Reinforce.
Why it matters for scheduling: Pensieve established a pattern that would recur throughout the field — train on historical traces, evaluate on a held-out window, beat the best human-designed policy. The same experimental design shows up in cluster scheduling, network routing, and inference serving.
2017–2018: PPO — The Algorithm That Took Over
Around the same time, OpenAI published Proximal Policy Optimization (Schulman et al., 2017). PPO is not a scheduling algorithm — it is a general-purpose policy gradient method. But it rapidly became the default choice for almost every scheduling research paper, for good reason.
The core insight of PPO: earlier policy gradient methods were unstable because a single large gradient update could push the policy into a bad region it would never recover from. PPO adds a clip on how much the policy can change in one update, keeping training stable without the complexity of earlier constrained methods.
For scheduling problems specifically, PPO is attractive because:
Scheduling episodes are discrete and structured (assign job, observe outcome, repeat)
The reward signal is relatively dense (you know after each decision whether it was good)
PPO’s sample efficiency is reasonable for medium-complexity problems
Every major RL scheduling paper since 2018 either uses PPO or explicitly benchmarks against it.
2018: Traffic Engineering at Scale — AuTO
AuTO (Chen et al., 2018) tackled network traffic engineering in large data centres — deciding how to route traffic flows across a network to minimise congestion and latency.
What made AuTO interesting was the scale problem. A real data centre has thousands of flows and thousands of paths. A single RL agent handling everything would have an action space too large to explore effectively. The AuTO team’s solution: a two-level architecture. A lightweight rule-based system handles short flows (the vast majority by count). A deeper RL agent handles elephant flows (rare but bandwidth-dominant). The two tiers work together without stepping on each other.
Algorithm used: Actor-Critic with custom state representation encoding flow statistics.
Why it matters: AuTO introduced the idea of hybrid scheduling — RL for the hard decisions, heuristics as a cheap fallback for the easy ones. This is directly analogous to how production inference schedulers now work: the RL agent makes routing decisions when there is genuine optionality, and falls back to First-Fit when all pods are equivalent.
2019: Graph Neural Networks Meet Scheduling — Decima
Decima (Mao et al., MIT CSAIL, 2019) is probably the most cited RL scheduling paper of the decade.
The problem: scheduling DAG-structured jobs in Apache Spark clusters. A Spark job is not a single task — it is a directed acyclic graph of stages, where later stages cannot start until earlier ones finish. Traditional schedulers treat these stages independently and miss the dependency structure entirely.
Decima’s contribution was to represent the cluster and job queue as a graph and use a Graph Neural Network (GNN) to generate node embeddings that captured the structural relationships. The GNN’s output fed into a policy network that decided which stage to run next.
The result: Decima reduced average job completion time by 21% over the best existing heuristic in a Spark production trace. More importantly, it discovered scheduling strategies that exploited the DAG structure in ways no human engineer had thought to encode.
Algorithm used: Policy gradient with GNN-based observation. The GNN allows the model to generalise across different job graph topologies rather than memorising fixed patterns.
Why it matters: Decima introduced graph-structured observation to scheduling — instead of giving the agent a flat feature vector, give it a relational structure so it can reason about dependencies directly. This idea is now being applied to GPU inference scheduling, where the “graph” is the relationship between jobs sharing base models.
2019: HPC Job Scheduling — RLScheduler
RLScheduler (Zhou et al., Argonne National Laboratory / University of Chicago, 2019) applied RL to HPC (High Performance Computing) batch job scheduling on supercomputers.
HPC scheduling is a distinctive problem. Jobs arrive over hours or days. Each job requests specific numbers of nodes, and the scheduler must decide which jobs to run now versus queue — while trying to maximise utilisation and minimise wait time for users.
RLScheduler used a 1D convolutional neural network to encode the job queue (ordered by arrival) and trained a policy to select which job to execute next. It outperformed EASY backfill — the industry-standard HPC scheduler algorithm — across multiple workload types.
Algorithm used: PPO with a CNN-based encoder for the job queue.
Why it matters: RLScheduler showed that RL could work in batch scheduling contexts where decisions span much longer time horizons than in real-time resource management. It also demonstrated the technique on production HPC traces from real supercomputing facilities.
2020–2022: Constrained RL Enters the Picture
A recurring problem across all these papers: RL agents optimise for whatever reward signal you give them, and nothing else. If you reward cache hit rate, the agent will sacrifice throughput to get it. If you reward throughput, it ignores cache hits.
Real scheduling problems have multiple objectives that must all be satisfied — not optimised jointly, but constrained. You might want to maximise cache hit rate subject to throughput staying above some minimum level.
This led to growing interest in Constrained Markov Decision Processes (CMDPs) and Lagrangian RL methods. The theoretical foundations go back to Altman (1999), but practical deep RL implementations for scheduling came largely from work at Stanford, MIT, and Berkeley between 2020 and 2022.
The key idea: instead of a single scalar reward, the agent has a primary objective (maximise X) and one or more cost functions (keep Y above threshold Z). A Lagrange multiplier — updated during training — automatically balances the trade-off. When the cost constraint is violated, the multiplier increases, penalising the agent; when the constraint is met, the multiplier relaxes.
Why it matters: Lagrangian RL is the theoretically correct way to handle multi-objective scheduling when some objectives are constraints rather than preferences. For GPU inference scheduling, this means: optimise cache hit rate, subject to throughput not falling below baseline.
2021–2023: Soft Actor-Critic for Inference Systems
Soft Actor-Critic (Haarnoja et al., UC Berkeley, 2018) was developed originally for continuous control problems — robotics, locomotion. It adds an entropy regularisation term to the reward, encouraging the agent to maintain a diverse, exploratory policy rather than collapsing to a single deterministic behaviour.
In scheduling research, SAC-Discrete (Christodoulou, 2019 — an adaptation of SAC to discrete action spaces) has attracted attention for inference systems where the reward signal is sparse. In a high-traffic serving cluster, cache hits are relatively rare events. Sparse rewards make standard PPO slow to learn — the agent sees too many steps with zero signal before getting useful feedback.
SAC’s entropy regularisation keeps the policy exploring longer, which is exactly what sparse-reward environments need. Several 2022–2023 papers on LLM and diffusion model serving have benchmarked SAC-Discrete against PPO variants and found it more sample-efficient in low-traffic, sparse-reward regimes.
Algorithm used: Actor-Critic with maximum entropy objective and replay buffer (off-policy). The replay buffer is a significant advantage — each experience can be reused many times, reducing the amount of online interaction needed.
Where This Is All Going: Inference Scheduling
The trajectory of this field leads naturally to one of the most pressing problems in AI infrastructure today: GPU inference scheduling for multi-model serving.
When an organisation runs dozens or hundreds of models simultaneously on shared GPU infrastructure — Stable Diffusion variants, LLMs, embedding models — the scheduling problem has a structure that none of the classic heuristics are equipped to handle. The key variable is not just which pod has capacity, but which pod already has the right base model loaded in VRAM.
A model loaded in VRAM is an asset. Routing a request to a pod that already holds that model costs ~1 second. Routing the same request to a cold pod costs 10–15 seconds of model loading time — pure latency with no useful work done.
The research trajectory from DeepRM through Decima through Constrained RL points directly at this problem. An agent that can learn co-location patterns — which model types tend to arrive together, which pods are likely to free up, when deferring a job briefly will result in a warm placement — will consistently outperform any heuristic.
Preliminary results from our own research using MaskablePPO trained on the Alibaba GenTD26 inference trace (26,128 jobs, 79 base models) show +55% improvement in cache hit rate over the strongest heuristic baseline (FGD). That translates directly to reduced cold start latency and lower GPU cost per served request.
The next frontier is applying the multi-objective techniques developed between 2020 and 2023 — Constrained RL, SAC-Discrete, GNN-based observation — to bring throughput, utilisation, and cache hit rate into alignment rather than in tension.
Where to Go Next
If you want to go deeper, the papers worth reading in order are:
Mao et al., “Resource Management with Deep Reinforcement Learning” (2016) — the foundation
Mao et al., “Decima” (SIGCOMM 2019) — GNN + scheduling, the state of the art for DAG jobs
Schulman et al., “Proximal Policy Optimization Algorithms” (2017) — understand the algorithm most papers use
Haarnoja et al., “Soft Actor-Critic” (2018) — the off-policy alternative worth knowing
Altman, “Constrained Markov Decision Processes” (1999) — the theory behind multi-objective RL
Tags: reinforcement learning, GPU scheduling, MLOps, inference serving, machine learning systems



