Ai2 swaps GPU priorities for budgets: 98% of owed hours delivered
The Allen Institute for AI (Ai2) published on 9 October how it rebuilt the scheduler for its training clusters — thousands of NVIDIA H100, B200 and B300 GPUs in clusters of 88 to 1,024 GPUs, shared by about 150 researchers, with demand running at two to three times the available capacity.
What was broken. The old system was priority-based, and teams could mark workloads as non-preemptible up to a concurrent-GPU limit. Ai2 describes the predictable results: GPU "squatting" (holding GPUs idle so as not to lose them), priority inflation because "high" cost nothing, and attempts to fix it by handing important projects fixed GPU monopolies — which left GPUs idle whenever a team was between experiments. The team calls it a tragedy of the commons it had built for itself.
What replaced it. Three pieces:
- Budgets, not schedules. Managers allocate a share of GPU time, not GPUs, down a project tree — a leaf project might hold a guaranteed 35% claim on capacity. Any request not funded by a budget is preemptible, so squatting now spends the squatter's own allocation.
- Hierarchical fair-share over a 7-day lookback window — the same lineage as the Hadoop Fair Scheduler (2009), SLURM's Fair Tree and YARN. Under-served allocations sort ahead of over-served ones; work beyond a budget can still run on idle GPUs, unprotected.
- A scheduling contract. Every workload declares a minimum runtime — the shortest occupancy that makes real progress — and whether it is resumable. It is protected for that window, then can be preempted and requeued. That is what turns fair-share into time-slicing for jobs that otherwise run for days.

Results from the 30-day test after a cluster-by-cluster rollout from late July. Teams received 98% of the GPU hours they were owed; 13 of 15 allocations got 95% or more, the worst 90%. Debug jobs' p90 queue wait fell from 2 hours to 30 seconds (a simulator had predicted 6 hours to 5 minutes). Because unhealthy hosts now drain automatically once jobs reach their minimum runtime, repairs needing a human fell by 74%.
What did not improve. Interactive sessions that used to last a week are now protected for only 8 hours. Ai2 is watching possible capacity fragmentation that could lengthen waits for the largest jobs, and says documentation alone did not explain the new rules — live walkthroughs and per-team usage charts did.
Why it matters beyond Ai2. Most organisations with a shared GPU pool are running some version of the old system. The useful lesson is the order of operations: decide strategy as budgets in advance, let a well-known fair-share algorithm enforce it, and make preemption a contract users sign rather than a punishment.
Source: Ai2 on the Hugging Face blog, "Impactful scheduling for GPU clusters", 9 October 2026 — https://huggingface.co/blog/allenai/impactful-scheduling