Kubernetes node swap on NVMe triples sandbox density in benchmarks, but only for idle memory
The Kubernetes blog published on 5 October a set of benchmarks for running nodes with swap enabled, a feature that reached general availability in Kubernetes 1.34. The argument is aimed at agentic AI workloads: sandboxes that need a lot of memory to start and run untrusted code, then sit idle waiting for the next prompt, holding RAM that caps how many pods fit on a node.
Why swap was discouraged, and what changed. Under cgroup v1, memory and swap shared one combined limit, so a container's real memory use was hard to predict and isolate. Kubernetes' swap support relies on cgroup v2, which accounts swap separately. The other objection, the latency of paging to spinning disks, is largely removed by fast local NVMe SSDs.

What was measured. Three workload types, with swap on Local SSD:
- Linux kernel build (CI-style). The minimum memory limit to avoid an out-of-memory kill fell from 600 MB to 300 MB, and the build ran in 374 s against 433 s without swap. Cutting the limit to 200 MB pushed the active working set into swap and made the build more than 40% slower.
- Headless Chrome sandboxes. Under gVisor, a node went from 80 to 160 concurrent pods; under Kata Containers, from 40 to 50. Plain runc containers on a 32-vCPU, 120 GB node went from failing past 512 pods to 768.
- Python sandboxes under gVisor. From 80 to 240 concurrent pods, the largest gain at 3x.
The caveat the authors make themselves. Swap works as insurance for bursts and for memory that is resident but idle; it is not a replacement for RAM the workload is actively using. The kernel-build result shows the boundary: halving the limit was free, cutting it by two thirds was not.
The tests ran on Google Cloud with Local SSD swap; the plain-runc sweep used a c4-standard-32 node.