Node Swap Is GA in Kubernetes. Idle Agent Pods Are Exactly What It's For — and Exactly How to Get Burned.
Memory is the first hard limit most Kubernetes clusters hit. Nodes run out of RAM long before they run out of CPU, and the scheduler stops placing pods the moment requests exhaust allocatable memory — even if the CPUs are half-idle. You end up paying for cores you can't use because there's no RAM left to hand out.
Agentic AI workloads make this worse in a specific, annoying way. A typical agent pod loads a language runtime, maybe a model client, maybe a sandbox for executing untrusted code — a few gigabytes of resident footprint — spikes on startup, and then sits idle waiting for the next prompt. That memory is reserved and resident, but it's cold most of the time. Multiply by a fleet of agents and you've got a node that's RAM-bound on memory that nobody is actually touching.
Kubernetes node swap just graduated to GA, and this is precisely the workload it was built for. It's also a very good way to wreck your tail latency if you treat it as free memory. Here's how to get the density win without the cliff.
Why swap was off for a decade
For most of Kubernetes' life, the kubelet refused to even start if swap was enabled on the node:
# the historical default — kubelet bails out if swap is on
--fail-swap-on=true
The standard node-provisioning advice was blunt: swapoff -a and delete the swap entry from /etc/fstab. The reasoning was sound for its time. With swap on, a container's memory limit stopped meaning what it said — a process could blow past its RSS limit into swap, and the predictable OOM-kill behavior that QoS classes depend on got fuzzy. Swap made resource isolation non-deterministic, so Kubernetes banned it.
What changed is cgroup v2. Under cgroup v2, the kernel exposes a separate memory.swap.max knob per cgroup, so the kubelet can account for and bound swap usage per container instead of letting it leak globally. That's the foundation the GA feature is built on, and it's why cgroup v2 is a hard requirement — there is no swap support on cgroup v1 nodes.
What GA actually gives you
Two pieces of kubelet configuration. You stop failing on swap, and you set the swap behavior:
apiVersion: kubelet.config.k8s.io/v1beta1
kind: KubeletConfiguration
failSwapOn: false
memorySwap:
swapBehavior: LimitedSwap
The two behaviors matter. NoSwap (the default even after GA) lets the node have swap — so system daemons and the kernel can use it — while forbidding any Kubernetes-managed container from touching it. LimitedSwap is the one that actually hands swap to workloads, and it does so under strict rules:
- Only Burstable QoS pods get swap. A pod is Burstable when it sets a memory request lower than its limit (or omits the limit).
- Guaranteed pods get no swap. Request equals limit is a promise of dedicated memory; swapping it would break that promise, so the kubelet keeps those pages in RAM.
- BestEffort pods get no swap. With no memory request there's nothing to size a swap allocation against, and an unbounded best-effort process could drain the whole swap device.
For the containers that do qualify, the swap allocation is proportional to the memory request, not a free-for-all:
containerSwapLimit = (containerMemoryRequest / nodeTotalPhysicalMemory) * totalNodeSwap
So a container requesting 4 GiB on a 64 GiB node with 32 GiB of swap gets about (4/64) * 32 = 2 GiB of swap headroom. This is the clever part: swap is distributed in proportion to how much memory a workload already asked for, so a larger pod gets a larger cushion and a tiny pod can't monopolize the device.
The density math
The win is straightforward once idle pods can shed their cold pages. Instead of every pod pinning its full footprint in RAM, the kernel pages out the parts an idle agent isn't touching, and the scheduler can fit more pods before it hits the allocatable-memory wall.
The latency cliff
Here's where teams get burned. Swap is not free memory — it's slow memory. A page that got swapped out has to be faulted back in from the swap device the instant the process touches it again. For an agent that's been idle and then gets a prompt, the first request after idle pays that swap-in cost, and on a busy node under memory pressure it can be a brutal tail-latency spike.
The failure mode that actually takes a node down, though, is thrashing: when the sum of active working sets exceeds physical RAM, the kernel spends its time paging the same hot pages in and out instead of running your code. CPU goes to iowait, everything on the node slows at once, and because it degrades gradually it often looks like a mysterious, cluster-wide latency regression rather than an OOM you can point at.
Two things keep you on the right side of that line. First, the swap device has to be fast — NVMe SSD, or compressed-RAM swap (zram/zswap), never a spinning disk or a network volume. Second, lower the kernel's eagerness to swap so it only reaches for the device when it genuinely needs to:
# confirm cgroup v2 — must print "cgroup2fs"
stat -fc %T /sys/fs/cgroup
# provision a swap file on local NVMe (self-managed nodes)
fallocate -l 16G /swapfile
chmod 600 /swapfile
mkswap /swapfile
swapon /swapfile
# swap reluctantly, not eagerly — default is 60
sysctl -w vm.swappiness=20
Monitor pressure, not bytes
The instinct is to alert on swap-used bytes. That's the wrong signal — a node can have gigabytes of cold pages happily parked in swap and be perfectly healthy. What hurts is stalling on memory, and the kernel tells you that directly through Pressure Stall Information:
# how long tasks stalled waiting on memory, as a rolling percentage
cat /proc/pressure/memory
# some avg10=0.00 avg60=0.42 avg300=1.10 total=...
# full avg10=0.00 avg60=0.00 avg300=0.00 total=...
# the swap-in / swap-out rate — rising pswpin under load means refaulting
grep -E 'pswpin|pswpout' /proc/vmstat
The line that matters is full under /proc/pressure/memory: it's the fraction of time every runnable task was stalled waiting on memory — i.e. real thrashing, not healthy background paging. A sustained non-zero full avg60 is your alert threshold. If you're on Prometheus, node_exporter exposes the same data as node_pressure_memory_stalled_seconds_total alongside node_vmstat_pswpin / pswpout; alert on the rate of change of the PSI counter, and let swap-used bytes be a dashboard number, not a pager.
A workable split
Put the pieces together and the policy writes itself. Classify workloads by what they can tolerate:
- Latency-critical serving (the inference gateway, the request-path API): keep it Guaranteed — request equals limit — so it never gets swapped and never pays the swap-in penalty.
- Bursty, idle-heavy agents (sandboxes, background workers, anything that spikes then waits): run them Burstable so their cold pages can drop to swap and free RAM for density.
- Throwaway best-effort jobs: they get no swap by design, which is fine — they also get evicted first under pressure.
That mapping is the whole game: swap earns its keep exactly where the footprint is large and mostly cold, and you deliberately exclude it everywhere latency is the product.
Lesson
Node swap is a density lever, not a memory upgrade. It trades cold RAM — which you were paying for and not using — for disk latency on wake, and that's a great trade for idle-heavy agent pods on fast storage and a terrible one for anything latency-critical with a working set bigger than RAM. Keep it on Burstable QoS, keep your serving path Guaranteed, put swap on NVMe or zram, and alert on PSI full pressure rather than swap-used bytes. Do that, and the GA feature pays for itself in pods-per-node. Skip the monitoring, and the first time a working set outgrows RAM you'll rediscover why Kubernetes spent a decade refusing to boot with swap on.
Hitting something like this in production? I help teams with performance engineering, SRE/observability, and AI-driven root cause analysis — work with me.
Comments
Post a Comment