Docker OOM-Killed With Free Host Memory: cgroup Limits Explained
Docker · Intermediate · 6 min read · published
This article was written by Claude (Anthropic) and published automatically.
What this solves: Explains why a container gets killed with exit code 137 even though `free -h` on the host shows plenty of RAM available.
The Problem
A background worker container keeps dying with exit code 137 every couple of hours. The team checks the host with free -h and sees 40GB free out of 64GB. No swapping, no CPU pressure, nothing in the host's system logs about memory pressure. Yet docker inspect confirms OOMKilled: true every single time. Someone bumps the host instance size from 64GB to 128GB RAM as a quick fix. The container still dies on the same schedule, at the same memory usage, like nothing changed — because nothing that mattered changed.
Why the Obvious Fix Falls Short
The instinct is to treat container memory like process memory on a bare-metal box: if the machine has memory, the process should be able to use it. But a container's memory usage is governed by its cgroup limit (--memory in docker run, or resources.limits.memory in Kubernetes), which is completely independent of host-wide availability. The kernel's cgroup OOM killer watches the container's own memory cgroup, not system-wide MemAvailable. If the container's cgroup limit is 512MB, it gets killed at 512MB even while sitting on a host with terabytes free. Upgrading the host changes nothing because the ceiling being hit is the cgroup's, not the host's.
A second obvious fix — just remove the memory limit entirely — trades one failure mode for a worse one: an unbounded container can now consume enough host memory to trigger the host kernel's OOM killer, which may kill unrelated processes, including the Docker daemon itself, based on OOM-score heuristics you don't control.
How It Actually Works
Docker enforces --memory using Linux cgroups (v1 or v2). Each container is a cgroup with its own memory accounting: RSS, page cache attributed to the container, kernel memory, and (depending on version) network buffers. When usage hits the limit, the kernel's cgroup-scoped OOM killer fires inside that cgroup, picking a process to kill within it — usually PID 1 of the container, which tears down the whole container. This happens without ever consulting host-wide free memory.
A common trap: many runtimes (JVM, Node's V8, Go's GC in older versions) historically sized their default heap/behavior based on host memory (via /proc/meminfo) rather than the cgroup limit, because cgroup-awareness had to be added explicitly. A JVM defaulting to 25% of a 64GB host heap inside a 512MB container is a guaranteed OOM kill.
flowchart TD
A[Process allocates memory in container] --> B{Cgroup memory usage vs limit}
B -- under limit --> A
B -- at/over limit --> C[cgroup OOM killer triggered]
C --> D[Kill PID 1 in container]
D --> E[Container exits 137]
F[Host free memory: 40GB] -.->|irrelevant to decision| B
Before and After
# BEFORE: no explicit memory awareness, JVM sizes heap off host RAM
FROM eclipse-temurin:17-jre
COPY app.jar /app.jar
CMD ["java", "-jar", "/app.jar"]
# On a 64GB host with --memory=512m, JVM may still try to use
# a heap sized from host memory, blowing past the cgroup limit.
# AFTER: explicit cgroup-aware heap sizing + right-sized limit
FROM eclipse-temurin:17-jre
COPY app.jar /app.jar
# Cap heap relative to the container's own cgroup limit, not host RAM
CMD ["java", "-XX:MaxRAMPercentage=70.0", "-jar", "/app.jar"]
# docker run: measured peak usage was ~900MB, so give headroom, not host RAM
docker run --memory=1536m --memory-swap=1536m myworker:latest
When NOT to Use This
Don't chase this rabbit hole if your container is being killed by the host's OOM killer instead of the cgroup one — check dmesg for which killer fired and against which cgroup. If several containers on a shared host are collectively overcommitting real RAM, the fix is capacity planning and setting sane limits/requests across all workloads (or using Kubernetes' QoS classes), not tuning one container in isolation. Also, if the workload is genuinely memory-hungry by design (large in-memory caches, big batch jobs), raising the limit is the correct move, not a compromise — don't starve it artificially just to
Key takeaway: A container is killed against its own cgroup memory limit, not the host's available RAM, so always size limits from measured working-set memory plus headroom, not from host capacity.
Real-world challenge
A worker container processing image uploads restarts randomly with exit code 137. `docker stats` shows memory usage climbing steadily before each restart, and `docker inspect` shows a `--memory=512m` limit. The host machine has 64GB of RAM and rarely goes above 30% utilization. The team's first instinct is to add more RAM to the host.
Adding host RAM won't help because the container is bounded by its own cgroup limit, not host capacity.
- Confirm the kill:
dmesg | grep -i 'killed process'ordocker inspect <id> --format '{{.State.OOMKilled}}'returnstrue. - Watch actual usage inside the limit:
docker stats <id>shows RSS climbing toward 512MB right before each restart. - Check what's consuming memory: is it a real leak (heap growing unbounded across requests) or just a limit set too low for legitimate peak usage (e.g. batch image resizing spikes)?
- Fix the right thing:
# If it's legitimate peak usage, raise the limit to cover real peak + headroom
docker run --memory=1536m --memory-swap=1536m myworker
# If it's a leak, fix the app and keep the limit tight to catch regressions early
- For JVM-based containers, explicitly cap heap relative to the container limit (e.g.
-XX:MaxRAMPercentage=70.0) instead of letting the JVM guess based on host memory.