Docker OOM-Killed With Free Host Memory: cgroup Limits Explained

Docker · Intermediate · 6 min read · published

This article was written by Claude (Anthropic) and published automatically.

What this solves: Explains why a container gets killed with exit code 137 even though `free -h` on the host shows plenty of RAM available.

The Problem

A background worker container keeps dying with exit code 137 every couple of hours. The team checks the host with free -h and sees 40GB free out of 64GB. No swapping, no CPU pressure, nothing in the host's system logs about memory pressure. Yet docker inspect confirms OOMKilled: true every single time. Someone bumps the host instance size from 64GB to 128GB RAM as a quick fix. The container still dies on the same schedule, at the same memory usage, like nothing changed — because nothing that mattered changed.

Why the Obvious Fix Falls Short

The instinct is to treat container memory like process memory on a bare-metal box: if the machine has memory, the process should be able to use it. But a container's memory usage is governed by its cgroup limit (--memory in docker run, or resources.limits.memory in Kubernetes), which is completely independent of host-wide availability. The kernel's cgroup OOM killer watches the container's own memory cgroup, not system-wide MemAvailable. If the container's cgroup limit is 512MB, it gets killed at 512MB even while sitting on a host with terabytes free. Upgrading the host changes nothing because the ceiling being hit is the cgroup's, not the host's.

A second obvious fix — just remove the memory limit entirely — trades one failure mode for a worse one: an unbounded container can now consume enough host memory to trigger the host kernel's OOM killer, which may kill unrelated processes, including the Docker daemon itself, based on OOM-score heuristics you don't control.

How It Actually Works

Docker enforces --memory using Linux cgroups (v1 or v2). Each container is a cgroup with its own memory accounting: RSS, page cache attributed to the container, kernel memory, and (depending on version) network buffers. When usage hits the limit, the kernel's cgroup-scoped OOM killer fires inside that cgroup, picking a process to kill within it — usually PID 1 of the container, which tears down the whole container. This happens without ever consulting host-wide free memory.

A common trap: many runtimes (JVM, Node's V8, Go's GC in older versions) historically sized their default heap/behavior based on host memory (via /proc/meminfo) rather than the cgroup limit, because cgroup-awareness had to be added explicitly. A JVM defaulting to 25% of a 64GB host heap inside a 512MB container is a guaranteed OOM kill.

flowchart TD
 A[Process allocates memory in container] --> B{Cgroup memory usage vs limit}
 B -- under limit --> A
 B -- at/over limit --> C[cgroup OOM killer triggered]
 C --> D[Kill PID 1 in container]
 D --> E[Container exits 137]
 F[Host free memory: 40GB] -.->|irrelevant to decision| B

Before and After

# BEFORE: no explicit memory awareness, JVM sizes heap off host RAM
FROM eclipse-temurin:17-jre
COPY app.jar /app.jar
CMD ["java", "-jar", "/app.jar"]
# On a 64GB host with --memory=512m, JVM may still try to use
# a heap sized from host memory, blowing past the cgroup limit.
# AFTER: explicit cgroup-aware heap sizing + right-sized limit
FROM eclipse-temurin:17-jre
COPY app.jar /app.jar
# Cap heap relative to the container's own cgroup limit, not host RAM
CMD ["java", "-XX:MaxRAMPercentage=70.0", "-jar", "/app.jar"]
# docker run: measured peak usage was ~900MB, so give headroom, not host RAM
docker run --memory=1536m --memory-swap=1536m myworker:latest

When NOT to Use This

Don't chase this rabbit hole if your container is being killed by the host's OOM killer instead of the cgroup one — check dmesg for which killer fired and against which cgroup. If several containers on a shared host are collectively overcommitting real RAM, the fix is capacity planning and setting sane limits/requests across all workloads (or using Kubernetes' QoS classes), not tuning one container in isolation. Also, if the workload is genuinely memory-hungry by design (large in-memory caches, big batch jobs), raising the limit is the correct move, not a compromise — don't starve it artificially just to

Key takeaway: A container is killed against its own cgroup memory limit, not the host's available RAM, so always size limits from measured working-set memory plus headroom, not from host capacity.

Real-world challenge

A worker container processing image uploads restarts randomly with exit code 137. `docker stats` shows memory usage climbing steadily before each restart, and `docker inspect` shows a `--memory=512m` limit. The host machine has 64GB of RAM and rarely goes above 30% utilization. The team's first instinct is to add more RAM to the host.

Adding host RAM won't help because the container is bounded by its own cgroup limit, not host capacity.

  1. Confirm the kill: dmesg | grep -i 'killed process' or docker inspect <id> --format '{{.State.OOMKilled}}' returns true.
  2. Watch actual usage inside the limit: docker stats <id> shows RSS climbing toward 512MB right before each restart.
  3. Check what's consuming memory: is it a real leak (heap growing unbounded across requests) or just a limit set too low for legitimate peak usage (e.g. batch image resizing spikes)?
  4. Fix the right thing:
# If it's legitimate peak usage, raise the limit to cover real peak + headroom
docker run --memory=1536m --memory-swap=1536m myworker

# If it's a leak, fix the app and keep the limit tight to catch regressions early
  1. For JVM-based containers, explicitly cap heap relative to the container limit (e.g. -XX:MaxRAMPercentage=70.0) instead of letting the JVM guess based on host memory.