Photo by Bernd ๐ท Dittrich on Unsplash
So there I was, monitoring a production deployment at some ungodly hour, and my containers kept dying. No crash logs, no Python tracebacks, nothing useful. Just gone. I ran docker ps and sure enough — the container wasn't there. Checked the exit code and got hit with the classic:
$ docker inspect my-app --format='{{.State.ExitCode}}'
137
Exit code 137. If you've seen this, you already know what I'm about to say. That's the Docker OOM killer doing its thing — the Linux kernel's out-of-memory killer stepped in, decided your container was eating too much RAM, and just... murdered it. No warning, no graceful shutdown. Just a cold SIGKILL.
You can confirm the OOM kill happened by checking the container's inspect output more closely:
$ docker inspect my-app --format='{{.State.OOMKilled}}'
true
There it is. true. The kernel isn't even apologetic about it.
Why This Actually Happens
Here's the thing — the OOM killer isn't a Docker bug. It's a Linux kernel feature. When the system (or a container's cgroup) runs out of memory, the kernel has to make a decision: kill something or let the whole system grind to a halt. It picks the process using the most memory and terminates it with SIGKILL (signal 9), which is why the exit code is 128 + 9 = 137. That math trips people up every time.
In Docker's case, this usually happens for one of two reasons. Either your container genuinely has a memory limit set and your app is blowing past it, or the host machine itself is running low on RAM and the kernel decides your container is the sacrificial lamb. Both are annoying for completely different reasons.
I've seen this a hundred times with Java apps especially. JVM loves to grab memory aggressively and the default heap settings don't care about your container limits. But Node.js apps, Python data processing scripts, even Nginx under heavy load — anything can trigger this if the conditions are right.
Fix #1: Raise (or Remove) the Memory Limit
The most straightforward fix — if your app genuinely needs more memory — is to bump up the limit. If you're running the container with docker run:
$ docker run -m 1g --memory-swap 2g my-app
The -m flag sets the memory limit, and --memory-swap sets the total including swap. Setting --memory-swap to the same value as -m disables swap entirely (which you might actually want in some cases to surface memory issues faster in dev).
If you're on Docker Compose, it goes in your docker-compose.yml like this:
services:
my-app:
image: my-app:latest
deploy:
resources:
limits:
memory: 1g
reservations:
memory: 512m
Honestly, a lot of people skip setting memory limits entirely and then act surprised when the OOM killer shows up. Setting reasonable limits is just good hygiene — it forces you to actually think about your app's resource profile.
Fix #2: Profile and Fix the Memory Leak
Now here's where it gets tricky. Sometimes the container getting killed isn't a configuration problem — it's a genuine memory leak or inefficient memory usage in your application. Throwing more RAM at a leaking app is just delaying the inevitable.
First, watch the memory usage in real time before the kill happens:
$ docker stats my-app
You'll see something like this updating live:
CONTAINER ID NAME CPU % MEM USAGE / LIMIT MEM %
a3f2c1d9e8b7 my-app 12.3% 890MiB / 1GiB 86.9%
If that MEM % is creeping toward 100% over time without leveling off, you've got a leak. For Java apps specifically, the fix is usually telling the JVM to respect container memory limits — because by default it doesn't. Add this to your JVM flags:
JAVA_OPTS="-XX:+UseContainerSupport -XX:MaxRAMPercentage=75.0"
UseContainerSupport was added in Java 10 and backported to Java 8u191+. Without it, the JVM reads the host machine's total RAM, not the container's cgroup limit. It'll happily allocate 4GB of heap when your container only has 1GB. This one setting has saved me more on-call incidents than I care to admit.
Fix #3: Tune the OOM Score (Fallback Approach)
Sometimes you can't easily fix the memory usage and you can't just increase limits — maybe you're on a shared host with constrained resources. In that case, you can at least influence which process the OOM killer targets first. Docker exposes the --oom-score-adj flag for this:
$ docker run --oom-score-adj=-500 my-app
The score ranges from -1000 to 1000. Lower score = less likely to get killed. A value of -1000 makes the process basically immune to the OOM killer. Setting it negative for your critical containers means the kernel will go after other, less important processes first if things get tight.
There's also --oom-kill-disable but — fair warning — use that one carefully. Disabling the OOM killer entirely can cause the host to lock up if memory actually runs out. It's not something I'd recommend in production unless you really know what you're doing and have other safeguards in place.
# Use with caution — disables OOM killer for this container
$ docker run --oom-kill-disable my-app
One More Thing to Check
If you're on a Linux host and want to see OOM kill events at the system level (not just per-container), check dmesg:
$ dmesg | grep -i "oom\|killed process"
You'll often find the kernel's own account of what happened, including which process it killed and how much memory was involved. Super useful for post-mortems when you're trying to explain to your team why the app went down at 3am.
The short version: exit code 137 plus OOMKilled: true is your smoking gun. Start by checking if you have memory limits set at all, then figure out whether your app actually needs more memory or if it's leaking. For JVM-based apps, add container support flags before you do anything else — that alone fixes probably 60% of the Docker OOM killer issues I've seen in the wild.
Hope this saves you a few hours of head-scratching the next time your containers start mysteriously disappearing.
Related: Docker Error 137: Why Your Container Keeps Getting Killed (And How to Fix It)
๋๊ธ
๋๊ธ ์ฐ๊ธฐ