502 Errors During Kubernetes Rolling Deployment: Fix the Race
Kubernetes · Intermediate · 6 min read · published
This article was written by Claude (Anthropic) and published automatically.
What this solves: Every deploy throws a burst of 502s or connection resets even with readiness probes and graceful shutdown in place. The cause is a race in pod termination, and a preStop delay fixes it.
The Problem
You ship a one-line change and run kubectl rollout restart. Your dashboards then show a spike of 502 errors during the Kubernetes rolling deployment: 40 to 200 failed requests, all inside a 3 to 5 second window per replaced pod. Nothing is wrong with the new code, and rolling back produces the same spike.
Some details make it confusing:
- It happens at 50 req/s and at 5,000 req/s. Only the count changes.
- It happens even though you have readiness probes.
- It happens even though your app handles SIGTERM.
The errors appear as 502 Bad Gateway from nginx-ingress, upstream connect error or disconnect/reset before headers from Envoy/Istio, or ECONNREFUSED in client logs.
Why the Obvious Fix Falls Short
There are two usual first moves.
Tune the readiness probe. This has no effect on termination. A pod being deleted is pulled from Endpoints whether or not its probe passes. Readiness controls when a new pod starts receiving traffic. The 502s come from the old pod.
Add graceful shutdown. You catch SIGTERM, stop accepting new connections, and finish in-flight requests. That sounds correct, and it is the step that makes the bug visible.
The flaw is the assumption that by the time your pod gets SIGTERM, nobody is sending it traffic anymore. That assumption is false. Kubernetes does not wait for traffic to stop before signalling your container. It starts two independent processes at the same moment:
- The kubelet runs preStop and sends SIGTERM to your container.
- The control plane removes the pod IP from EndpointSlices.
The removal then has to reach every kube-proxy, ingress controller, service mesh sidecar, and cloud load balancer. Each one watches the API and reprograms itself on its own schedule, which takes anywhere from about 100ms to several seconds. During that window, routers still send new connections to a process that has just closed its listener. Your graceful shutdown handled the requests already in flight correctly, and refused all of the new ones.
How It Actually Works
Pod deletion fans out in parallel:
sequenceDiagram
participant API as API Server
participant K as Kubelet
participant App as App container
participant EP as EndpointSlice controller
participant R as kube-proxy / ingress / LB
API->>K: pod marked Terminating
API->>EP: pod marked Terminating
par Termination path
K->>App: run preStop hook
K->>App: SIGTERM (after preStop finishes)
App->>App: close listener, drain
and Routing path
EP->>API: remove pod IP from endpoints
API-->>R: watch event (seconds later)
R->>R: reprogram iptables / upstreams
end
Note over App,R: Gap: R still routes new requests to the pod, App refuses them, client sees 502
The fix is to make the termination path deliberately slower than the routing path:
- preStop sleep (5 to 15s). The container keeps serving normally while every router learns the pod is gone. SIGTERM is not sent until preStop returns.
- Then drain on SIGTERM. Stop accepting connections, close idle keep-alive sockets, and finish in-flight work.
- Grace period covers both.
terminationGracePeriodSecondsstarts counting when termination begins and includes preStop. If preStop plus drain exceeds it, the kubelet sends SIGKILL.
The sleep length should match your slowest router. Use about 5s for in-cluster kube-proxy. Use 10 to 20s when a cloud load balancer such as an ALB or GCLB is in the path, because their deregistration is slower.
Before and After
Before: graceful shutdown alone, which races the endpoint removal.
apiVersion: apps/v1
kind: Deployment
spec:
template:
spec:
# default terminationGracePeriodSeconds: 30
containers:
- name: api
image: registry/api:1.4
readinessProbe: # does NOT help on termination
httpGet: { path: /healthz, port: 8080 }
# App closes its listener immediately on SIGTERM,
# while ingress still sends it new connections. Result: 502s.
After: keep serving until routing catches up, then drain.
apiVersion: apps/v1
kind: Deployment
spec:
strategy:
rollingUpdate: { maxSurge: 25%, maxUnavailable: 0 }
template:
spec:
terminationGracePeriodSeconds: 45 # CHANGED: 10 preStop + up to 30 drain + margin
containers:
- name: api
image: registry/api:1.5
lifecycle:
preStop:
sleep: # CHANGED: native sleep action (on by default since 1.30)
seconds: 10 # no shell needed, works on distroless
readinessProbe:
httpGet: { path: /healthz, port: 8080 }
// App side: drain only AFTER the preStop sleep has elapsed (that is when SIGTERM arrives).
ctx, stop := signal.NotifyContext(context.Background(), syscall.SIGTERM)
defer stop()
<-ctx.Done()
shutdownCtx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
defer cancel()
srv.Shutdown(shutdownCtx) // closes listener + idle keep-alives, waits for in-flight requests
When NOT to Use This
- Clients that retry idempotently. If every caller is a gRPC client with retry policies or a mesh with automatic retries on connection failure, the race may already be invisible. Measure before you add 10s to every pod termination.
- Batch workers and queue consumers. They receive no routed traffic, so the sleep only slows deploys. Stop pulling from the queue on SIGTERM instead.
- Very long-lived connections such as WebSockets or streaming gRPC. A preStop sleep fixes new connections only. Existing ones need application-level draining: send a goaway or reconnect message, and use a grace period sized to how long you are willing to wait.
- Fast scale-down matters more than errors. A sleep slows every eviction, including node drains. If you are on spot instances with a 2-minute reclaim notice, keep the whole budget well inside that limit.
Gotchas
exec: ["sleep", "10"]on distroless or scratch images fails silently. There is nosleepbinary, so the hook errors out and SIGTERM is sent immediately. UsepreStop.sleep(Kubernetes 1.30+) or compile the sleep into your binary.- The grace period includes preStop. With a 30s default and a 20s sleep, you have only 10s to drain before SIGKILL.
- PID 1 is a shell. With
CMD npm startorsh -c ..., SIGTERM goes to the shell and is not forwarded to your app, so the app is killed at the deadline. Use exec-formCMD ["node", "server.js"]ortini. - Keep-alive sockets outlive
close(). Node'sserver.close()leaves idle keep-alive connections open, and upstream proxies keep reusing them. Callserver.closeIdleConnections()(Node 18.2+), or sendConnection: closeonce draining starts. - Cloud LB deregistration delay. The ALB's
deregistration_delay(300s by default) can keep a target in the draining state much longer than your pod lives. Lower it to roughly preStop plus drain time. maxUnavailable: 25%with few replicas. With 2 pods, one being drained means 50% of capacity is gone, and the overload looks like the same 502 spike. UsemaxUnavailable: 0with a surge instead.
Key takeaway: On SIGTERM, keep serving for a few seconds so endpoint removal can propagate. Then drain. Size terminationGracePeriodSeconds to cover preStop sleep plus drain time.
Real-world challenge
A Node.js API sits behind an AWS ALB using the AWS Load Balancer Controller in IP target mode. After adding a 15s preStop sleep, the 502s during deploys mostly stopped. However, roughly one deploy in three still logs a few 502s about 30 seconds into pod termination. The pods have the default terminationGracePeriodSeconds. The app's SIGTERM handler calls server.close() and waits up to 20s for in-flight requests.
Diagnose the timing. The grace period clock starts when termination begins, and it includes the preStop hook. The budget looks like this:
- 15s preStop sleep
- up to 20s drain
- = 35s total, against a 30s default
At 30s the kubelet sends SIGKILL, and any requests still in flight get reset. The ALB reports those resets as 502s. This only happens when slow requests are in flight, which explains why it hits one deploy in three rather than every deploy.
Second issue: ALB keep-alive connections. server.close() stops accepting new connections. It does not close idle keep-alive sockets the ALB is reusing. The ALB may keep sending requests on those sockets until the process dies.
Fix:
spec:
terminationGracePeriodSeconds: 45 # 15 preStop + 20 drain + margin
containers:
- name: api
lifecycle:
preStop:
sleep:
seconds: 15
process.on('SIGTERM', () => {
server.close(() => process.exit(0));
server.closeIdleConnections(); // Node 18.2+
setTimeout(() => process.exit(1), 20_000).unref();
});
Also check that the target group's deregistration_delay.timeout_seconds is not longer than your preStop sleep plus drain. Otherwise the ALB can keep routing to a target that is already shutting down.