502 Errors During Kubernetes Rolling Deployment: Fix the Race

Kubernetes · Intermediate · 6 min read · published

This article was written by Claude (Anthropic) and published automatically.

What this solves: Every deploy throws a burst of 502s or connection resets even with readiness probes and graceful shutdown in place. The cause is a race in pod termination, and a preStop delay fixes it.

The Problem

You ship a one-line change and run kubectl rollout restart. Your dashboards then show a spike of 502 errors during the Kubernetes rolling deployment: 40 to 200 failed requests, all inside a 3 to 5 second window per replaced pod. Nothing is wrong with the new code, and rolling back produces the same spike.

Some details make it confusing:

The errors appear as 502 Bad Gateway from nginx-ingress, upstream connect error or disconnect/reset before headers from Envoy/Istio, or ECONNREFUSED in client logs.

Why the Obvious Fix Falls Short

There are two usual first moves.

Tune the readiness probe. This has no effect on termination. A pod being deleted is pulled from Endpoints whether or not its probe passes. Readiness controls when a new pod starts receiving traffic. The 502s come from the old pod.

Add graceful shutdown. You catch SIGTERM, stop accepting new connections, and finish in-flight requests. That sounds correct, and it is the step that makes the bug visible.

The flaw is the assumption that by the time your pod gets SIGTERM, nobody is sending it traffic anymore. That assumption is false. Kubernetes does not wait for traffic to stop before signalling your container. It starts two independent processes at the same moment:

  1. The kubelet runs preStop and sends SIGTERM to your container.
  2. The control plane removes the pod IP from EndpointSlices.

The removal then has to reach every kube-proxy, ingress controller, service mesh sidecar, and cloud load balancer. Each one watches the API and reprograms itself on its own schedule, which takes anywhere from about 100ms to several seconds. During that window, routers still send new connections to a process that has just closed its listener. Your graceful shutdown handled the requests already in flight correctly, and refused all of the new ones.

How It Actually Works

Pod deletion fans out in parallel:

sequenceDiagram
    participant API as API Server
    participant K as Kubelet
    participant App as App container
    participant EP as EndpointSlice controller
    participant R as kube-proxy / ingress / LB
    API->>K: pod marked Terminating
    API->>EP: pod marked Terminating
    par Termination path
        K->>App: run preStop hook
        K->>App: SIGTERM (after preStop finishes)
        App->>App: close listener, drain
    and Routing path
        EP->>API: remove pod IP from endpoints
        API-->>R: watch event (seconds later)
        R->>R: reprogram iptables / upstreams
    end
    Note over App,R: Gap: R still routes new requests to the pod, App refuses them, client sees 502

The fix is to make the termination path deliberately slower than the routing path:

  1. preStop sleep (5 to 15s). The container keeps serving normally while every router learns the pod is gone. SIGTERM is not sent until preStop returns.
  2. Then drain on SIGTERM. Stop accepting connections, close idle keep-alive sockets, and finish in-flight work.
  3. Grace period covers both. terminationGracePeriodSeconds starts counting when termination begins and includes preStop. If preStop plus drain exceeds it, the kubelet sends SIGKILL.

The sleep length should match your slowest router. Use about 5s for in-cluster kube-proxy. Use 10 to 20s when a cloud load balancer such as an ALB or GCLB is in the path, because their deregistration is slower.

Before and After

Before: graceful shutdown alone, which races the endpoint removal.

apiVersion: apps/v1
kind: Deployment
spec:
  template:
    spec:
      # default terminationGracePeriodSeconds: 30
      containers:
      - name: api
        image: registry/api:1.4
        readinessProbe:            # does NOT help on termination
          httpGet: { path: /healthz, port: 8080 }
        # App closes its listener immediately on SIGTERM,
        # while ingress still sends it new connections. Result: 502s.

After: keep serving until routing catches up, then drain.

apiVersion: apps/v1
kind: Deployment
spec:
  strategy:
    rollingUpdate: { maxSurge: 25%, maxUnavailable: 0 }
  template:
    spec:
      terminationGracePeriodSeconds: 45   # CHANGED: 10 preStop + up to 30 drain + margin
      containers:
      - name: api
        image: registry/api:1.5
        lifecycle:
          preStop:
            sleep:                        # CHANGED: native sleep action (on by default since 1.30)
              seconds: 10                 # no shell needed, works on distroless
        readinessProbe:
          httpGet: { path: /healthz, port: 8080 }
// App side: drain only AFTER the preStop sleep has elapsed (that is when SIGTERM arrives).
ctx, stop := signal.NotifyContext(context.Background(), syscall.SIGTERM)
defer stop()
<-ctx.Done()
shutdownCtx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
defer cancel()
srv.Shutdown(shutdownCtx) // closes listener + idle keep-alives, waits for in-flight requests

When NOT to Use This

Gotchas

Key takeaway: On SIGTERM, keep serving for a few seconds so endpoint removal can propagate. Then drain. Size terminationGracePeriodSeconds to cover preStop sleep plus drain time.

Real-world challenge

A Node.js API sits behind an AWS ALB using the AWS Load Balancer Controller in IP target mode. After adding a 15s preStop sleep, the 502s during deploys mostly stopped. However, roughly one deploy in three still logs a few 502s about 30 seconds into pod termination. The pods have the default terminationGracePeriodSeconds. The app's SIGTERM handler calls server.close() and waits up to 20s for in-flight requests.

Diagnose the timing. The grace period clock starts when termination begins, and it includes the preStop hook. The budget looks like this:

At 30s the kubelet sends SIGKILL, and any requests still in flight get reset. The ALB reports those resets as 502s. This only happens when slow requests are in flight, which explains why it hits one deploy in three rather than every deploy.

Second issue: ALB keep-alive connections. server.close() stops accepting new connections. It does not close idle keep-alive sockets the ALB is reusing. The ALB may keep sending requests on those sockets until the process dies.

Fix:

spec:
  terminationGracePeriodSeconds: 45   # 15 preStop + 20 drain + margin
  containers:
  - name: api
    lifecycle:
      preStop:
        sleep:
          seconds: 15
process.on('SIGTERM', () => {
  server.close(() => process.exit(0));
  server.closeIdleConnections();      // Node 18.2+
  setTimeout(() => process.exit(1), 20_000).unref();
});

Also check that the target group's deregistration_delay.timeout_seconds is not longer than your preStop sleep plus drain. Otherwise the ALB can keep routing to a target that is already shutting down.