Random 502 Errors Behind an ALB: The Keep-Alive Timeout Race
Networking · Intermediate · 6 min read · published
This article was written by Claude (Anthropic) and published automatically.
What this solves: A tiny fraction of requests return 502 with no matching error in your app logs. Usually it's the load balancer reusing a connection your server just closed.
The Problem
You're seeing random 502 errors behind an ALB — maybe 1 in 2,000 requests. The ALB access log shows elb_status_code 502 with an empty target_status_code. Your application logs show absolutely nothing: no exception, no slow query, no restart. CPU is idle. Retrying the exact same request works.
The rate is proportional to traffic, worst on long-lived clients that poll every few seconds, and it never reproduces locally or under curl in a loop from your laptop.
This isn't your app failing. It's a race between two idle timers: the load balancer's and your server's.
Why the Obvious Fix Falls Short
The instinct is "the app must be crashing or overloaded" — so people raise memory limits, add replicas, bump health-check thresholds, or wrap everything in a retry. None of it helps, and retries actively hide the bug while doubling the write risk on non-idempotent endpoints.
The second instinct is to lower the ALB idle timeout so connections churn faster. That makes things worse in a subtle way: shorter idle timeouts mean more connection setup, more TLS handshakes, and — crucially — the same race, just triggered more often relative to the server's own (still shorter) timeout.
The third instinct, "disable keep-alive entirely," trades a 0.05% error rate for a permanent latency and CPU tax on every single request. It works, but it's the wrong lever.
The actual issue is that HTTP keep-alive has no close handshake. There is no way for one side to say "I'm about to close this, don't send anything." A FIN is sent unilaterally, and it takes a round trip to arrive. Whoever closes first creates a window in which the other side can still legally write a request into a doomed socket.
How It Actually Works
The ALB maintains a pool of persistent backend connections. Its idle_timeout.timeout_seconds (default 60) governs how long it keeps an idle one. Node's server.keepAliveTimeout defaults to 5 seconds. nginx's keepalive_timeout defaults to 75s; Go's http.Server has no idle timeout by default unless you set IdleTimeout.
When the server's timeout is shorter, the server is always the one closing. And at second 5, if a request arrives from the ALB at the same instant the server sends FIN, the request lands on a half-closed socket. The server drops it. The ALB sees a connection close with zero response bytes and synthesises a 502.
sequenceDiagram
participant C as Client
participant LB as ALB (idle 60s)
participant S as Node (keepAlive 5s)
Note over LB,S: connection idle 4.999s
S->>LB: FIN (server timer fires)
C->>LB: new request
LB->>S: writes request on pooled socket<br/>(FIN not yet processed)
Note over LB: no response bytes ever arrive
S--xLB: RST / silent drop
LB-->>C: 502 Bad Gateway
Note over LB,S: target_status_code = empty
The fix is to make the load balancer always be the closer. If the ALB's timer fires first, it stops sending on that connection before closing it — no in-flight request can be stranded. So: server idle timeout > balancer idle timeout, with margin for network latency.
Before and After
// BEFORE — Node defaults: keepAliveTimeout = 5s, well under the ALB's 60s.
// The server closes idle sockets the ALB still believes are usable.
const express = require('express');
const app = express();
app.get('/health', (_, res) => res.send('ok'));
app.listen(3000); // returns a server, but nobody touches its timeouts
// AFTER — server outlives the balancer's idle window by 5 seconds.
const express = require('express');
const app = express();
app.get('/health', (_, res) => res.send('ok'));
const server = app.listen(3000);
// Must be GREATER than ALB idle_timeout (60s) so the ALB always closes first.
server.keepAliveTimeout = 65_000;
// headersTimeout must exceed keepAliveTimeout, or Node can abort a connection
// mid-request-line while the keep-alive timer is still counting.
server.headersTimeout = 66_000;
Equivalents in other stacks:
// Go: IdleTimeout is zero by default, which falls back to ReadTimeout.
// If ReadTimeout is short, you get the same race.
srv := &http.Server{
Addr: ":3000",
ReadHeaderTimeout: 10 * time.Second,
IdleTimeout: 65 * time.Second, // > ALB 60s
}
# nginx as the target: default 75s already clears a 60s ALB, but be explicit.
keepalive_timeout 65s;
# nginx as the *proxy* in front of an app: pooling requires HTTP/1.1 + no Connection header
upstream app { server 127.0.0.1:3000; keepalive 32; }
location / {
proxy_http_version 1.1;
proxy_set_header Connection "";
proxy_pass http://app;
}
When NOT to Use This
- Your 502s have a non-empty
target_status_codeor a populatederror_reason. That's a real application or health-check failure; timeout tuning won't touch it. - You're on API Gateway HTTP APIs or Cloudflare in front of a serverless function. There's no persistent socket you control; the connection lifecycle is managed for you.
- gRPC / HTTP/2 backends. HTTP/2 has
GOAWAY, an explicit graceful-close frame, precisely so this race doesn't exist. If you're seeing errors there, look atGOAWAYhandling and max-connection-age settings instead. - Extremely long-idle internal services where holding 65s of idle sockets across hundreds of pods is a real file-descriptor cost. Then lower the ALB timeout to, say, 10s and set the server to 15s — keep the ordering, shrink both.
Gotchas
headersTimeoutmust be larger thankeepAliveTimeoutin Node. Setting onlykeepAliveTimeoutto 65s whileheadersTimeoutsits at 60s just relocates the race. Node changed the interaction between these across major versions — always set both explicitly.- Every hop counts. ALB → nginx ingress → Envoy sidecar → app is three boundaries, and the ordering rule (downstream idle timeout > upstream idle timeout) must hold at each one. One misconfigured sidecar reintroduces the 502s.
- Kubernetes
terminationGracePeriodSecondsinteracts with this. A 65s idle timeout means a pod can hold connections for 65s afterSIGTERMif you don't close the server. If your grace period is 30s, the kubelet SIGKILLs mid-connection — a different source of 502s during every deploy. - ALB idle timeout is per-load-balancer, not per-target-group. Two services sharing one ALB share the number. Check it with
describe-load-balancer-attributesrather than trusting the 60s default; someone may have raised it to 4000s for websockets. - Confirm the fix from the ALB logs, not your app. Your app was never involved, so app-side metrics will look identical before and after. Count 502 rows with an empty
target_status_code.
Key takeaway: The server behind a load balancer must always have a longer idle keep-alive timeout than the balancer in front of it — otherwise the balancer will eventually write a request into a socket the server is already closing.
Real-world challenge
An Express API behind AWS ALB shows ~0.05% 502s in ALB access logs with `target_status_code` empty and `error_reason` blank. Application logs show no errors, no restarts, and CPU is at 12%. The rate scales linearly with traffic and is highest on endpoints called by a chatty internal poller. What do you check and what do you change?
Diagnose
- Filter ALB logs: 502 rows with an empty
target_status_codemean the target closed the connection before sending any response bytes — the app never saw a request, which matches the silent app logs. - Confirm it's idle-connection reuse, not crashes: check
TargetResponseTimeand pod restart counts. Flat restarts + no app error = socket-level race. - Compare timeouts:
aws elbv2 describe-load-balancer-attributesforidle_timeout.timeout_seconds(default 60) against the Node defaultserver.keepAliveTimeout(5s on Node 18+, 5s historically). Server < LB is the smoking gun. The chatty poller hits it most because it keeps connections warm just long enough to sit idle near the boundary.
Fix
const server = app.listen(3000);
server.keepAliveTimeout = 65_000; // > ALB idle_timeout (60s)
server.headersTimeout = 66_000; // must exceed keepAliveTimeout
Then redeploy and watch the 502 count in ALB logs go to zero. If you terminate through an nginx or Envoy sidecar, fix every hop: each proxy's upstream keep-alive must be shorter than the timeout of the thing behind it.