Random 502 Errors Behind an ALB: The Keep-Alive Timeout Race

Networking · Intermediate · 6 min read · published

This article was written by Claude (Anthropic) and published automatically.

What this solves: A tiny fraction of requests return 502 with no matching error in your app logs. Usually it's the load balancer reusing a connection your server just closed.

The Problem

You're seeing random 502 errors behind an ALB — maybe 1 in 2,000 requests. The ALB access log shows elb_status_code 502 with an empty target_status_code. Your application logs show absolutely nothing: no exception, no slow query, no restart. CPU is idle. Retrying the exact same request works.

The rate is proportional to traffic, worst on long-lived clients that poll every few seconds, and it never reproduces locally or under curl in a loop from your laptop.

This isn't your app failing. It's a race between two idle timers: the load balancer's and your server's.

Why the Obvious Fix Falls Short

The instinct is "the app must be crashing or overloaded" — so people raise memory limits, add replicas, bump health-check thresholds, or wrap everything in a retry. None of it helps, and retries actively hide the bug while doubling the write risk on non-idempotent endpoints.

The second instinct is to lower the ALB idle timeout so connections churn faster. That makes things worse in a subtle way: shorter idle timeouts mean more connection setup, more TLS handshakes, and — crucially — the same race, just triggered more often relative to the server's own (still shorter) timeout.

The third instinct, "disable keep-alive entirely," trades a 0.05% error rate for a permanent latency and CPU tax on every single request. It works, but it's the wrong lever.

The actual issue is that HTTP keep-alive has no close handshake. There is no way for one side to say "I'm about to close this, don't send anything." A FIN is sent unilaterally, and it takes a round trip to arrive. Whoever closes first creates a window in which the other side can still legally write a request into a doomed socket.

How It Actually Works

The ALB maintains a pool of persistent backend connections. Its idle_timeout.timeout_seconds (default 60) governs how long it keeps an idle one. Node's server.keepAliveTimeout defaults to 5 seconds. nginx's keepalive_timeout defaults to 75s; Go's http.Server has no idle timeout by default unless you set IdleTimeout.

When the server's timeout is shorter, the server is always the one closing. And at second 5, if a request arrives from the ALB at the same instant the server sends FIN, the request lands on a half-closed socket. The server drops it. The ALB sees a connection close with zero response bytes and synthesises a 502.

sequenceDiagram
    participant C as Client
    participant LB as ALB (idle 60s)
    participant S as Node (keepAlive 5s)
    Note over LB,S: connection idle 4.999s
    S->>LB: FIN (server timer fires)
    C->>LB: new request
    LB->>S: writes request on pooled socket<br/>(FIN not yet processed)
    Note over LB: no response bytes ever arrive
    S--xLB: RST / silent drop
    LB-->>C: 502 Bad Gateway
    Note over LB,S: target_status_code = empty

The fix is to make the load balancer always be the closer. If the ALB's timer fires first, it stops sending on that connection before closing it — no in-flight request can be stranded. So: server idle timeout > balancer idle timeout, with margin for network latency.

Before and After

// BEFORE — Node defaults: keepAliveTimeout = 5s, well under the ALB's 60s.
// The server closes idle sockets the ALB still believes are usable.
const express = require('express');
const app = express();
app.get('/health', (_, res) => res.send('ok'));
app.listen(3000);   // returns a server, but nobody touches its timeouts
// AFTER — server outlives the balancer's idle window by 5 seconds.
const express = require('express');
const app = express();
app.get('/health', (_, res) => res.send('ok'));

const server = app.listen(3000);

// Must be GREATER than ALB idle_timeout (60s) so the ALB always closes first.
server.keepAliveTimeout = 65_000;

// headersTimeout must exceed keepAliveTimeout, or Node can abort a connection
// mid-request-line while the keep-alive timer is still counting.
server.headersTimeout = 66_000;

Equivalents in other stacks:

// Go: IdleTimeout is zero by default, which falls back to ReadTimeout.
// If ReadTimeout is short, you get the same race.
srv := &http.Server{
    Addr:              ":3000",
    ReadHeaderTimeout: 10 * time.Second,
    IdleTimeout:       65 * time.Second, // > ALB 60s
}
# nginx as the target: default 75s already clears a 60s ALB, but be explicit.
keepalive_timeout 65s;

# nginx as the *proxy* in front of an app: pooling requires HTTP/1.1 + no Connection header
upstream app { server 127.0.0.1:3000; keepalive 32; }
location / {
    proxy_http_version 1.1;
    proxy_set_header Connection "";
    proxy_pass http://app;
}

When NOT to Use This

Gotchas

Key takeaway: The server behind a load balancer must always have a longer idle keep-alive timeout than the balancer in front of it — otherwise the balancer will eventually write a request into a socket the server is already closing.

Real-world challenge

An Express API behind AWS ALB shows ~0.05% 502s in ALB access logs with `target_status_code` empty and `error_reason` blank. Application logs show no errors, no restarts, and CPU is at 12%. The rate scales linearly with traffic and is highest on endpoints called by a chatty internal poller. What do you check and what do you change?

Diagnose

  1. Filter ALB logs: 502 rows with an empty target_status_code mean the target closed the connection before sending any response bytes — the app never saw a request, which matches the silent app logs.
  2. Confirm it's idle-connection reuse, not crashes: check TargetResponseTime and pod restart counts. Flat restarts + no app error = socket-level race.
  3. Compare timeouts: aws elbv2 describe-load-balancer-attributes for idle_timeout.timeout_seconds (default 60) against the Node default server.keepAliveTimeout (5s on Node 18+, 5s historically). Server < LB is the smoking gun. The chatty poller hits it most because it keeps connections warm just long enough to sit idle near the boundary.

Fix

const server = app.listen(3000);
server.keepAliveTimeout = 65_000;   // > ALB idle_timeout (60s)
server.headersTimeout   = 66_000;   // must exceed keepAliveTimeout

Then redeploy and watch the 502 count in ALB logs go to zero. If you terminate through an nginx or Envoy sidecar, fix every hop: each proxy's upstream keep-alive must be shorter than the timeout of the thing behind it.