HTTP Request Times Out on Long-Running Tasks: The 202 + Job Store Pattern

Architecture · Intermediate · 7 min read · published

This article was written by Claude (Anthropic) and published automatically.

What this solves: Your endpoint does real work — a report, a video encode, an LLM call — and the proxy kills it at 30 or 60 seconds. Here's the async job architecture that replaces it.

The Forces at Play

An HTTP request times out on a long-running task and there is nothing you can do inside the handler to fix it. The work genuinely takes 90 seconds — a 40-page PDF render, a Whisper transcription, a bulk CSV import — but an ALB idles you out at 60s, API Gateway hard-caps at 29s, Cloudflare at 100s, and the browser gives up somewhere too. You'll see 504 Gateway Timeout, upstream request timeout, or a bare net::ERR_EMPTY_RESPONSE with no server log at all, because your handler was still running happily when the proxy hung up on it.

The forces pulling against each other:

The resolution is to stop treating the HTTP connection as the container for the work. The connection becomes a way to start work and ask about work. The work itself lives in a durable job record.

The Shape

flowchart TD
    C["Client"]
    subgraph edge["Synchronous edge (< 100ms)"]
      API["API: POST /exports<br/>validate + persist + enqueue<br/>returns 202 + job id"]
      STAT["API: GET /exports/{id}<br/>reads job record only"]
    end
    DB[("Job store<br/>id, status, attempts,<br/>lease_expires_at, result_url")]
    Q[["Queue<br/>at-least-once delivery"]]
    W1["Worker 1"]
    W2["Worker N"]
    BLOB[("Object store<br/>large results")]
    DLQ[["Dead-letter queue"]]
    REAP["Reaper<br/>expired leases -> queued"]

    C -->|1. POST + Idempotency-Key| API
    API -->|2. INSERT job status=queued| DB
    API -->|3. enqueue job id| Q
    API -->|4. 202 Accepted, Location| C
    Q --> W1
    Q --> W2
    W1 -->|5. claim: queued -> processing<br/>heartbeat lease| DB
    W1 -->|6. write bytes| BLOB
    W1 -->|7. status=done, result_url| DB
    W1 -.->|poison message| DLQ
    REAP -->|abandoned jobs| DB
    C -->|8. poll / SSE| STAT
    STAT --> DB
    STAT -->|9. 303 redirect to presigned URL| BLOB

The key structural fact: the queue carries only an ID, never the payload or the result. The job store is the single source of truth; the queue is just a wake-up signal that happens to be durable.

How Data Flows Through It

One export, end to end.

1. Submit. POST /exports with Idempotency-Key: 8f3c.... The API validates input, then does one transaction: insert the job row (unique index on the idempotency key) and enqueue the ID. If the key already exists, it returns the existing job unchanged. Total handler time: 20ms.

2. Respond. 202 Accepted, with Location: /exports/j_9021 and a suggested poll interval. The connection closes. No proxy timeout can touch this.

HTTP/1.1 202 Accepted
Location: /exports/j_9021
Retry-After: 2
Content-Type: application/json

{ "id": "j_9021", "status": "queued", "progress": 0 }

3. Claim. A worker pulls the ID and atomically transitions the row queued -> processing, stamping worker_id and lease_expires_at = now() + 60s. The conditional update is what makes duplicate queue deliveries harmless — the second worker's UPDATE ... WHERE status='queued' affects zero rows and it drops the message.

4. Work and heartbeat. Every 15 seconds the worker extends the lease and writes progress. That single column is what lets the UI show "page 12 of 40" instead of a shrug.

5. Publish. The rendered PDF goes to object storage first. Then the job row flips to done with result_url. Ordering matters: a done row whose bytes don't exist yet is a bug the client sees; bytes with no done row are just garbage a lifecycle rule cleans up.

6. Collect. The client polls GET /exports/j_9021. On done, the endpoint issues a short-lived presigned URL — the API never proxies the payload.

The job record is the contract between all of them:

{
  "id": "j_9021",
  "type": "pdf_export",
  "status": "queued|processing|done|failed|cancelled",
  "progress": 0.35,
  "attempts": 1,
  "worker_id": "worker-7c2",
  "lease_expires_at": "2025-03-04T10:14:22Z",
  "result_url": null,
  "error": null,
  "idempotency_key": "8f3c...",
  "created_at": "2025-03-04T10:13:02Z",
  "updated_at": "2025-03-04T10:13:47Z"
}

What Each Piece Owns

Submit endpoint owns validation, idempotency, and durability of the intent. It deliberately does not own the work, and must never do "just a little" of it inline — the moment it does, you're back under the timeout ceiling.

Job store owns state, progress, attempt count, and the lease. It does not own the result bytes; anything above a few KB belongs in object storage, or your jobs table becomes an accidental blob store with terrible vacuum behaviour.

Queue owns delivery and retry timing. It does not own job state. Resist encoding status in queue semantics — "is it done?" must be answerable without touching the queue, because queues are write-optimised and generally unqueryable.

Worker owns execution and heartbeating. It does not own deciding whether it's allowed to run — the conditional claim in the job store decides that. And it must not assume it runs exactly once.

Status endpoint owns a cheap point read and result handoff. It does not own streaming the file.

Reaper owns liveness: any processing job whose lease expired goes back to queued. This is the only component that can tell "slow" from "dead", and systems that skip it end up with jobs stuck in processing forever.

Where It Breaks Down

Polling becomes the load. 5,000 concurrent clients polling every second is 5,000 rps of point reads on your job store — often more traffic than the real work. First mitigations: honour Retry-After with exponential backoff (1s, 2s, 4s, cap 15s), serve status from a cache with a 1-second TTL, or move to SSE/WebSocket for the hot path with polling as fallback.

At-least-once bites the side effects. Lease expiry during a genuinely slow job means a second worker starts while the first is still running. If the job charges a card or sends email, the conditional claim isn't enough — the side effect itself needs its own idempotency key. Assume every job body can run twice.

Backlog without backpressure. The submit endpoint stays fast while the queue grows to 400,000 items and p99 completion goes from 90 seconds to four hours. Nothing errors; it just quietly stops being useful. Alert on queue age of the oldest message, not queue depth — depth tells you nothing about whether you're keeping up.

Poison jobs. One malformed 900MB CSV that OOM-kills a worker gets redelivered forever, taking down a worker each time. You need a hard attempts cap that routes to a dead-letter queue, and a human-visible list of failed jobs — otherwise failures are invisible until a customer complains.

Operational surface. You've gone from one deployable to three (API, workers, reaper) plus a queue. The workers are the burden: they need their own autoscaling signal (queue age), graceful SIGTERM handling to release leases, and separate dashboards. Budget for that.

When This Is Overkill

If p99 work time is under ~10 seconds, don't build this. Raise the proxy's idle timeout, keep the request synchronous, and move on. Two smaller designs also cover a lot of ground:

The signals you've outgrown them: you're seeing 504 Gateway Timeout in production and can't raise the limit because it's a platform hard cap; users retry and cause duplicate work; a rolling deploy loses in-flight work someone paid for; or you need to answer "where is my export?" after the browser tab closed. Any one of those, and the job record has to exist.

Key takeaway: When work outlives an HTTP timeout, return a job ID immediately and make the job record — not the connection — the source of truth for progress, results, and retries.

Real-world challenge

Your video transcode service has 40 jobs sitting in status `processing` — some for six hours, with no worker touching them. The queue is empty, workers are idle, and the client UIs spin forever. It started after you enabled aggressive spot-instance reclamation on the worker node group. Restarting workers doesn't recover the jobs.

Diagnose. processing is a state only a live worker can justify, but nothing in your design proves a worker is still alive. When a spot node is reclaimed the worker gets SIGKILL: the queue message eventually reappears (visibility timeout), but if your worker deleted the message before finishing, or if the job row was already flipped to processing and never flipped back, the job is orphaned. There is no timestamp that lets anyone tell a stuck job from a slow one.

Check: SELECT id, updated_at, now() - updated_at FROM jobs WHERE status='processing' ORDER BY 2; If updated_at is hours old, nobody is working them.

Fix — make processing a lease, not a label.

ALTER TABLE jobs ADD COLUMN lease_expires_at timestamptz;

-- worker heartbeats every 15s while working
UPDATE jobs SET lease_expires_at = now() + interval '60 seconds',
                progress = $2
 WHERE id = $1 AND worker_id = $3;

-- reaper, every 30s
UPDATE jobs SET status='queued', attempts = attempts + 1,
                worker_id = NULL, lease_expires_at = NULL
 WHERE status='processing' AND lease_expires_at < now()
   AND attempts < 3;

Jobs past attempts >= 3 go to failed with a reason so they surface in a dead-letter view instead of silently looping.

Also: delete/ack the queue message only after the result is committed, and handle SIGTERM to release the lease immediately on graceful drain so reclaimed nodes recover in seconds, not a minute.