Webhook Delivered Twice? Build an Idempotent Receiver That Dedupes Safely
Architecture · Intermediate · 7 min read · published
This article was written by Claude (Anthropic) and published automatically.
What this solves: Providers retry webhooks on timeouts, so the same event hits your endpoint twice and you double-charge or double-email. Here's the receiver shape that stops it.
The Forces at Play
A webhook gets delivered twice and your customer gets charged twice, or emailed twice, or their order ships twice. This isn't a bug in the provider — every serious webhook sender (payment processors, Git hosts, messaging platforms) guarantees at-least-once delivery, because the alternative is silently losing events when your endpoint blips.
The tension is this: the sender cannot distinguish "your server never got it" from "your server got it, did the work, and then the TCP connection died before the 200 came back". From the sender's side both look identical — a timeout. Given that ambiguity it has exactly two choices: retry (risk duplicates) or don't (risk loss). It picks duplicates, because duplicates are a problem you can solve and loss isn't.
So the burden shifts to you. Three forces pull against each other in the receiver:
- Fast acknowledgement. The longer your handler runs, the higher the chance of a timeout-triggered redelivery. Slow handlers manufacture duplicates.
- Durability before ack. You can't return 200 until the event is somewhere that survives a process restart, or you've told the sender to stop retrying something you then dropped.
- Exactly-once side effects. The external world — Stripe charges, emails, Slack messages — has no rollback. Idempotency has to be enforced at the boundary, not wished for.
The pattern that resolves these is the idempotent receiver: a thin, fast ingest tier whose only job is to durably claim an event, plus an async worker tier that does the real work.
The Shape
flowchart TB
P[Provider<br/>at-least-once sender]
subgraph Ingest["Ingest tier — must finish in <100ms"]
V[Signature verify<br/>+ raw body hash]
C{{"INSERT event_id<br/>ON CONFLICT DO NOTHING"}}
end
L[(webhook_events<br/>UNIQUE event_id<br/>status, payload)]
Q[[Durable queue<br/>job key = event_id]]
subgraph Work["Worker tier — slow, retryable"]
H[Handler]
E[External effect<br/>idempotency-key: event_id]
end
D[(Domain tables)]
P -->|POST attempt 1| V
P -.->|POST attempt 2 retry| V
V --> C
C -->|row inserted: first sighting| L
C -->|conflict: duplicate| R200b["200 OK — no-op"]
L --> Q
C --> R200a["200 OK"]
Q --> H
H --> E
H --> D
H -->|status = processed| L
H -.->|throw| Q
The centre of gravity is the ON CONFLICT DO NOTHING insert. That single atomic operation is the whole dedupe mechanism. Everything else is plumbing around it.
Notice what is not in the ingest path: no business logic, no external calls, no joins. The ingest tier is deliberately boring so that it is deliberately fast.
How Data Flows Through It
Take a payment_intent.succeeded event for intent pi_9f2.
Attempt 1, 10:00:00.000. Provider POSTs. Ingest verifies the HMAC signature over the raw body (before any JSON parsing — reserialising changes bytes and breaks the signature). It extracts the provider's event id, evt_8Kq, and runs:
INSERT INTO webhook_events (event_id, source, payload, status)
VALUES ('evt_8Kq', 'payments', $1, 'received')
ON CONFLICT (event_id) DO NOTHING
RETURNING id;
A row comes back. This request owns the event. It enqueues a job keyed evt_8Kq and returns 200. Total elapsed: 22ms.
Attempt 1 worker, 10:00:00.4. Worker picks up the job, loads the payload, calls the ledger service to credit the account, marks status='processed'. Done.
Now the interesting path. Suppose instead the network hiccups and the 200 never reaches the provider. At 10:00:30 it retries the same evt_8Kq.
Ingest verifies the signature again, runs the same insert — conflict, zero rows returned. There is nothing to do. It returns 200 immediately, without enqueueing anything. The provider stops retrying. The worker never sees a second job.
And the nastiest path: both attempts arrive concurrently, three milliseconds apart. Two transactions race on the unique index. One commits, one hits the constraint. Postgres serialises this for you at the index level — there is no window where both succeed. This is precisely why a SELECT ... IF NOT EXISTS ... INSERT in application code is not equivalent: that version has a window, and under retry storms you will land in it.
What Each Piece Owns
Ingest tier owns authenticity and the claim. It verifies the signature, records the raw payload, and makes the atomic claim decision. It does not own correctness of the business outcome, validation of payload semantics, or ordering. It never inspects whether the event "makes sense" — a malformed event is still stored, then failed loudly by a worker where you have retries and alerting.
The event ledger table owns the answer to "have we seen this before?" and "what was the exact bytes we received?". It does not own domain state. Resist the urge to denormalise order status into it; the moment two things own the same fact you get drift.
The queue owns delivery to workers and backpressure. It does not own dedupe — the queue's own at-least-once semantics mean the same job can be delivered twice, which is fine because the worker is idempotent against domain state.
Worker owns the side effect and the retry policy. It does not own the HTTP response to the provider; that already went out. Crucially, the worker must be independently idempotent, because queue redelivery bypasses the ingest dedupe entirely. Pass event_id as the downstream idempotency key:
POST /v1/refunds
Idempotency-Key: evt_8Kq
Ordering is owned by nobody, and that's intentional. Providers do not guarantee order. Handlers must be commutative or version-checked — e.g. ignore an event whose occurred_at is older than the row's updated_at.
Where It Breaks Down
The ledger table becomes the bottleneck first. Every inbound webhook is a write to one table with a unique index, and you're keeping payloads. At a few thousand events/sec the index and the table bloat. Symptoms: ingest p99 creeps up, autovacuum falls behind. Fixes in order: stop storing full payloads in the same row (move to object storage, keep a pointer), partition by month, and aggressively drop partitions older than your provider's maximum retry window plus your investigation window — usually 30–90 days.
Dedupe keys expire. If you prune the ledger at 7 days but the provider replays a 30-day-old event during an incident recovery, your "idempotent" receiver re-processes it. Prune window must exceed the maximum replay window, not the typical one.
Partial failure between claim and enqueue. Ingest inserts the row, then crashes before the queue accepts the job. You returned nothing, so the provider retries — but now the insert conflicts and you no-op. The event is claimed and orphaned. Two mitigations: use the transactional-outbox trick (enqueue by inserting into the same DB transaction, with a poller draining it), or run a reconciliation sweeper that finds rows stuck in status='received' for more than N minutes and re-enqueues. Most teams need the sweeper regardless — it's also your poison-message safety net.
Providers that reuse or omit event ids. Some senders give you no stable id, only a payload. Then your dedupe key is a hash of the canonical payload plus a timestamp bucket — and two legitimately identical events (two $5 charges, same second) become indistinguishable. Push back on the provider first; if you can't, dedupe on a business key (order_id + status_transition) rather than payload bytes.
Operational burden concentrates in the worker DLQ. The dead-letter queue is where every schema surprise, every downstream outage, every bad assumption piles up. Budget for someone owning DLQ triage; an unwatched DLQ is silent data loss with extra steps.
When This Is Overkill
If you're handling fewer than a few events per minute and the side effect is genuinely idempotent already — "set order status to shipped", "upsert a user record" — then a synchronous handler with a unique constraint on the domain table is the right design. No ledger, no queue, no workers:
INSERT INTO orders (id, status, updated_at) VALUES ($1, 'shipped', $2)
ON CONFLICT (id) DO UPDATE SET status = EXCLUDED.status, updated_at = EXCLUDED.updated_at
WHERE orders.updated_at < EXCLUDED.updated_at;
That one statement is idempotent and out-of-order safe. Ship it and move on.
Three signals tell you you've outgrown it:
- Your handler makes an external call that isn't idempotent. The moment a duplicate costs real money, you need an explicit claim before the effect.
- Your p99 handler latency approaches the provider's timeout (often 10s, sometimes 5s). You're now generating your own duplicates — split ingest from work.
- You need replay. Someone asks "can you re-run last Tuesday's events after we fixed the bug?" If you didn't store raw payloads, the answer is no. That question tends to arrive exactly once, at the worst possible time.
Key takeaway: Acknowledge fast on a durable dedupe insert, not on business logic — the unique constraint, not your `if already processed` check, is what makes redelivery safe.
Real-world challenge
Your payments webhook endpoint returns 200 in ~40ms on average, but your provider's dashboard shows a 6% retry rate and customers report duplicate confirmation emails. Tracing shows the slow tail: 4% of requests take 12–30s because the handler calls an external tax API inline. The dedupe table has a unique index on event_id and you insert into it after the email send. What's happening and what do you change?
Diagnosis. Two defects compound:
- Slow tail causes redelivery. The provider times out around 10s, retries, and the retry arrives while the first attempt is still inside the tax API call. Nothing has been recorded yet.
- The dedupe row is written last. The window between "email sent" and "event_id inserted" is the entire duration of the handler — that's exactly where the duplicate lands.
Fix, in order of leverage.
First, invert the order: claim the event before doing any work.
INSERT INTO webhook_events (event_id, status, received_at)
VALUES ($1, 'received', now())
ON CONFLICT (event_id) DO NOTHING
RETURNING event_id;
No row returned ⇒ someone else owns it ⇒ return 200 immediately.
Second, get the tax API out of the request path. The handler should do exactly: verify signature → claim → enqueue → 200. The tax call moves to a worker that reads the claimed row, with the job keyed by event_id so requeues are also idempotent.
Third, make the email itself idempotent as defence in depth — pass event_id as the provider's idempotency key so a worker crash between send and status update can't double-send.
After this, your p99 endpoint latency drops under 100ms and retry rate should fall toward the provider's baseline; any remaining retries are harmless because the claim absorbs them.