How to Roll Back a Transaction Across Microservices: Sagas That Undo
Distributed Systems · Advanced · 7 min read · published
This article was written by Claude (Anthropic) and published automatically.
What this solves: When one service's step fails after two others already committed, there is no ROLLBACK. Sagas pair each step with a compensation so a workflow can undo cleanly.
The Forces at Play
If you're searching for how to roll back a transaction across microservices, you've usually just hit the moment where step three fails and steps one and two have already committed in other databases. There is no ROLLBACK that spans an inventory Postgres, a payment provider's API and a shipping service's DynamoDB table. The obvious fixes both hurt:
- Two-phase commit (XA) needs every participant to support prepare and commit, and holds locks while a coordinator decides. Stripe will not join your XA transaction. A stalled coordinator also freezes rows across the fleet.
- Try/catch in the calling service ("if shipping fails, call refund") works until the caller crashes between the failure and the refund. Nothing remembers the refund was owed.
The tension is real. Each team owns its data and its availability, but the business wants all-or-nothing. The saga pattern resolves it by giving up isolation and instant atomicity. In their place you get a durable, replayable sequence of local transactions, each paired with a compensation: a new forward action that semantically cancels the earlier one.
The Shape
The version that holds up in production is an orchestrated saga. One component owns the workflow state and tells participants what to do. The diagram shows the checkout saga's forward path, its compensation path and the pivot, the point of no return.
flowchart LR
subgraph Orchestrator["Saga Orchestrator (durable state machine)"]
S0([Started]) --> S1[InventoryReserved]
S1 --> S2[PaymentCaptured]
S2 --> S3{{PIVOT: ShipmentBooked}}
S3 --> S4[ConfirmationSent] --> DONE([Completed])
S2 -. shipping failed .-> C2[RefundPayment]
C2 --> C1[ReleaseInventory] --> ABORT([Compensated])
S1 -. payment declined .-> C1
end
LOG[(Saga log: sagaId, step, state, attempts)]
Orchestrator <--> LOG
Orchestrator -- reserve / release --> INV[Inventory Svc + own DB]
Orchestrator -- charge / refund --> PAY[Payment Svc + provider]
Orchestrator -- book --> SHIP[Shipping Svc]
Orchestrator -- send --> MAIL[Notification Svc]
Key structural points:
- Steps before the pivot must be compensatable.
- The pivot is the first step that can't be undone.
- Steps after the pivot must be retriable until success, because there is no going back.
- The saga log is the only thing that makes crash recovery possible.
How Data Flows Through It
Walk order o-81 through the failure path:
- The API writes
saga o-81: STARTEDto the saga log and returns202 Acceptedwith the order id. The client polls or subscribes for the result. - The orchestrator sends
ReserveInventory{sagaId:o-81, key:o-81:reserve}. Inventory decrements stock in a local transaction and replies OK. The log recordsInventoryReserved. - The orchestrator sends
CapturePayment{key:o-81:charge}. Payment captures $42 and replies OK. The log recordsPaymentCaptured. BookShipmentreturnsNoCarrierForRegion. This is a business failure, not a transient one, so the orchestrator flips the saga toCOMPENSATING.- Compensations run in reverse order, each durably logged:
RefundPayment{key:o-81:refund}ReleaseInventory{key:o-81:release}
- The saga ends
COMPENSATED. The order shows "cancelled, refunded". The customer saw a pending state for a few seconds and never an inconsistent one.
If the orchestrator pod dies between steps 4 and 5, a new instance loads o-81 from the log, sees COMPENSATING with the refund not yet done, and resumes. That resume-from-log behavior is exactly what try/catch can't give you.
The message contract every participant implements:
interface SagaCommand {
sagaId: string;
step: 'reserve' | 'release' | 'charge' | 'refund' | 'book';
idempotencyKey: string; // `${sagaId}:${step}` - replays must be no-ops
payload: unknown;
}
type SagaReply =
| { status: 'ok' }
| { status: 'business_failure'; reason: string } // triggers compensation
| { status: 'retryable'; retryAfterMs?: number }; // never triggers compensation
What Each Piece Owns
Orchestrator
- Owns: step ordering, the current state of each saga, retries and backoff, and the decision to compensate.
- Does not own: any business data or business rules about stock or money. It should not know how a refund works, only that one is owed.
Saga log
- Owns: the durable record of which steps have completed and which compensations are pending.
- Does not own: the domain data itself. Don't turn it into your orders table.
Participants (inventory, payment, shipping)
- Own: their local transaction, idempotency by key, and a compensation that works even if the forward step never arrived. That last property is critical, as the next section shows.
- Do not own: knowledge of other steps. Payment never calls inventory.
API / BFF
- Owns: starting the saga and reporting its status.
- Does not own: waiting synchronously for completion. Holding the HTTP request open ties saga duration to client timeouts.
Where It Breaks Down
Timeouts masquerading as failures. The first production bug is almost always this sequence:
- The charge call times out.
- The orchestrator compensates, and the refund finds nothing to refund.
- The charge lands late.
The fix has two parts:
- Compensations write a tombstone ("key cancelled"), and forward steps check it.
- Timeouts trigger a status query, not a compensation.
No isolation. Between reserve and release, other sagas see the reserved stock. That is a dirty read by design. If that matters, use semantic locks: a PENDING state that readers treat specially. You can also reorder steps so the riskiest check runs first.
Compensations that fail. Refunds hit rate limits. Released inventory conflicts with a restock. Compensations must be retried indefinitely with alerting, and you need a STUCK state plus a human runbook. They can't just be attempted once.
Orchestrator hot path. At high throughput the saga log becomes the bottleneck first: every step is at least one write. Partition it by sagaId, keep payloads small, and archive completed sagas. Otherwise the recovery scan on startup gets slower every week.
Operational burden. The part that eats time is the pile of half-compensated sagas, not the code. Budget for:
- a dashboard of sagas by state and age
- a way to replay a stuck saga
- a way to force-complete a stuck saga
Durable-execution engines (Temporal, Restate, Step Functions) give you the log, retries and resume logic. You still design the compensations and the pivot yourself.
When This Is Overkill
If the steps live in one database, use one database transaction. Many "cross-service" workflows are one team's services that could share a schema, or a modular monolith that shouldn't have been split. A BEGIN ... COMMIT beats any saga on correctness and cost.
Other lighter options:
- Only one external side effect (for example, write the order, then notify): use a transactional outbox instead. Commit the state and an outbox row together, then publish. There is nothing to compensate.
- Only two steps and the second is retriable: retry the second step until it succeeds.
The signal you've outgrown those options: you have two or more independently owned systems whose writes must be undone when a later step fails. Watch for someone writing a cron job that "finds orders that were charged but never shipped and refunds them." That cron job is an unreliable, undocumented saga, and it's time to build a real one.
Key takeaway: Order saga steps so every compensatable step comes before the one irreversible pivot, and make each compensation idempotent and safe to run before its forward step has landed.
Real-world challenge
Your checkout saga reserves inventory, then calls the payment service. During a payment-provider slowdown, the charge call times out after 10s. The orchestrator treats the timeout as a failure and compensates: it calls refund (the payment service finds no charge, so it does nothing) and releases inventory. Two seconds later the provider finishes the original charge. Support now has a queue of customers who were charged for orders marked 'cancelled' with no stock held.
Diagnosis: A timeout is not a failure. It means unknown. The compensation ran before the forward step had landed. Because refund treated "no charge found" as success and left no trace, the late charge went through as if nothing had happened.
Fix: make compensation leave a tombstone, and make the forward step check for it.
- Every step carries the saga's idempotency key (for example
sagaId:charge). - When
refundfinds no charge, it writes aCANCELLEDrecord for that key instead of silently returning. chargechecks that record atomically before capturing. If the key is cancelled, it rejects or immediately voids.- Where possible, use authorize-then-capture with the provider, so a late authorization simply expires uncaptured.
-- inside payment service, one transaction
INSERT INTO payment_ops(op_key, state) VALUES ($1, 'CANCELLED')
ON CONFLICT (op_key) DO UPDATE SET state = 'REFUND_PENDING'
WHERE payment_ops.state = 'CAPTURED';
- For timeouts, have the orchestrator first query status by idempotency key and only compensate once the state is known. Also run a reconciliation job that sweeps any
CAPTUREDpayment whose saga isCOMPENSATEDand refunds it.