How to Roll Back a Transaction Across Microservices: Sagas That Undo

Distributed Systems · Advanced · 7 min read · published

This article was written by Claude (Anthropic) and published automatically.

What this solves: When one service's step fails after two others already committed, there is no ROLLBACK. Sagas pair each step with a compensation so a workflow can undo cleanly.

The Forces at Play

If you're searching for how to roll back a transaction across microservices, you've usually just hit the moment where step three fails and steps one and two have already committed in other databases. There is no ROLLBACK that spans an inventory Postgres, a payment provider's API and a shipping service's DynamoDB table. The obvious fixes both hurt:

The tension is real. Each team owns its data and its availability, but the business wants all-or-nothing. The saga pattern resolves it by giving up isolation and instant atomicity. In their place you get a durable, replayable sequence of local transactions, each paired with a compensation: a new forward action that semantically cancels the earlier one.

The Shape

The version that holds up in production is an orchestrated saga. One component owns the workflow state and tells participants what to do. The diagram shows the checkout saga's forward path, its compensation path and the pivot, the point of no return.

flowchart LR
  subgraph Orchestrator["Saga Orchestrator (durable state machine)"]
    S0([Started]) --> S1[InventoryReserved]
    S1 --> S2[PaymentCaptured]
    S2 --> S3{{PIVOT: ShipmentBooked}}
    S3 --> S4[ConfirmationSent] --> DONE([Completed])
    S2 -. shipping failed .-> C2[RefundPayment]
    C2 --> C1[ReleaseInventory] --> ABORT([Compensated])
    S1 -. payment declined .-> C1
  end
  LOG[(Saga log: sagaId, step, state, attempts)]
  Orchestrator <--> LOG
  Orchestrator -- reserve / release --> INV[Inventory Svc + own DB]
  Orchestrator -- charge / refund --> PAY[Payment Svc + provider]
  Orchestrator -- book --> SHIP[Shipping Svc]
  Orchestrator -- send --> MAIL[Notification Svc]

Key structural points:

How Data Flows Through It

Walk order o-81 through the failure path:

  1. The API writes saga o-81: STARTED to the saga log and returns 202 Accepted with the order id. The client polls or subscribes for the result.
  2. The orchestrator sends ReserveInventory{sagaId:o-81, key:o-81:reserve}. Inventory decrements stock in a local transaction and replies OK. The log records InventoryReserved.
  3. The orchestrator sends CapturePayment{key:o-81:charge}. Payment captures $42 and replies OK. The log records PaymentCaptured.
  4. BookShipment returns NoCarrierForRegion. This is a business failure, not a transient one, so the orchestrator flips the saga to COMPENSATING.
  5. Compensations run in reverse order, each durably logged:
    • RefundPayment{key:o-81:refund}
    • ReleaseInventory{key:o-81:release}
  6. The saga ends COMPENSATED. The order shows "cancelled, refunded". The customer saw a pending state for a few seconds and never an inconsistent one.

If the orchestrator pod dies between steps 4 and 5, a new instance loads o-81 from the log, sees COMPENSATING with the refund not yet done, and resumes. That resume-from-log behavior is exactly what try/catch can't give you.

The message contract every participant implements:

interface SagaCommand {
  sagaId: string;
  step: 'reserve' | 'release' | 'charge' | 'refund' | 'book';
  idempotencyKey: string; // `${sagaId}:${step}` - replays must be no-ops
  payload: unknown;
}

type SagaReply =
  | { status: 'ok' }
  | { status: 'business_failure'; reason: string } // triggers compensation
  | { status: 'retryable'; retryAfterMs?: number }; // never triggers compensation

What Each Piece Owns

Orchestrator

Saga log

Participants (inventory, payment, shipping)

API / BFF

Where It Breaks Down

Timeouts masquerading as failures. The first production bug is almost always this sequence:

  1. The charge call times out.
  2. The orchestrator compensates, and the refund finds nothing to refund.
  3. The charge lands late.

The fix has two parts:

No isolation. Between reserve and release, other sagas see the reserved stock. That is a dirty read by design. If that matters, use semantic locks: a PENDING state that readers treat specially. You can also reorder steps so the riskiest check runs first.

Compensations that fail. Refunds hit rate limits. Released inventory conflicts with a restock. Compensations must be retried indefinitely with alerting, and you need a STUCK state plus a human runbook. They can't just be attempted once.

Orchestrator hot path. At high throughput the saga log becomes the bottleneck first: every step is at least one write. Partition it by sagaId, keep payloads small, and archive completed sagas. Otherwise the recovery scan on startup gets slower every week.

Operational burden. The part that eats time is the pile of half-compensated sagas, not the code. Budget for:

Durable-execution engines (Temporal, Restate, Step Functions) give you the log, retries and resume logic. You still design the compensations and the pivot yourself.

When This Is Overkill

If the steps live in one database, use one database transaction. Many "cross-service" workflows are one team's services that could share a schema, or a modular monolith that shouldn't have been split. A BEGIN ... COMMIT beats any saga on correctness and cost.

Other lighter options:

The signal you've outgrown those options: you have two or more independently owned systems whose writes must be undone when a later step fails. Watch for someone writing a cron job that "finds orders that were charged but never shipped and refunds them." That cron job is an unreliable, undocumented saga, and it's time to build a real one.

Key takeaway: Order saga steps so every compensatable step comes before the one irreversible pivot, and make each compensation idempotent and safe to run before its forward step has landed.

Real-world challenge

Your checkout saga reserves inventory, then calls the payment service. During a payment-provider slowdown, the charge call times out after 10s. The orchestrator treats the timeout as a failure and compensates: it calls refund (the payment service finds no charge, so it does nothing) and releases inventory. Two seconds later the provider finishes the original charge. Support now has a queue of customers who were charged for orders marked 'cancelled' with no stock held.

Diagnosis: A timeout is not a failure. It means unknown. The compensation ran before the forward step had landed. Because refund treated "no charge found" as success and left no trace, the late charge went through as if nothing had happened.

Fix: make compensation leave a tombstone, and make the forward step check for it.

  1. Every step carries the saga's idempotency key (for example sagaId:charge).
  2. When refund finds no charge, it writes a CANCELLED record for that key instead of silently returning.
  3. charge checks that record atomically before capturing. If the key is cancelled, it rejects or immediately voids.
  4. Where possible, use authorize-then-capture with the provider, so a late authorization simply expires uncaptured.
-- inside payment service, one transaction
INSERT INTO payment_ops(op_key, state) VALUES ($1, 'CANCELLED')
ON CONFLICT (op_key) DO UPDATE SET state = 'REFUND_PENDING'
  WHERE payment_ops.state = 'CAPTURED';
  1. For timeouts, have the orchestrator first query status by idempotency key and only compensate once the state is known. Also run a reconciliation job that sweeps any CAPTURED payment whose saga is COMPENSATED and refunds it.