Skip to content
José da Cruz

Enterprise Architect, and Life Thinker

José da Cruz

Enterprise Architect, and Life Thinker

The Saga Pattern: How to Undo a Transaction That Spans Five Services

josedacruz, September 16, 2026September 16, 2026

TL;DR: A saga is how you handle a “transaction” that spans multiple services, since you can’t wrap them all in one database commit. Instead of one atomic rollback, you run a chain of local steps and, if something downstream fails, you undo the earlier steps with separate “compensating” actions. It’s not magic and it’s not automatic — you have to design the undo path as carefully as the happy path.

If you’ve worked with a single database, you’ve probably leaned on transactions without thinking much about it. You update three tables, and either all three changes stick, or none of them do. That guarantee disappears the moment those three updates live in three different services with three different databases. There’s no single “commit” button that spans all of them. The saga pattern is the answer most teams reach for — and it comes with a pile of myths that trip people up the first time they use it. Let’s walk through the biggest ones.

Myth 1: A saga is basically a distributed transaction

This is the most common misunderstanding, and it’s an easy one to make because the word “transaction” is right there in the pattern’s use case. A real database transaction gives you atomicity (all steps happen or none do) and isolation (nobody else sees the in-between state). A saga gives you neither.

Picture an online store called ShopFlow. Placing an order means: reserve the item in Inventory, charge the card in Payment, and book a slot in Shipping. Each of those is its own local transaction inside its own service, and each one commits on its own the moment it succeeds. There is no single lock holding all three together. If Shipping fails after Payment already succeeded, you don’t get an automatic rollback — you get a completed charge sitting there, and it’s now your job to undo it. A saga is really a sequence of separate, already-committed steps, plus a plan for cleaning up if a later step fails.

Sequence diagram showing an orchestrator calling Inventory, Payment, and Shipping, with a shipping failure triggering a refund and inventory release
The orchestrator style of saga: one coordinator calls each service in order and triggers compensations when a step fails.

Myth 2: You need a special saga framework to do this correctly

Frameworks like Temporal or Camunda are genuinely useful once you have dozens of these multi-step flows and need retries, timeouts, and visibility handled for you. But the pattern itself doesn’t require one. There are two plain approaches, and either can be hand-rolled.

Orchestration means one piece of code — the orchestrator — calls each service in order and decides what to do when a step fails. This is what the diagram above shows: it’s just a function (or a small state machine) that calls Inventory, then Payment, then Shipping, and calls the “undo” version of each earlier step if a later one fails.

Choreography means there’s no central coordinator. Each service listens for events and reacts. Payment succeeds, publishes a “payment charged” event; Shipping listens for that event, tries to book a slot, and if it fails, it publishes a “shipping failed” event that Payment listens for so it can refund itself. Both approaches are just event handlers and function calls — the “framework” is optional scaffolding, not a requirement.

Myth 3: Compensating actions fully undo what happened, like a rollback

This is the myth that causes the most painful surprises in production. A database rollback erases a change as if it never happened. A compensating action in a saga does not do that — it can only take a new action that cancels out the effect of the old one, and sometimes it can’t even do that cleanly.

Back to ShopFlow: if Payment already charged the customer’s card and Shipping then fails, the compensating action isn’t “un-charge the card.” It’s “issue a refund.” Those look similar but aren’t the same thing — the customer’s bank statement will show a charge and a separate refund line, not a clean nothing. It gets worse for some steps: if a “your order shipped” email already went out before the failure was detected, there is no compensating action that unsends an email. The best you can do is send a follow-up apology email. When you design a saga, ask of every step: “if I have to undo this later, what does undoing it actually look like to the customer?” Some steps you can only partially undo, and you need to plan for that up front, not discover it during an incident.

Myth 4: Sagas are a microservices-only concept

The word “saga” gets used almost exclusively in microservices talks, so it’s easy to assume it doesn’t apply anywhere else. But the actual trigger for needing this pattern isn’t “do I have microservices” — it’s “do I have a multi-step process where each step has a side effect I can’t roll back for free.”

A single monolithic backend can hit this just as easily. Say your one application calls a third-party payment gateway, then a third-party shipping API, then a third-party email service. Those are three external systems, none of which know about each other, and none of which are inside your database’s transaction boundary. If the shipping API call fails, you still need to go back and refund the payment through the gateway’s API. That’s a saga, full stop, even though your own code is a single deployable monolith. The pattern is about crossing transaction boundaries, not about how many services you’ve split your code into.

Myth 5: If every step is idempotent, your saga is safe

Idempotency is a real and important property here — it means calling an action twice has the same effect as calling it once, which matters because network calls fail and get retried. If your “charge card” step isn’t idempotent, a retry after a timeout could charge the customer twice. So yes, make every step in your saga idempotent. But that solves exactly one problem: safe retries. It does not solve everything else that can go wrong.

Idempotency doesn’t tell you what state to show the customer while the saga is still in progress or being compensated. If someone checks their order status right after Shipping fails but before the refund has gone through, what do they see? “Confirmed”? “Processing”? An order stuck in limbo looks broken even if the underlying logic is technically correct. This is why it helps to model the whole thing as an explicit state machine, with a real “Compensating” state, rather than just a pass/fail flag.

State diagram showing an order moving from Pending through Reserved, Charged, Shipped, or into Compensating and Cancelled on failure
Modeling the saga as explicit states — including a Compensating state — makes the in-between periods visible instead of hidden inside a boolean flag.

Idempotency also doesn’t protect you from a saga that just hangs. If the Shipping service never responds at all, you need a timeout that decides “this step has failed” and kicks off compensation — idempotency has nothing to say about that. It’s a necessary ingredient, not a complete safety net.

The takeaway

The saga pattern is less about a clever trick and more about being honest that distributed systems don’t get free rollbacks. Once you accept that every step needs its own undo plan, that undoing isn’t always clean, and that the in-between states need to be visible rather than hidden, the pattern stops feeling exotic. It’s just careful bookkeeping, spread across services instead of tables.

Related

architecture compensating transactionsdistributed transactionsmicroservicessaga patternsystem design

Post navigation

Previous post
Next post

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

  • architecture (154)
  • artificial-intelligence (6)
  • books (2)
  • courses (4)
  • decision-frameworks (4)
  • finances (1)
  • java (8)
  • life-improvement (20)
  • metrics (3)
  • observability (2)
  • puzzles and challenges (3)
  • reviews (1)
  • security (8)
  • spring-boot (36)
  • Systems Thinking (3)
  • GitHub Weekly: MCP Tools Trending Right Now (September 28 – October 4, 2026)
  • The Bulkhead Pattern: Why One Slow Dependency Shouldn’t Sink the Whole Ship
  • Essential Reading: A Philosophy of Software Design — Fighting Complexity One Module at a Time
  • GitHub Weekly: Top 10 Trending Repos (September 21 – September 27, 2026)
  • The Volunteer Tax: Why Raising Your Hand in a Meeting Makes You the Permanent Owner
©2026 José da Cruz | WordPress Theme by SuperbThemes