Skip to content
José da Cruz

Enterprise Architect, and Life Thinker

José da Cruz

Enterprise Architect, and Life Thinker

Reinforcing vs. Balancing Feedback Loops: Why Retries Can Make an Outage Worse

josedacruz, August 3, 2026August 3, 2026

You get paged at 2am. A payment service is running slow. You watch the dashboard, and something strange happens: the more you look at it, the worse it gets. Error rates climb. Latency climbs. Then, ten minutes after the first alert, the whole checkout system goes down — not because the original problem was that bad, but because of what your own system did in response to it.

This is not bad luck. It is a predictable pattern called a feedback loop, and once you can see it, you start noticing it everywhere: in retry logic, in autoscaling, in alerting, even in team meetings. This post explains what a feedback loop is, why one specific kind of it is responsible for a huge share of “small problem became a huge outage” stories, and what you can do about it.

What a feedback loop actually is

A feedback loop is just a fancy name for a simple idea: the output of a system becomes an input back into that same system. Something happens, that something changes the conditions, and the changed conditions influence what happens next.

There are two flavors, and the difference between them matters a lot.

  • A balancing loop pushes back toward a stable point. Think of a thermostat: when the room gets too hot, the thermostat turns the heating off; when it gets too cold, it turns the heating back on. The system corrects itself.
  • A reinforcing loop pushes further in the same direction. The more something happens, the more it keeps happening, faster and faster, until something breaks. A microphone placed too close to its own speaker is the classic example: a tiny bit of sound gets picked up, amplified, played back louder, picked up again, amplified again — and within a second you have an ear-splitting screech.

Software systems are full of both kinds. The dangerous ones are the reinforcing loops you didn’t know you built.

A concrete example: the retry storm

Imagine you work on the backend team for an online store called Northwind Goods. Northwind is running a flash sale, and traffic is three times higher than a normal day. There is a payment-service that talks to a third-party payment processor to charge customers’ cards.

Here is the chain of events, step by step:

  1. The payment processor, under the extra load, starts responding a bit slower than usual — say, 2 seconds instead of 200 milliseconds.
  2. Your checkout-service calls payment-service with a 1-second timeout. Because the processor is now slower than that, a chunk of those calls start timing out.
  3. Like most reasonable-sounding code, checkout-service is written to retry on timeout — nobody wants to tell a customer “checkout failed” just because of one slow response. So every timed-out request gets retried automatically, often more than once.
  4. Those retries are new requests. They land on top of the requests that are already queued up and struggling to get through. The payment processor now has even more work to do than before.
  5. Because there is more work to do, the payment processor gets even slower. Which causes more timeouts. Which causes more retries. Which adds more load.

Notice what happened: nothing about this loop returns the system to normal. Each trip around the loop makes the next trip worse. That is the signature of a reinforcing loop. Fifteen minutes in, what started as “the payment processor is a bit slow today” has turned into “checkout is completely down,” even though the underlying processor might only be running at, say, 40% worse latency than usual. The retries didn’t fix the slowness — they turned a moderate slowdown into a total outage.

Here is what that loop looks like when you draw it out:

Diagram of a reinforcing feedback loop: payment service slows down, requests time out, clients retry, load increases, and the service slows down further
A reinforcing loop: each step makes the next step worse, with nothing pulling the system back toward normal.

Why this is so easy to miss

The reason retry storms sneak up on so many teams is that retry logic looks like a purely local decision. When you’re writing the code for checkout-service, retrying a failed call to payment-service seems like simple defensive programming — of course you should try again rather than immediately giving up. Nothing about that one line of code looks dangerous.

The danger only shows up when you zoom out and ask: what happens if every caller does this at the same time, against a system that is already struggling? A single retry is harmless. Thousands of retries, all firing within the same few seconds because they all had similar timeouts, is a wave of extra load hitting a system exactly when it can least afford it. This is a core lesson of systems thinking in general: the behavior of a system is often not obvious from looking at any one piece of it. You have to trace the loop all the way around before the danger becomes visible.

Turning a reinforcing loop into a balancing one

The fix is not “don’t retry.” Retries are still useful — a single request can genuinely fail for a one-off reason and succeed on the second try. The fix is to add mechanisms that behave like the thermostat: something that senses things are getting worse and actively pushes back, instead of blindly adding more load.

A few standard tools do exactly this:

  • Exponential backoff with jitter. Instead of retrying immediately, wait a bit longer after each failure (1 second, then 2, then 4…), and add a small random delay (“jitter”) so that thousands of clients don’t all retry at the exact same moment and create a new spike together.
  • Circuit breakers. A circuit breaker is a piece of code that watches the error rate for a downstream call. If too many calls to payment-service start failing, the circuit “opens” and, for a short period, checkout-service stops calling it at all — failing fast instead of piling on more requests. This gives the struggling service room to recover, which is exactly the pushback a thermostat provides.
  • Load shedding. When a service is overwhelmed, it deliberately rejects some incoming requests early and cheaply (often with a clear “try again later” response) rather than accepting every request and slowly drowning under all of them.
  • Retry budgets. Cap the total number of retries a service is allowed to send in a given window, so retries can never multiply traffic without limit.

Every one of these tools does the same conceptual job: it turns a loop that amplifies trouble into a loop that dampens it.

The habit worth building

You don’t need a systems-thinking degree to use this. The next time you add a retry, a queue, an autoscaling rule, or an alert threshold, ask yourself one question: if this action fires, does it make the underlying condition more likely to happen again, or less?

Autoscaling that adds servers when CPU is high is a balancing loop — more capacity relieves the pressure that triggered it. But autoscaling that’s misconfigured to scale based on request queue length, when the new servers themselves add coordination overhead that grows the queue further, can quietly become a reinforcing loop. Alerts that page an already-overloaded on-call engineer with twenty duplicate notifications instead of one clear one are doing the same thing to a human system instead of a technical one.

Once you start asking that question by habit, you start catching these loops on paper, before they cost you a 2am page. That’s the real value of thinking in feedback loops: it’s a cheap mental model that catches expensive problems early.

Related

architecture feedback loopsmental modelsresilience engineeringretry logicsystems thinking

Post navigation

Previous post
Next post

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

  • architecture (154)
  • artificial-intelligence (6)
  • books (2)
  • courses (4)
  • decision-frameworks (4)
  • finances (1)
  • java (8)
  • life-improvement (20)
  • metrics (3)
  • observability (2)
  • puzzles and challenges (3)
  • reviews (1)
  • security (8)
  • spring-boot (36)
  • Systems Thinking (3)
  • GitHub Weekly: MCP Tools Trending Right Now (September 28 – October 4, 2026)
  • The Bulkhead Pattern: Why One Slow Dependency Shouldn’t Sink the Whole Ship
  • Essential Reading: A Philosophy of Software Design — Fighting Complexity One Module at a Time
  • GitHub Weekly: Top 10 Trending Repos (September 21 – September 27, 2026)
  • The Volunteer Tax: Why Raising Your Hand in a Meeting Makes You the Permanent Owner
©2026 José da Cruz | WordPress Theme by SuperbThemes