Skip to content
José da Cruz

Enterprise Architect, and Life Thinker

José da Cruz

Enterprise Architect, and Life Thinker

The Bulkhead Pattern: Why One Slow Dependency Shouldn’t Sink the Whole Ship

josedacruz, October 2, 2026October 2, 2026

TL;DR: The bulkhead pattern means giving each dependency — a downstream API, a database, a queue — its own separate pool of resources instead of making everyone share one big pool. That way, when one dependency gets slow or fails, it can only eat its own pool, not every other request in your app. It’s a small, cheap change that stops the most common way a minor slowdown turns into a full outage.

What is the bulkhead pattern, exactly?

Imagine your app talks to three other services: a payment API, a shipping API, and a recommendations API that suggests “you might also like” products. If all three calls share the same pool of, say, 30 worker threads, then any one of them can grab as many of those 30 threads as it wants. If the recommendations API starts responding slowly, more and more threads sit around waiting for it to answer. Eventually all 30 threads are stuck waiting on recommendations, and now payment and shipping calls — which were working fine the whole time — can’t get a thread either. Your checkout flow is down because of a “you might also like” widget.

The bulkhead pattern fixes this by splitting that one shared pool into several smaller, separate pools — one per dependency. Recommendations gets its own 10 threads. Payment gets its own 10. Shipping gets its own 10. Now if recommendations gets slow, it can only use up its own 10 threads. Payment and shipping keep working normally, because they were never sharing a pool with recommendations in the first place.

Where does the name come from?

It’s borrowed straight from shipbuilding. A ship’s hull is divided into separate watertight compartments called bulkheads. If the hull gets punctured in one compartment, that compartment floods, but watertight doors keep the water from spreading into the rest of the ship. The ship stays afloat with one damaged section instead of sinking entirely.

Software architects borrowed the term for exactly the same idea: contain the damage to one part of the system instead of letting it spread everywhere. A “leak” in one dependency — slowness, errors, timeouts — gets sealed off instead of flooding every other request.

What does this actually look like when it goes wrong, in practice?

Here’s a realistic version of the story. A mid-sized e-commerce team runs a checkout service that calls three downstream APIs on every page load: payments, shipping estimates, and product recommendations. All three calls go through the same shared HTTP client, which under the hood uses one connection pool of 50 connections for the whole service.

One afternoon, the recommendations API — owned by a different team — gets slow because of an unrelated database migration. Each call that used to take 50 milliseconds now takes 8 seconds. Nothing about recommendations is “broken,” it’s just slow. But every checkout request still calls it, and every one of those calls now holds a connection from the shared pool for 8 seconds instead of 50 milliseconds. Within a couple of minutes, all 50 connections are tied up waiting on recommendations. New checkout requests — including ones that only needed payments and shipping, not recommendations at all — can’t get a connection. The whole checkout flow grinds to a halt, and revenue stops, over a feature that was supposed to be a nice-to-have.
Journey comparing an on-call incident with and without bulkheads
Same incident, two outcomes: without isolation, one slow dependency drags everything down with it; with it, the damage stays contained.

If recommendations had been calling out through its own separate connection pool, the story stops after “recommendations gets slow.” Payment and shipping keep running on their own pools, checkout keeps working, and the team fixes recommendations without anyone outside that team even noticing.

What does a bulkhead actually look like in code?

Concretely, it usually means giving each outbound dependency its own:

  • Thread pool or worker pool (if you’re making blocking calls)
  • Connection pool (database connections, HTTP connections)
  • Or, in async code, a separate semaphore limiting how many concurrent calls to that dependency are allowed at once

So instead of one shared HttpClient with one pool of 50 connections for every outbound call, you’d configure three clients, each scoped to one dependency: one for payments with its own pool, one for shipping with its own pool, one for recommendations with its own pool. Most HTTP client libraries and many resilience libraries (Resilience4j in Java, Polly in .NET, and similar tools in other languages) have a bulkhead concept built in — you usually just declare a named pool with a max size per dependency and route calls through it.
Checkout service calling three dependencies, each through its own isolated thread pool
Each dependency gets its own lane — nothing shared, nothing to starve.

You don’t need a fancy library to get the basic idea working, either. Even just instantiating a separate, smaller thread pool per external call and refusing to let them share is most of the benefit.

Isn’t this basically the same thing as a circuit breaker?

They’re related, but they solve different problems, and the strongest setups use both together. A circuit breaker watches a dependency’s error rate or latency and, once it crosses a threshold, stops sending it calls for a while — it’s a reaction to a dependency that’s already clearly unhealthy. A bulkhead doesn’t wait for anything to be detected as unhealthy at all. It just makes sure that, healthy or not, one dependency can never consume more than its fair share of resources.

Think of it this way: the bulkhead limits the damage a slow dependency can do while it’s still slow. The circuit breaker decides when to stop calling it altogether. Bulkheads protect your resources; circuit breakers protect your dependency (and save you the wasted work of calling something that’s going to fail anyway). Used together, the bulkhead keeps the blast radius small, and the circuit breaker shortens how long the problem lasts.

How do I decide how big each pool should be?

Start from how important and how risky each dependency is, not from an even split. A dependency that’s critical to your main flow (payments, in the example above) can usually justify a larger pool, because you want it to keep up with real traffic. A dependency that’s optional or a “nice to have” (recommendations) should get a smaller pool on purpose — if it’s going to fail, you want it to fail small.

A reasonable starting point is to look at your current traffic and latency numbers for each dependency, size the pool to comfortably handle normal load with some headroom, and then deliberately cap it below what the whole service could ever starve on. You don’t need to get this perfect on day one. The real win isn’t the exact number, it’s the fact that a ceiling exists at all for each dependency separately, instead of one shared, uncapped free-for-all.

Are there any downsides to adding bulkheads everywhere?

Yes, a couple worth knowing before you reach for this on every single call. First, it adds operational complexity: more pools to configure, monitor, and tune, and more places where you can get a number wrong (set a pool too small and you’ll see artificial rejections even when the dependency is perfectly healthy). Second, it can mean slightly less efficient use of resources overall — a shared pool can, in theory, use idle capacity more fully than several smaller fixed pools, since an idle slot in the payments pool can’t be borrowed by a busy shipping pool.

In practice, for most teams, that tradeoff is well worth it. A little wasted idle capacity is a much smaller cost than an outage that takes down your whole service over a non-critical dependency. But you don’t need to bulkhead every single thing your service touches on day one — see the next question.

How do I know if I actually need this?

The clearest signal is asking: “if this dependency got slow right now, would it take anything else down with it?” If your service calls more than one external thing through a shared pool of any kind — threads, connections, whatever — the answer is probably yes, and that’s your cue. You don’t have to bulkhead everything at once. Start with the dependency that’s least critical to your core flow but still shares resources with critical calls — that’s exactly the “recommendations API” situation above, and it’s usually the highest-value, lowest-risk place to start. From there, widen it out to your other external calls as you have time.

Related

architecture architecture patternsbulkhead patternfault isolationmicroservicesresilience

Post navigation

Previous post
Next post

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

  • architecture (154)
  • artificial-intelligence (6)
  • books (2)
  • courses (4)
  • decision-frameworks (4)
  • finances (1)
  • java (8)
  • life-improvement (20)
  • metrics (3)
  • observability (2)
  • puzzles and challenges (3)
  • reviews (1)
  • security (8)
  • spring-boot (36)
  • Systems Thinking (3)
  • GitHub Weekly: MCP Tools Trending Right Now (September 28 – October 4, 2026)
  • The Bulkhead Pattern: Why One Slow Dependency Shouldn’t Sink the Whole Ship
  • Essential Reading: A Philosophy of Software Design — Fighting Complexity One Module at a Time
  • GitHub Weekly: Top 10 Trending Repos (September 21 – September 27, 2026)
  • The Volunteer Tax: Why Raising Your Hand in a Meeting Makes You the Permanent Owner
©2026 José da Cruz | WordPress Theme by SuperbThemes