Skip to content
José da Cruz

Enterprise Architect, and Life Thinker

José da Cruz

Enterprise Architect, and Life Thinker

How Distributed Locks Actually Work (And Why They’re Harder Than They Look)

josedacruz, August 5, 2026August 5, 2026

TL;DR: A distributed lock stops two servers from doing the same risky thing at the same time, like selling the last concert ticket twice. The tricky part isn’t getting the lock — it’s what happens if the worker holding it crashes, freezes, or runs slower than expected. Get the timeout and release logic wrong and your lock either deadlocks everything or stops protecting anything at all.

The problem locks were built to solve

Imagine you’re on the backend team for a ticketing site called TicketNest. A popular band just announced a show, and there’s exactly one ticket left. Two fans click “Buy” within the same second.

If your app runs on a single server, this is easy. You can use a regular lock (sometimes called a mutex, short for “mutual exclusion”) built into your programming language. It makes sure only one piece of code touches the ticket count at a time. The first buyer’s code runs, finishes, and only then does the second buyer’s code get a turn. No double-selling.

But TicketNest doesn’t run on one server. It runs on ten, behind a load balancer, so it can handle traffic spikes. The two fans’ requests might land on two completely different servers at the exact same moment. A regular in-process lock only protects code running inside one process on one machine. Server A has no idea what Server B is doing. Both servers check “is this ticket available?”, both see yes, and both sell it. Now you’ve oversold a one-of-a-kind ticket and have an angry customer to deal with.

This is exactly the situation a distributed lock is for: a lock that every server in your fleet can see and respect, usually backed by a shared, fast piece of infrastructure like Redis, a database, or a dedicated coordination service.

The basic idea: a shared “key is taken” flag

The simplest version works like this. Before touching the ticket, a server tries to write a special key into a shared store, something like lock:ticket:8842. The store only lets one server succeed at creating that key. Whoever wins the race gets to do the work — check the ticket, sell it, update the count. Everyone else gets told “not right now” and can retry a moment later or show the customer a friendly error.

When the winning server finishes, it deletes the key, freeing the lock for the next request.
Flow diagram showing a worker requesting a lock, the lock service checking availability, setting a key with a TTL, the worker doing protected work, then releasing the lock
The basic lock lifecycle: ask, acquire with a timeout, do the work, release.

This sounds simple, and for a lot of real systems it’s genuinely enough. The hard part shows up when you ask: what happens if the server holding the lock never gets to the “delete the key” step?

Why locks that never expire are dangerous

Say the server that won the race crashes right after creating the lock key, before it finishes selling the ticket. If nothing ever removes that key, every other server is now permanently locked out of ticket 8842. The ticket becomes unsellable forever, even though it’s actually still available. One crash just took your last ticket off the market.

This is why real distributed locks almost always come with a TTL, short for “time to live”: an expiration time set when the lock is created. If the key isn’t deleted within, say, 10 seconds, the lock store removes it automatically. Now a crash only costs you 10 seconds of unavailability instead of forever.

But a TTL introduces a new problem, and it’s the one that trips up most teams the first time they build this.

The problem with timeouts: slow doesn’t mean dead

Picture this sequence at TicketNest:

  • Server A acquires the lock on ticket 8842, with a 10-second TTL.
  • Server A starts doing the work: checking inventory, charging the card, writing the sale to the database.
  • Server A hits an unlucky garbage collection pause (a normal, if annoying, part of how many languages manage memory) that freezes it for 12 seconds.
  • While Server A is frozen, its lock expires. Server B now acquires the same lock and starts selling the same ticket.
  • Server A wakes up, has no idea time has passed, finishes its own sale, and deletes the lock — which by now is actually Server B’s lock.

You’ve now sold the ticket twice, and Server A even helpfully deleted Server B’s still-valid lock on its way out. The TTL that was supposed to prevent a permanent deadlock just caused a double-sale instead. This is the core trade-off of distributed locks: a short TTL risks two workers overlapping, a long TTL risks a long outage when a worker dies. There’s no TTL value that makes this problem disappear, because the real issue is that “the lock expired” and “the worker is dead” are not the same fact, even though it’s tempting to treat them as one.

Fencing tokens: catching the overlap

The standard fix is called a fencing token. Instead of just handing out a lock, the lock service hands out a lock plus a number that increases every time the lock changes hands — 1, 2, 3, and so on. Server A gets token 7. When its lock expires and Server B acquires the lock, Server B gets token 8.

The trick is that the ticket database itself is taught to reject any write that shows up with an old token. So when Server A wakes from its pause and tries to write its sale using token 7, the database says “I’ve already seen token 8, I’m ignoring anything numbered 7 or lower.” Server A’s stale write is rejected. It’s not the lock that protects you here — it’s the fact that the last line of defense, the actual write, knows how to spot a straggler and refuse it.

This matters because a lock by itself is really just a hint that says “probably safe to proceed.” It can’t physically stop a slow server from continuing to run. Fencing tokens push the real safety check down to the one place that can enforce it: the system actually recording the result.

Choosing a locking strategy without overbuilding

Not every situation needs the full Redis-plus-fencing-token setup, and it’s worth resisting the urge to reach for the fanciest tool by default. A few questions help narrow it down:

  • Do you actually need a lock at all? Sometimes you can sidestep the whole problem. If your database supports an atomic “decrement the ticket count, but only if it’s above zero” operation, you may not need a separate lock — the database is already doing the coordination for you.
  • Is the protected work extremely short? If it’s a single, fast database update, a plain row lock (SELECT ... FOR UPDATE in SQL) is often simpler and safer than a separate lock service, because the database already handles crashes and timeouts for you.
  • Do you already run Redis? If so, a TTL-based Redis lock with fencing tokens is a well-worn, well-documented path and probably your best default for anything longer-running.
  • Do correctness requirements go beyond what Redis comfortably guarantees? Systems like etcd or ZooKeeper are built specifically for coordination and handle edge cases (like a lock service itself losing data) more rigorously. They’re heavier to run and operate, so they’re usually reserved for infrastructure-level problems — leader election in a cluster, for example — rather than everyday application logic.

Decision tree for choosing a locking strategy based on whether concurrent access is safe, how long the work takes, and what infrastructure is already available
A rough guide for picking the lightest tool that’s still safe enough.

The takeaway

A distributed lock’s job looks simple from the outside: let one worker in, keep everyone else out. What makes it genuinely hard is that “the lock expired” doesn’t mean “the worker stopped.” Any design that assumes those are the same thing will eventually double-sell a ticket, double-charge a card, or run a batch job twice. The fix isn’t a cleverer lock — it’s accepting that the lock can only ever be a hint, and putting the real check at the place where the damage would actually happen.

Before reaching for a distributed lock, it’s worth asking whether you can design the operation to be safe even when it runs twice, an idea called idempotency. If “sell this ticket” can be rewritten as “sell this ticket, but only if it hasn’t already been sold in this exact way,” you may not need a lock’s complexity at all. Locks are a powerful tool, but the simplest fix for a race condition is often to make the race harmless instead of trying to prevent it from happening.

Related

architecture concurrencydistributed systemsrace conditionsredissystem design

Post navigation

Previous post
Next post

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

  • architecture (154)
  • artificial-intelligence (6)
  • books (2)
  • courses (4)
  • decision-frameworks (4)
  • finances (1)
  • java (8)
  • life-improvement (20)
  • metrics (3)
  • observability (2)
  • puzzles and challenges (3)
  • reviews (1)
  • security (8)
  • spring-boot (36)
  • Systems Thinking (3)
  • GitHub Weekly: MCP Tools Trending Right Now (September 28 – October 4, 2026)
  • The Bulkhead Pattern: Why One Slow Dependency Shouldn’t Sink the Whole Ship
  • Essential Reading: A Philosophy of Software Design — Fighting Complexity One Module at a Time
  • GitHub Weekly: Top 10 Trending Repos (September 21 – September 27, 2026)
  • The Volunteer Tax: Why Raising Your Hand in a Meeting Makes You the Permanent Owner
©2026 José da Cruz | WordPress Theme by SuperbThemes