Designing a Webhook Delivery System: Why Your Retries Need an Idempotency Key josedacruz, September 13, 2026September 13, 2026 TL;DR: A webhook is just an HTTP POST your system sends to someone else’s server when something happens, but “just an HTTP POST” hides a lot of hard problems. Networks drop packets, receivers time out, and retries create duplicates — so a webhook delivery system that actually works needs retries with backoff, a way for the receiver to safely ignore duplicates (an idempotency key), and a plan for what happens when delivery keeps failing. This post walks through five common myths about webhook delivery and what’s actually true underneath each one. Myth 1: “If I send the webhook and get no error, it arrived.” This is the myth that trips up almost every team building their first webhook system. Say you have an e-commerce platform, and every time an order is marked “paid,” you send a webhook to the merchant’s server so their system can start packing the order. Your code calls an HTTP client, sends a POST request, and moves on. No exception was thrown, so surely it worked, right? Not necessarily. An HTTP request can fail in ways that never raise a clean error on your end. The request might time out after the receiver’s server already processed it but before the response made it back to you. A load balancer in between might drop the connection. Your own server might crash mid-request. From your point of view, all of these can look identical: “no confirmed success.” The only way to know for sure that a webhook was received is an explicit acknowledgment — typically an HTTP 200 status code that the receiver returns only after it has actually processed (or safely queued) the event. This is why real webhook systems are built around a retry queue instead of a single fire-and-forget call. If you don’t get a clear 2xx response within a reasonable timeout, you treat the delivery as failed and try again later, even though the first attempt might have technically succeeded. A typical delivery: attempt one times out, attempt two gets a server error, attempt three succeeds — and the idempotency key protects against any of the earlier attempts having secretly gone through. Myth 2: “Retrying immediately after a failure is the fastest way to recover.” It feels intuitive: something failed, so try again right away. But immediate retries can actually make things worse, especially at scale. Imagine your platform sends webhooks to a thousand merchants at once because you just ran a big batch of order updates. If a merchant’s server is briefly overloaded and starts failing requests, and your system responds by immediately retrying every failed request, you’ve just added more load to a server that was already struggling. That’s called a retry storm, and it can turn a small hiccup into a real outage for the receiver. The fix is exponential backoff with jitter. Exponential backoff means each retry waits longer than the last — say, 30 seconds, then 2 minutes, then 10 minutes, then an hour. Jitter means you add a small random amount of extra delay to each wait, so that if a hundred webhooks failed at the exact same moment, they don’t all retry at the exact same moment too. Spreading retries out in time gives the receiver’s server room to recover instead of getting hit by a second wave right as it’s coming back up. Most production systems also cap the number of retries — somewhere between 5 and 20 attempts over a period of hours or a day or two — rather than retrying forever. Myth 3: “The receiver will only ever get each event once.” This is probably the most important myth to unlearn. Because you can’t always tell the difference between “the request failed” and “the request succeeded but the response got lost,” any retry-based system has to accept that the same event might get delivered more than once. This is usually called “at-least-once delivery”: you’re guaranteed the receiver gets the event at least once, but you can’t guarantee exactly once. Going back to the order-paid example: if your system sends the webhook, the merchant’s server processes it and starts packing the order, but the acknowledgment gets lost on the way back to you, your system will see that as a failure and retry. The merchant’s server now gets the same “order paid” event a second time. If their code isn’t ready for that, they might pack and ship the order twice. The fix lives on the receiving side, and it’s simpler than it sounds: attach a unique ID to every event — an idempotency key — and have the receiver keep a short-term record of which IDs it has already processed. When a webhook arrives, the receiver checks the ID first. If it’s already been handled, it just returns a 200 and does nothing else. This turns “duplicate delivery” from a bug into a non-event. Myth 4: “Webhooks always arrive in the order they were sent.” It’s tempting to assume order is preserved, especially because the events usually happen in order on your side. But once retries enter the picture, order can’t be guaranteed. Picture two events for the same order: “payment received” and then, a few seconds later, “order cancelled.” If the first webhook happens to fail and go into a retry queue while the second one succeeds on the first try, the receiver could see “order cancelled” before it ever sees “payment received.” If a receiver assumes strict ordering, this kind of out-of-order delivery can cause real damage — like an order that gets marked as active again because a late “payment received” event arrived after the cancellation. There are two common ways to guard against this. One is to include a timestamp or a sequence number in every event’s payload, so the receiver can check “is this older than the last thing I processed for this order?” before acting on it. The other is to key retries per resource (for example, per order ID) so that events about the same object are always delivered in sequence, even if events about different objects can interleave. Myth 5: “A simple queue and a POST request is basically the whole system.” Once you accept that failures, retries, duplicates, and out-of-order delivery are all normal parts of the job, it becomes clear that a webhook system is closer to a small piece of distributed infrastructure than a single API call bolted onto your codebase. It needs a queue to hold pending deliveries, a worker process that attempts delivery and applies backoff on failure, a cap on retry attempts, and — critically — somewhere for events to go once they’ve exhausted all their retries. That last piece is called a dead letter queue. Instead of silently giving up after the final retry, a well-built system moves the failed event into a separate holding area where a human (or an automated alert) can notice it and decide what to do — resend it manually once the receiver’s server is back up, or investigate why it kept failing. Without a dead letter queue, permanently failing webhooks just vanish, and the merchant in our example might never find out their “order shipped” webhook silently failed twenty times and got dropped. The full lifecycle of a single webhook delivery, including the path to a dead letter queue when retries run out. Putting it together None of these five pieces — acknowledgment-based success, backoff with jitter, idempotency keys, ordering guards, and a dead letter queue — are exotic. Each one is a fairly small, well-understood piece of engineering on its own. What makes webhook delivery feel harder than it should be is that skipping any single one of them tends to produce a bug that only shows up under load, or only for one particular customer, or only once every few thousand events — which makes it painful to track down after the fact. If you’re building or reviewing a webhook system, the five myths above double as a checklist. Does a delivery only count as successful when you get an explicit acknowledgment? Do failed deliveries back off instead of hammering the receiver? Can the receiver safely ignore a duplicate event? Is there a way to detect events arriving out of order? And when an event exhausts every retry, does it land somewhere visible instead of disappearing? If you can answer yes to all five, you’ve built a webhook system that will hold up once real traffic — and real network failures — show up. Related architecture distributed systemsidempotencyreliabilitysystem designwebhooks