The Dead Letter Queue Pattern: Why Nobody Ever Looks Inside It josedacruz, September 18, 2026September 18, 2026 TL;DR: A dead letter queue (DLQ) is just a holding pen for messages your system tried to process and failed. It’s not a fix by itself — it just moves the failure somewhere quieter. Whether you actually need one depends on how much it costs you to lose a message versus how much it costs you to build (and staff) a process that actually looks at what lands in there. If nobody’s job is to check it, don’t bother building it — log and alert instead. What actually matters here A dead letter queue sounds like a safety net. Something fails, instead of vanishing it gets caught and set aside. That sounds strictly better than just dropping it, right? Usually. But “strictly better” only holds if someone actually does something with what gets caught. Otherwise you’ve just built a more expensive place for failures to be silently ignored. Before deciding whether to add one, three things actually decide the answer for you: Can you afford to lose this message? A password-reset email you can afford to lose (the user just clicks “resend”). A payment confirmation, an inventory decrement, or a compliance record you probably can’t. Do you have a real process for reprocessing? Not “someone could write a script” — an actual runbook, an owner, and a cadence for checking the queue. How often do messages actually fail, and why? A trickle of failures from a flaky third-party API is a different problem than a flood from a bad deploy that broke your message schema. Picture a mid-size online store. When someone places an order, a message goes onto a queue, and a background worker picks it up to send a confirmation email, update inventory, and notify the warehouse. One day, a teammate renames a field in the order message — from customerEmail to buyerEmail — and forgets that the email-sending worker still expects the old name. Every order message the worker touches now throws an exception. The basic shape of the pattern: most messages succeed, but the ones that don’t need somewhere to land instead of disappearing. Option A: No DLQ — just retry and log The simplest option: the worker retries a failed message a few times (in case it’s a temporary blip, like the database being briefly unavailable), and if it keeps failing, it logs the error and moves on. The message is gone for good after that. This is fine when losing a message is genuinely not a big deal, or when your failure rate is so low that a human glancing at error logs once in a while is enough of a safety net. It’s also the cheapest option — there’s no extra queue to provision, monitor, or explain to the next engineer who joins the team. The risk: “genuinely not a big deal” is a judgment call people get wrong under pressure. A retry-and-drop system quietly trains your team to treat every failure as disposable, including the ones that weren’t supposed to be. Option B: A DLQ with manual review Here, failed messages get routed to a separate queue after retries are exhausted. Nothing happens to them automatically — an engineer is expected to look at the DLQ periodically, figure out why messages landed there, and decide whether to fix and replay them or throw them away. This buys you two things retry-and-drop doesn’t: nothing is silently lost, and you get a natural place to notice patterns (if the DLQ suddenly fills with 10,000 messages at 2:14pm, that’s a signal something broke, not 10,000 unrelated coincidences). The catch is the word “periodically.” A DLQ with no owner and no alert threshold behaves exactly like Option A, except now you’re also paying for the extra queue. This is the single most common way teams end up with a DLQ that just accumulates messages for months. Someone built it with good intentions, nobody assigned who checks it, and it became a graveyard nobody visits. The uncomfortable relationship most DLQs end up with: an entity called “Engineer” is technically connected to it, but the connection is thin. Back at the online store: the DLQ fills up with every order-confirmation message from the moment the field got renamed. Nobody notices for three weeks, because nothing alerts on DLQ depth — it’s just a number nobody’s watching. When someone finally opens it during an unrelated audit, there are 40,000 messages sitting there, and the team has to figure out, after the fact, which customers never got a confirmation email and whether any orders were affected too. Option C: A DLQ with alerting and an owner This is Option B plus two additions: an alert that fires when the DLQ depth crosses a threshold or grows too fast, and a named owner (a person or a team, not “whoever notices”) responsible for triaging it on a schedule. Some teams go further and automate the common cases — for example, automatically replaying messages that failed due to a downstream timeout, since those are usually safe to retry once the dependency recovers. This is the version that actually delivers on the safety-net promise. It costs more to set up — someone has to build the alert, write the runbook, and agree to be paged — but it’s the only version where “we have a DLQ” is actually true in practice, not just in the architecture diagram. This is the journey Option B quietly sets your team up for. Option C is what prevents it. The rule of thumb Don’t ask “should we add a DLQ” in isolation — ask “who checks it, and how would we know if it’s filling up.” If you can’t answer both, you’re not choosing between Option A and Option B. You’re choosing between an honest Option A (log it, drop it, move on) and a fake Option B that costs more but behaves identically. So, in short: If losing the message is genuinely fine and failures are rare: skip the DLQ, log and move on. If losing the message matters, but you’re not ready to commit an owner and an alert: you’re better off spending that effort making the consumer itself more reliable (better retries, idempotency, fixing the root cause of failures) than building a DLQ that will quietly rot. If losing the message matters and you’re willing to name an owner and wire up an alert on queue depth: build it. This is the only version of the pattern that earns its keep. A dead letter queue is not a solution. It’s a deferral. The only question worth asking before you build one is whether you’ve also built the part where someone actually shows up. Related architecture architecture patternsdead letter queuedistributed systemsmessaging patternsreliability