architecture Essential Reading: Normal Accidents — Why Complex, Tightly Coupled Systems Fail by Design josedacruz, September 18, 2026September 18, 2026 TL;DR: Charles Perrow’s Normal Accidents argues that in sufficiently complex, tightly-connected systems, catastrophic failures aren’t rare flukes caused by bad luck or careless operators — they are a structural inevitability the system was built to eventually produce. For anyone running distributed systems, that reframes “we had an outage” from a… Continue Reading
architecture The Dead Letter Queue Pattern: Why Nobody Ever Looks Inside It josedacruz, September 18, 2026September 18, 2026 TL;DR: A dead letter queue (DLQ) is just a holding pen for messages your system tried to process and failed. It’s not a fix by itself — it just moves the failure somewhere quieter. Whether you actually need one depends on how much it costs you to lose a message… Continue Reading
architecture Designing a Webhook Delivery System: Why Your Retries Need an Idempotency Key josedacruz, September 13, 2026September 13, 2026 TL;DR: A webhook is just an HTTP POST your system sends to someone else’s server when something happens, but “just an HTTP POST” hides a lot of hard problems. Networks drop packets, receivers time out, and retries create duplicates — so a webhook delivery system that actually works needs retries… Continue Reading
architecture Designing a Distributed Job Scheduler: Why “Just Use Cron” Doesn’t Scale josedacruz, August 26, 2026August 26, 2026 TL;DR: A single cron job looks simple until the server it lives on goes down, and “just add a second cron server” quietly turns into a duplicate-execution bug. A real distributed job scheduler needs a shared source of truth for what’s due, a way for exactly one worker to claim… Continue Reading