Skip to content
José da Cruz

Enterprise Architect, and Life Thinker

José da Cruz

Enterprise Architect, and Life Thinker

Designing a Distributed Job Scheduler: Why “Just Use Cron” Doesn’t Scale

josedacruz, August 26, 2026August 26, 2026

TL;DR: A single cron job looks simple until the server it lives on goes down, and “just add a second cron server” quietly turns into a duplicate-execution bug. A real distributed job scheduler needs a shared source of truth for what’s due, a way for exactly one worker to claim a job at a time, and a plan for what happens when a worker dies mid-run. None of that requires a huge system — it just requires replacing a few wrong assumptions with the right ones.

Every team eventually needs something to run on a schedule: a nightly report, a reminder email, a job that reconciles inventory counts. The first version is almost always a single line in crontab on whatever server is closest at hand. That works great for months. Then the team grows, the server gets replaced, or someone adds a second instance for redundancy — and suddenly “the scheduled job” isn’t simple anymore. Here are the myths that trip teams up along the way, and what’s actually true instead.

Myth: “Cron on one server is fine — it’s simple and it works.”

It works right up until that one server has a bad night. Maybe it gets rebooted for a security patch, maybe the disk fills up, maybe it’s replaced during a routine infrastructure upgrade and nobody remembers to copy the crontab over. Cron itself has no concept of “this job didn’t run, someone should know.” If the server is down at 2am when the job was supposed to fire, the job simply never happens, and the first sign of trouble is a customer asking why they never got their invoice. A single cron job isn’t wrong for a low-stakes, easily-noticed task. But treating it as your scheduling strategy for anything the business actually depends on is a bet that nothing will ever go wrong with that one box, forever.

Myth: “Just run cron on multiple servers for redundancy.”

This is the natural next step, and it’s where things get worse instead of better. If Server A and Server B both have the same crontab entry, and both are healthy at 2am, they will both decide the job is due and both run it. Nothing stops them from talking to each other, because plain cron doesn’t know the other server exists. So instead of “sometimes the job silently doesn’t run,” you now have “sometimes the job silently runs twice.” A nightly report emailed twice is annoying. A reconciliation job that adjusts inventory twice, or a billing job that charges a customer twice, is a real incident.

Sequence diagram showing two cron servers both checking the jobs table and both running the same nightly report, resulting in it being sent twice
Two servers, no coordination between them: both see the same due job and both run it.

The root cause isn’t “we have two servers.” It’s that neither server has any way to tell the other “I’ve got this one, back off.” Redundancy without coordination just means you now have two independent chances for the same mistake to happen.

Myth: “A simple lock will fix duplicate runs.”

This is usually the next fix a team reaches for, and it’s a real improvement, but it’s easy to get wrong in ways that only show up later. A common first attempt: before running the job, a server writes a row to a “lock” table; if the row already exists, it skips the job. That handles the easy case. But what happens if the server that holds the lock crashes mid-job? If the lock has no expiration, the job is now stuck “locked” forever, and it will never run again until a human notices and manually deletes the row. If the lock does expire, you need to pick a timeout, and that timeout is a guess — too short and a slow-but-healthy job gets its lock stolen out from under it while it’s still running; too long and a genuinely crashed job sits unclaimed for ages before anyone else picks it up. Locking is necessary, but a lock by itself doesn’t answer “what do we do when the lock holder disappears,” and that question is the actual hard part of distributed scheduling.

Myth: “A distributed scheduler is way more complex than we actually need.”

This is the myth that keeps teams stuck patching cron indefinitely, and it’s usually based on imagining something far heavier than what’s actually required. You don’t need to build (or even run) a huge platform. You need three things: a shared table that’s the single source of truth for what jobs exist and when they’re next due, a lease instead of a permanent lock (a claim that automatically expires after N minutes unless the worker renews it), and a pool of workers that pull jobs from that table instead of each deciding independently what’s due. That’s it. The lease is the piece that solves the crash problem from the myth above: if a worker dies mid-job, its lease simply expires on its own, and another worker picks the job back up — no stuck lock, no manual cleanup, no guessing.

State diagram showing a job moving from Scheduled to Leased to Running, then to Done or Failed, with a lease-expiry path back to Scheduled and a retry path from Failed back to Scheduled
Every job’s life is one of these small number of states — the lease-expiry path is what makes crashes self-healing instead of stuck.

Notice what’s missing from that list: no consensus algorithm, no leader election protocol, no custom infrastructure most teams need to hand-roll. A plain relational database with row-level locking (most databases support something like “claim this row if nobody else has, atomically”) can implement leases just fine at moderate scale. If your job volume is small, this is a few hundred lines of code, not a distributed-systems research project.

Myth: “As long as jobs are idempotent, none of this matters.”

Idempotency — designing a job so that running it twice has the same effect as running it once — is genuinely valuable, and it’s worth building into any job that touches money, inventory, or anything a customer would notice. But it’s a safety net, not a substitute for coordination. An idempotent billing job that runs twice won’t double-charge anyone, which is great. It will still burn double the compute, double the database load, and double the third-party API calls (some of which cost real money per call and don’t care that the second charge attempt was harmlessly rejected). Idempotency limits the blast radius of a coordination failure. It doesn’t make the coordination failure free, and it does nothing at all for the opposite problem — a job that silently never runs because the one server holding it went down.

The rule of thumb

If a scheduled job would cause a real problem the one time it runs twice, or the one time it silently doesn’t run at all, plain cron on a single box — or worse, plain cron on several uncoordinated boxes — isn’t a scheduling strategy, it’s a bet. The fix isn’t a heavyweight platform. It’s a shared table of what’s due, leases instead of permanent locks so crashes self-heal, and idempotent job logic as backup insurance rather than the whole plan. Once those three pieces are in place, “just use cron” stops being the risky shortcut and the distributed version stops feeling like overkill — it’s just the version that actually matches how many servers you really have.

Related

architecture case studydistributed systemsjob schedulingreliabilitysystem design

Post navigation

Previous post
Next post

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

  • architecture (154)
  • artificial-intelligence (6)
  • books (2)
  • courses (4)
  • decision-frameworks (4)
  • finances (1)
  • java (8)
  • life-improvement (20)
  • metrics (3)
  • observability (2)
  • puzzles and challenges (3)
  • reviews (1)
  • security (8)
  • spring-boot (36)
  • Systems Thinking (3)
  • GitHub Weekly: MCP Tools Trending Right Now (September 28 – October 4, 2026)
  • The Bulkhead Pattern: Why One Slow Dependency Shouldn’t Sink the Whole Ship
  • Essential Reading: A Philosophy of Software Design — Fighting Complexity One Module at a Time
  • GitHub Weekly: Top 10 Trending Repos (September 21 – September 27, 2026)
  • The Volunteer Tax: Why Raising Your Hand in a Meeting Makes You the Permanent Owner
©2026 José da Cruz | WordPress Theme by SuperbThemes