All systems
Delivery

Scheduled Tasks System Design

Running something every night, exactly once, across however many replicas happen to be up.

internal/croninternal/jobs

1.Problem statement

Every application accumulates recurring work. Expire old sessions at midnight. Send the digest on Monday morning. Prune soft-deleted rows after thirty days. Reconcile with the payment provider hourly.

A system crontab does it, outside the application, in a different language, with no access to the models, no logging the application can see, and no record on any machine that is not that machine. It also stops existing the moment deployment becomes containers.

A ticker inside the process is worse in a specific way that only appears in production. It works perfectly on one instance. Scale to three replicas and the nightly job runs three times: three digests to each subscriber, three reconciliation passes, three prunes racing each other. Nothing errors, because each replica did exactly what it was told.

The requirement is not "run on a schedule". It is "run on a schedule, once, regardless of how many copies of the application are running", and that needs something that can be held by one replica at a time.

The system has to be able to:

  • Declare a schedule in the application, next to the work it runs.
  • Run each occurrence exactly once across all replicas.
  • Survive the replica holding the schedule dying, by having another take over.
  • Enqueue a job rather than doing the work inline, so a slow task does not delay the next tick.
  • Show the schedule, the last run and the next, in the admin.
  • Allow a manual run, for the case where last night failed.
  • Run on standard cron expressions, because that is the notation everybody already knows.

2.System requirements

Functional requirements

  • Schedules declared in Go with a cron expression and a task type.
  • A scheduler that enqueues rather than executes, so the work goes through the job queue.
  • A distributed lock, so exactly one replica schedules.
  • Automatic takeover when the lock holder goes away.
  • An admin page listing each task, its expression, its last run and its next.
  • A trigger button, enqueuing one occurrence immediately.
  • The scheduler disabled with no Redis, logged plainly rather than failing.

Non-functional requirements

  • Exactly once per occurrence. this is the whole requirement. A digest sent three times is worse than a digest not sent, because the second is noticed and fixed.
  • Survives a replica loss. the lock is leased, not held. A replica that dies releases it by not renewing, and another picks it up without an operator.
  • Scheduling is not executing. the scheduler enqueues. Running the work inline means a task that takes ten minutes delays the next tick and a crash loses the occurrence.
  • Visible. the last run and the next are in the admin, because the failure mode of scheduled work is that it quietly stops and nobody notices for a fortnight.
  • Honest about time zones. the expression is evaluated in a declared zone. "Midnight" without one is midnight in whatever zone the container happened to get.

3.Capacity estimation

Numbers for a mid-sized deployment. They are here to size the thing, not to predict your traffic: change an assumption and the sums below move with it.

Assumptions

ParameterValue
Scheduled tasks per project5 to 20
Finest useful granularityone minute
API replicas2 to 12
Lock leasetens of seconds, renewed
Heaviest tasknightly cleanup over millions of rows

Scheduler cost

20 tasks, evaluated once a minute
= 1,200 expression evaluations an hour
each one an arithmetic comparison

Negligible. The scheduler is cheap precisely because it only enqueues. All the real cost is in the workers.

Lock traffic

12 replicas, each attempting or renewing every few seconds
a handful of Redis operations per second, total

The midnight spike

8 tasks all written as 0 0 * * *
all enqueue in the same second
the worker pool absorbs them as a short backlog, not a stall

This is the argument for enqueuing rather than executing: eight simultaneous occurrences are a queue depth of eight, not eight things running in the scheduler at once.

Occurrences missed by downtime

no replica up at midnight = one occurrence missed
the schedule does not backfill: the next run is the next occurrence

Deliberate. A nightly task catching up on four missed nights at once is usually worse than skipping them, and the manual trigger exists for when it is not.

4.High level design

Every replica runs a scheduler. Exactly one of them holds the lock, and only the holder enqueues.

Core components

  • Schedule registry. the declared tasks: an expression, a task type and a payload. In Go, next to the work, so adding one is a code change that gets reviewed.
  • Scheduler. evaluates the expressions and enqueues what is due. Runs in every replica, acts in one.
  • Distributed lock. a leased key in Redis. Whoever holds it schedules; whoever does not, waits. The lease is what makes a dead holder recoverable without intervention.
  • Job queue. where the occurrence goes. The scheduler’s output is a queued job, which inherits retries, timeouts and the dashboard from the queue.
  • Admin page. the tasks, their expressions, their last and next runs, and a trigger.

Request flow

Three replicas, one occurrence

12345678Replica 1Cron LockSchedulerJob QueueReplica 2Replica 3WorkerCron AdminTask Runs Once
  1. 1Every replica starts a scheduler and tries to take the cron lock. One of them gets it.
  2. 2Replica 2 does not. It keeps trying on an interval, so that if the holder disappears it is already a candidate rather than needing a deploy.
  3. 3Replica 3 likewise. The scheduler is running in all three and acting in one, which is what makes takeover automatic rather than operational.
  4. 4The holder renews its lease while it keeps running. A lease rather than a permanent key is the whole mechanism: a replica that is killed stops renewing, the key expires, and the next attempt succeeds.
  5. 5When an expression comes due the holder enqueues a job. It does not run the work, so a ten minute cleanup does not delay the next evaluation and a crash during it loses nothing.
  6. 6A worker picks the job up like any other, with the same retries, the same timeout and the same dashboard.
  7. 7The task runs once, because only one replica enqueued it.
  8. 8A manual trigger from the admin enqueues the same job directly, for the night that failed or the digest somebody needs now.

Data flow

  • The lock is a lease with an expiry, renewed by the holder. That is what distinguishes "another replica takes over in thirty seconds" from "a human notices tomorrow".
  • Schedules do not backfill. A missed occurrence stays missed, and the manual trigger is the deliberate way to make one up.
  • The scheduler enqueues, which means every property the job queue has applies to scheduled work for free: retries, timeout, dead letter, the dashboard.
  • Expressions are evaluated in a declared zone, so a nightly task does not drift by an hour twice a year.

5.Technology stack

ComponentWhat it is
Schedulerasynq periodic tasks
Coordinationa leased lock in Redis
Notationstandard cron expressions
Granularityone minute
Executionenqueued to the job queue, never run inline
Admina page listing tasks, last run, next run, and a trigger
Absent Redisscheduler disabled, logged

6.API design

Schedule administration

MethodEndpointWhat it does
GET/api/v1/admin/cronDeclared tasks with last and next run
POST/api/v1/admin/cron/:name/runEnqueue one occurrence now

7.Low level design

Core types

cron.Registerinternal/cron/cron.go

Declares a schedule: an expression, a task type and a payload. In code next to the work rather than in a configuration file, so adding one goes through review.

cron.Start

Starts the scheduler in this replica. Competes for the lock and acts only while holding it. With no Redis it logs that scheduling is off and returns.

CronHandler

The admin endpoints: list the tasks, trigger one. The trigger enqueues the same job the scheduler would, so a manual run is not a different code path.

Design principles applied

  • Schedule, do not execute. the scheduler enqueues and nothing more. Everything else follows from that: a slow task cannot delay a tick, and a scheduler crash cannot lose work that was already queued.
  • Assume more than one replica. a design that is correct on one instance and wrong on three is a design that is wrong, because the third instance arrives without a decision being made.
  • Lease, do not hold. a permanent lock needs a human when its holder dies. An expiring lease needs nobody.
  • Do not catch up. missed occurrences stay missed, and a manual trigger makes one up deliberately. Automatic backfill means a deploy outage sends four days of digests at once.

Patterns

PatternWhere it is used
Leader electionby leased lock, so exactly one replica schedules
Scheduler and executor splitthe tick enqueues, the worker runs
Declarative scheduleexpressions in code, next to the work
Manual overridethe same enqueue path, triggered by a person

8.Scalability and performance

  • Adding replicas does not add occurrences, which is the single property the design exists to provide.
  • The scheduler itself never becomes a bottleneck, because its work is an arithmetic comparison per task per minute.
  • All the real cost is in the workers, which scale independently of the schedule.
  • Tasks written as midnight all fire at midnight. Spreading the expressions turns one spike into a flat few minutes, and costs nothing.
  • A heavy nightly task over millions of rows should process in batches and be resumable, because it will eventually not finish inside its window.
  • Takeover time is the lease duration. Shorter means faster recovery and more lock traffic, and tens of seconds is the right end of that trade for nightly work.

9.Bottlenecks and improvements

What breaks first

  • Clock skew between replicas. a few seconds of drift can mean an occurrence evaluated twice at a boundary, with the lock the only thing preventing a double enqueue.
  • A task that outgrows its window. a nightly cleanup that takes nine hours overlaps the next night, and two copies of it then compete over the same rows.
  • Silent stoppage. the characteristic failure of scheduled work is that it stops and nobody finds out. No error is raised by something not happening.
  • Everything at midnight. eight tasks on the same expression make a spike that can push a worker pool into a backlog that lasts into the morning.
  • Time zone and daylight saving. an hourly task runs twice or not at all on the two days a year the clock changes, in whichever zone the expression is evaluated.

What to do about it

  • Make tasks idempotent by occurrence. key the work on the date it is for, so a second enqueue for the same night is a no-op. Cheap, and it makes the lock a performance feature rather than a correctness one.
  • Batch and checkpoint long tasks. process in batches, record progress, and let the next occurrence resume. Then overrunning is slow rather than broken.
  • Alert on a missing run. a heartbeat per task and an alert when the last run is older than the expression allows. This is the only way the silent failure becomes visible.
  • Spread the expressions. move the eight midnight tasks to eight different minutes. It is a one-line change and it removes the spike entirely.
  • Schedule in UTC. and convert for display. The twice-yearly ambiguity disappears, at the cost of a nightly task happening an hour later for half the year.

Read next