All systems
Delivery

Background Jobs System Design

Getting slow work out of the request path, and having it survive a retry, a duplicate and a deploy.

internal/jobsinternal/handlers

1.Problem statement

A request that sends an email, resizes an image, builds a PDF or calls a third-party API is a request held open by something that has nothing to do with the response. The user waits for work they did not ask to wait for, a connection and a database handle are held for seconds, and if the third party is slow then so is the application.

The obvious fix is a goroutine, and it is wrong in four ways that all look fine in development. Nothing bounds it, so a burst of sign-ups is a burst of goroutines each holding a database connection and a provider socket. Nothing retries it, so a provider down for sixty seconds loses that work permanently. A deploy in the middle drops whatever was in flight. And no error surfaces, because from the request handler’s point of view nothing failed: it already returned 200.

This project shipped exactly that bug. Every email left the API from a bare goroutine started by the request, while the job client that handles all of it properly was never called from anywhere.

A job queue replaces the goroutine with a durable record. The work is written down, a worker picks it up, failures retry with backoff, and the thing that is holding up delivery is visible rather than lost.

The system has to be able to:

  • Enqueue work from a handler and return immediately.
  • Survive a process restart: queued work is not in memory.
  • Retry with exponential backoff, a bounded number of times.
  • Time out a stuck task so it does not hold a worker slot forever.
  • Refuse a duplicate enqueue of the same logical task within a window.
  • Separate queues by urgency, so a backlog of thumbnails does not delay a password reset.
  • Show what is queued, running, retrying and dead, with the error.
  • Run with no queue at all, in a deployment that has no Redis.

2.System requirements

Functional requirements

  • An enqueue call per job type, typed rather than a map of strings.
  • Workers registered per task type, started as their own process or alongside the API.
  • Five retry attempts by default, with exponential backoff.
  • A five minute task timeout by default, overridable per enqueue.
  • An idempotency key with a 24 hour window, deduplicating repeat enqueues.
  • Delayed enqueue, for work that should happen later rather than now.
  • Three queues: critical, default and low.
  • Completed tasks retained for 24 hours so a success can be inspected, not just a failure.
  • An admin dashboard over the queues, with retry and delete.

Non-functional requirements

  • Durable. enqueued work is in Redis, not in a goroutine. A deploy mid-flight loses nothing that was accepted.
  • Bounded. worker concurrency is configured, so load is a queue depth rather than a connection exhaustion. This is the single biggest difference from a goroutine.
  • At least once. a task can run twice if a worker dies after doing the work and before acknowledging. Handlers must be idempotent, and the framework says so rather than implying exactly-once.
  • Observable. a failed job and its error are in the dashboard. The failure mode of the goroutine version was that nobody ever found out.
  • Optional. with no Redis configured, jobs, cron and the cache are all disabled and the application runs. The alternative is requiring Redis for a project that wanted SQLite and a single binary.

3.Capacity estimation

Numbers for a mid-sized deployment. They are here to size the thing, not to predict your traffic: change an assumption and the sums below move with it.

Assumptions

ParameterValue
Jobs enqueued per second~45 average, 280 peak
Average job duration400 ms
Slowest common jobimage processing, 3 to 8 seconds
Worker concurrency per process10
Payload size~1 KB
Retry attempts5, exponential

Worker capacity

10 concurrent slots / 0.4 s average = 25 jobs/s per worker process
280 jobs/s at peak / 25 = 12 worker processes
or fewer processes with higher concurrency per process

Concurrency is the number to turn, not the process count, until the database connection pool becomes the limit. Each concurrent job may hold a connection.

Queue memory

a backlog of 100,000 jobs x 1 KB = 100 MB
plus completed retention: 45 jobs/s x 86,400 s = ~3.9M/day
at 24 hours retention and 1 KB that is ~3.9 GB

Retention is the number that surprises people. It is what makes a success inspectable, and it is also what fills Redis. Lower it per task type for anything high volume.

Backoff schedule

attempt 1 immediate, then roughly 1 s, 4 s, 15 s, 60 s
five attempts spans about 80 seconds of downstream outage
longer outages need a higher MaxRetry on that task type

Five was chosen to ride out a short blip without filling the dead queue with hours-old work that is no longer worth doing.

Database connections

12 workers x 10 concurrency = 120 potential connections
plus the API processes
against a pool that is usually configured at 25 per process

This is the real constraint on scaling workers, and it is the thing that breaks when somebody raises concurrency to fix a backlog.

4.High level design

A client that writes the task, a broker that holds it, workers that run it, and a dashboard over the broker.

Core components

  • Job client. what a handler calls. One typed method per job type, so a payload cannot be wrong at the call site and discovered in the worker.
  • Broker. Redis, through asynq. Holds the pending, scheduled, retrying, completed and dead sets.
  • Workers. a process with a handler registered per task type and a concurrency setting. Runs as its own container in production and alongside the dev server locally.
  • Retry policy. attempts, backoff and timeout, with defaults that each task type can override. Set at enqueue rather than in the worker, so the caller decides how important the work is.
  • Idempotency. a key plus a window. The same logical task enqueued twice inside the window is refused, and the refusal is a success for the caller because the original is already on its way.
  • Dashboard. the admin view over the queues: depth, failures with their errors, retry and delete.

Request flow

A handler enqueuing work, and a worker running it

12345678ClientHandlerEnqueueRedis QueueIdempotency KeyWorker PoolTask HandlerRetry or DeadJobs Dashboard
  1. 1A request arrives that would otherwise do slow work inline: a sign-up that sends a welcome email, an upload that needs thumbnails.
  2. 2The handler calls the typed enqueue method and returns. The response does not wait for the work, and the user does not wait for a mail provider.
  3. 3If the caller supplied an idempotency key, a duplicate inside the window is refused. The caller treats that as success, because the original task is already queued and running it twice is the thing being avoided.
  4. 4The task is written to Redis with its retry count, its timeout, its queue and its retention. It is now durable: a deploy, a crash or a restart does not lose it.
  5. 5A worker process pulls from the queues in priority order. Critical before default before low, so a backlog of thumbnails never delays a password reset.
  6. 6The registered handler for that task type runs, inside the timeout. It must be safe to run twice, because at-least-once delivery is what a durable queue buys and exactly-once is not available.
  7. 7On failure the task goes back with exponential backoff, up to five attempts by default. After the last attempt it moves to the dead set rather than disappearing.
  8. 8The dashboard shows depth, failures and the error text, with retry and delete. The whole point of the design over a goroutine is that a failure is a row somebody can see rather than a log line nobody wrote.

Data flow

  • A payload is small and references data rather than carrying it. An image job carries an upload identifier, not the image.
  • Nothing in a payload can be assumed still to exist. By the time a task runs the row may have been deleted, so a handler loads and tolerates absence.
  • Retry decisions are made at enqueue, not in the worker, because the caller is the one who knows whether this work is worth five attempts.
  • Completed tasks are retained so a success can be inspected. That retention is also the main consumer of Redis memory at volume.

5.Technology stack

ComponentWhat it is
Queueasynq over Redis
Queuescritical, default, low
Defaults5 retries, 5 minute timeout, 24 hour idempotency window
Retentioncompleted tasks kept 24 hours
Workerstheir own process in production, with the dev server locally
Dashboardadmin pages over the broker
Absent Redisjobs, cron and cache all disabled, application runs

6.API design

Queue administration

MethodEndpointWhat it does
GET/api/v1/admin/jobsQueue depths and worker state
GET/api/v1/admin/jobs/failedDead tasks with their last error
POST/api/v1/admin/jobs/:id/retryRequeue one dead task
DELETE/api/v1/admin/jobs/:idDiscard a dead task

Enqueuing from a handler, with the duplicate case handled

err := h.Jobs.EnqueueSendEmail(ctx, jobs.SendEmailPayload{
To: user.Email,
Template: "welcome",
Data: map[string]any{"name": user.Name},
}, jobs.EnqueueOption{
Queue: "critical",
IdempotencyKey: "welcome:" + user.ID,
})
// A duplicate inside the window is not a failure: the original is on its way.
if err != nil && !errors.Is(err, jobs.ErrDuplicateTask) {
return fmt.Errorf("queueing welcome mail: %w", err)
}

7.Low level design

Core types

jobs.Clientinternal/jobs/jobs.go

Wraps the broker client with one typed method per job type. Typed on purpose: a string task name and a map payload moves every mistake from compile time to the worker.

EnqueueSendEmailEnqueueProcessImageEnqueueCleanup
jobs.EnqueueOption

Per-task overrides: queue, retries, timeout, delay, retention, idempotency key and window. Zero values fall back to the package defaults.

jobs.ErrDuplicateTask

Returned when the same idempotency key is enqueued inside its window. Named so callers can distinguish it, because treating it as an error would make a retry look like a failure.

mail_dispatch.gointernal/handlers/mail_dispatch.go

The single place mail leaves a handler. Added after the review found every email going out from a bare goroutine while the job client sat unused.

Design principles applied

  • One way out of a handler. all mail goes through one dispatch function. Twelve handlers each starting a goroutine is twelve chances to get bounding, retry and shutdown wrong.
  • Typed payloads. a wrong field is a compile error rather than a nil map access inside a worker at three in the morning.
  • Say at-least-once out loud. documenting the real guarantee is what makes handlers idempotent. Implying exactly-once produces handlers that charge a card twice.
  • Defaults tuned for safety. five retries and a five minute timeout are about surviving a blip without hoarding stale work, not about maximum throughput.

Patterns

PatternWhere it is used
Work queuea durable record instead of an in-process goroutine
Exponential backoffretries that do not hammer a struggling dependency
Dead letterexhausted tasks kept rather than dropped
Idempotency keydeduplicating the enqueue, not just the execution
Priority queuesurgency separated so a backlog is not a blockage

8.Scalability and performance

  • Workers scale horizontally: more processes consume the same queues with no coordination.
  • The first number to turn is concurrency per process, and the first thing it breaks is the database connection pool. Worker concurrency times worker count is a connection count.
  • Priority queues mean a backlog is localised. Thumbnails piling up does not delay a password reset, which is the difference between a slow feature and a broken one.
  • Completed retention dominates Redis memory at volume. It should be lowered per task type for anything high frequency.
  • A long running job should be split. One task that takes an hour is a task that cannot be retried cheaply and that blocks a worker slot for an hour.
  • With no Redis the application degrades to doing nothing in the background, which is correct for a small single-binary deployment and documented rather than surprising.

9.Bottlenecks and improvements

What breaks first

  • Connection pool exhaustion. raising worker concurrency to clear a backlog takes connections from the API, and the symptom shows up in request latency rather than in the queue.
  • Non-idempotent handlers. at-least-once plus a handler that charges a card is a double charge, and it happens on the day the worker is killed mid-task.
  • Payloads carrying data. an image in a payload is an image in Redis, several times over with retries, and a queue that falls over on memory rather than depth.
  • Poison tasks. a task that fails deterministically burns five attempts and a backoff schedule every time, and a flood of them crowds out real work.
  • Redis as a single point. the broker is the durable store. Losing it loses the queue, and no amount of worker redundancy helps.

What to do about it

  • Size the pool for workers plus API. count the real maximum: worker processes times concurrency, plus API processes times their pool. Then set the database limit above it.
  • Make handlers idempotent with a key. a processed-tasks table or a unique constraint on the effect, so the second run is a no-op rather than a second charge.
  • Pass references, not payloads. an identifier and let the handler load. The row is the source of truth and it may have changed since the enqueue anyway.
  • Separate queue for the known-flaky. a task type that fails often belongs on its own queue, where its backoff cannot crowd out anything that matters.
  • Persist and replicate Redis. append-only persistence and a replica, because the queue is durable storage whether or not it was planned as such.

Read next