All systems
Delivery

Transactional Outbox System Design

Saving a row and telling the world about it as one atomic act, when the database and the message broker cannot share a transaction.

internal/outboxinternal/models

1.Problem statement

An order is placed. A row is written and a webhook goes out. Those are two systems, and there is no transaction that spans both, so one of them happens first and the other can fail.

Publish first, then commit: the webhook succeeds, the commit fails, and a downstream system has been told about an order that does not exist. It acts on it. There is nothing to reconcile against, because the row was never written.

Commit first, then publish: the commit succeeds, the process is killed before the publish, and the order exists with nobody told. Nothing is logged anywhere, because from the process’s point of view nothing failed.

Both orderings are wrong, and which one is wrong in a worse way depends on the integration. The outbox pattern removes the choice: the message is written to a table in the same transaction as the business data, so it commits or rolls back with it, and a separate relay delivers committed messages afterwards. Either both happened or neither did.

The system has to be able to:

  • Write a message in the caller’s transaction, so it shares the fate of the data.
  • Refuse to enqueue outside a transaction, because that is the bug with extra steps.
  • Deliver committed messages from a relay, separately from the request.
  • Claim a message before delivering, so two relays do not deliver it twice.
  • Retry a failed delivery with backoff, and stop after enough attempts.
  • Keep delivered messages as a record of what was sent and when.
  • Be honest that delivery is at-least-once, and say what consumers must do about it.

2.System requirements

Functional requirements

  • An enqueue that takes the transaction and returns an error if it is not one.
  • A message row with a type, a payload, a status and an attempt count.
  • Four statuses: pending, claimed, delivered, failed.
  • A relay that claims pending messages, delivers them and records the outcome.
  • Backoff between attempts, and a cap after which a message is marked failed.
  • The table declared as a model, so migrations create it and backups include it.
  • Delivered messages retained, as the audit trail of what left the system.

Non-functional requirements

  • Atomic with the data. the message is in the same transaction as the row. This is the entire property, and everything else is in service of it.
  • At least once, stated plainly. a message can be delivered twice if the process dies between the send and the status update. Consumers must be idempotent, and a dedup key on the receiving end is what that means in practice.
  • Ordered enough. messages are relayed oldest first. Strict global ordering is not offered, because it would mean a single relay and no parallelism.
  • Durable. the outbox is a database table, so it inherits the database’s durability, replication and backups rather than needing its own.
  • Refuses the wrong usage. enqueuing without a transaction is rejected rather than accepted. It would work in every test and fail only in production, which is the worst possible failure mode to allow.

3.Capacity estimation

Numbers for a mid-sized deployment. They are here to size the thing, not to predict your traffic: change an assumption and the sums below move with it.

Assumptions

ParameterValue
Messages per second~30 average, 200 peak
Payload size~2 KB of JSON
Relay poll intervalseconds
Delivery target latencyunder a second at normal load
Retention of delivered rows30 days

Table growth

30 messages/s x 86,400 = ~2.6M rows/day
at ~2.2 KB per row = ~5.7 GB/day
x 30 days retention = ~170 GB

This is the number that decides whether the outbox needs pruning, and at this volume it very much does. A smaller project at 1 message/s is 190 MB a month and needs nothing.

Relay throughput

batch of 100 claimed per poll
delivery at ~50 ms each, 10 in parallel = ~0.5 s per batch
= ~200 messages/s per relay process

One relay keeps up with peak. More than one requires the claim step to be correct, which is why claiming is a status transition and not a read.

Added write cost

one extra INSERT inside an existing transaction
no extra round trip to a broker in the request path
the transaction is marginally longer, not fundamentally slower

Delivery latency

commit to next poll: up to the poll interval
plus delivery time
so seconds, not milliseconds

The outbox trades latency for correctness. Anything that genuinely needs sub-second delivery needs a different mechanism, and usually does not actually need it.

4.High level design

One extra insert inside the transaction, and one process reading committed rows.

Core components

  • Outbox table. declared as a model like any other, so AutoMigrate creates it and the backup writer includes it. Calling code still refers to it through the outbox package.
  • Enqueue. writes the message using the caller’s transaction. Takes the transaction as an argument and errors if handed anything else.
  • Claim. moves a batch from pending to claimed in one statement. This is what makes more than one relay safe.
  • Relay. delivers claimed messages and records the result. Its own process, or a scheduled job, independent of the request that created the message.
  • Status machine. pending, claimed, delivered, failed. Strings rather than an enum so a person reading the table can see what happened without a lookup.

Request flow

A row and its message, committing together

12345678HandlerTransactionOrder RowOutbox RowCommitRelayClaim BatchWebhook TargetDelivered or Retry
  1. 1A handler starts a transaction, because the outbox needs one and will say so if it does not get one.
  2. 2The business data is written: the order, its lines, whatever the operation is.
  3. 3The message is written in the same transaction. This is the whole design. From here, the row and the announcement of the row have exactly one fate between them.
  4. 4The transaction commits, or it does not. If it rolls back, the message rolls back with it, so nothing was announced that did not happen.
  5. 5Some time later, within seconds, the relay looks for pending messages. It is a separate process, so a crash in the request path after the commit changes nothing.
  6. 6A batch is claimed: pending to claimed in one statement. A second relay reading at the same moment sees nothing to claim, which is what makes running two of them safe.
  7. 7Each message is delivered. The attempt count rises so a persistently failing target backs off rather than being retried in a tight loop.
  8. 8The outcome is recorded. Delivered stays as the record of what was sent; a failure goes back for another attempt until the cap, after which it is marked failed and waits for a person. A process killed between the send and this update means the message is delivered again later, which is exactly why consumers need to be idempotent.

Data flow

  • The message payload is a snapshot of what was true at commit time. It does not re-read the row at delivery, because the row may have changed and the message is about what happened, not about what is.
  • Claiming is a status transition rather than a read, so concurrent relays partition the work instead of duplicating it.
  • Delivered rows are kept. They are the answer to "did we tell them, and when", which is the question asked during every integration dispute.
  • At-least-once is a consequence of the crash window between sending and recording. It cannot be closed without a transaction spanning the database and the target, which is the thing that does not exist.

5.Technology stack

ComponentWhat it is
Storagea database table, in the application database
Atomicitythe caller’s transaction
Statusespending, claimed, delivered, failed
Relayits own process or a scheduled job
Guaranteeat least once
Orderingoldest first, not strictly global

6.Data model

outbox_messages

The index that matters is on status and created_at, because every relay poll is a query for the oldest pending rows and it runs several times a minute forever.

ColumnHolds
idprimary key
typewhat happened, for example order.created
payloadJSON, a snapshot at commit time
statuspending, claimed, delivered or failed
attemptshow many deliveries have been tried
last_errorwhy the last attempt failed
created_atwhen the transaction committed it
delivered_atwhen it was accepted

Where it lives

  • In the application database on purpose. A separate store would need its own transaction, which is the problem being solved.
  • The table grows with every message and needs pruning at volume. Delivered rows older than the retention period are the ones to go.
  • It is included in backups because it is declared as a model, which also means a restore brings undelivered messages back and they will be delivered.

7.Low level design

Core types

outbox.Enqueueinternal/outbox/outbox.go

Writes a message using the transaction it is handed. The signature takes the transaction rather than a database handle, so the correct usage is the only one that compiles cleanly.

outbox.ErrNoTransaction

Returned when enqueue is given something that is not a transaction. Refused rather than tolerated, because enqueuing outside a transaction is the commit-then-publish bug and it fails only in production.

outbox.Message

An alias for the model type. The table is declared in internal/models with every other table so migrations and backups cover it, while calling code still reads outbox.Message.

outbox.Relayinternal/outbox/relay.go

Claims a batch, delivers each message, records the outcome, backs off on failure.

Design principles applied

  • Refuse the usage that fails silently. enqueue without a transaction is not a degraded version of the right thing, it is the original bug. Returning an error is better than accepting it.
  • One table, declared like any other. putting it in models means AutoMigrate, backups and the studio all cover it with no special cases.
  • Claim before delivering. a status transition rather than a read is the difference between two relays sharing the work and two relays doing all of it.
  • Say at-least-once. the crash window is real and cannot be closed. Documenting it is what makes consumers idempotent instead of surprised.

Patterns

PatternWhere it is used
Transactional outboxthe message committed with the data
Polling publishera relay reading committed rows
Claim checka status transition partitioning work between relays
At-least-once deliverywith idempotent consumers as the stated requirement

8.Scalability and performance

  • The write path cost is one extra insert inside a transaction that already exists, which is as cheap as this problem gets.
  • Relay throughput scales by running more relays, which is safe because claiming is a status transition.
  • The table is the thing that needs managing. At thousands of messages a second it grows by gigabytes a day and needs a prune job.
  • The index on status and created_at is not optional. Without it, every poll is a scan of a table that only grows.
  • Delivery latency is bounded by the poll interval. Shortening it costs queries and buys seconds, and a listen-notify trigger removes the polling entirely when the latency genuinely matters.
  • Because the outbox lives in the application database, it inherits replication and backup rather than adding a second durable system to operate.

9.Bottlenecks and improvements

What breaks first

  • Table growth. retaining delivered messages is what makes the audit trail useful and what fills the disk. Both are true at once.
  • Polling latency. a message waits up to the poll interval before anybody looks at it, which is fine for a webhook and not for anything interactive.
  • A stuck claim. a relay that dies after claiming leaves rows in claimed that no other relay will touch, and they stay there until something notices.
  • Duplicate delivery. the crash window between sending and recording guarantees it will happen eventually, and a consumer that is not idempotent will act twice.
  • Head of line blocking. one target that is down and retrying can monopolise a relay batch, delaying messages for targets that are perfectly healthy.

What to do about it

  • Prune delivered rows on a schedule. a nightly job deleting delivered messages past the retention window, in batches. The retention length is an integration decision, not a technical one.
  • Use listen and notify where latency matters. the commit signals the relay directly, so delivery is immediate, with the poll kept as the fallback that makes it correct.
  • Time out claims. treat a row claimed longer than a threshold as pending again. The cost is a possible duplicate delivery, which at-least-once already allows for.
  • Require a dedup key at the consumer. the message identifier in the delivery, and a uniqueness check on the receiving side. This is the only real answer to duplicates.
  • Partition by target. claim per destination so a failing target backs off on its own without holding up anybody else’s messages.

Read next