All systems
Delivery

Webhooks System Design

Receiving a call from someone else’s system and believing it, and making one to theirs that survives their outage.

internal/webhooksinternal/outbox

1.Problem statement

Webhooks are two systems with the same name. Inbound, somebody else’s server posts to an endpoint of yours: a payment succeeded, a repository was pushed. Outbound, your server posts to theirs when something happens here.

The inbound problem is belief. The request arrives unauthenticated by any session, from an address you do not control, and it says a payment was captured. Acting on an unverified webhook means anybody who learns the URL can tell you they paid. Verification is a signature, and getting it wrong is subtle: comparing with string equality leaks timing, verifying a parsed and re-serialised body verifies something other than what was sent, and accepting any timestamp lets a captured request be replayed forever.

The inbound problem also includes duplicates. Every sender worth integrating with retries, so the same event will arrive twice, and a handler that is not idempotent will act twice.

The outbound problem is that the receiver will be down. A send inside the request that created the event either delays the response or is lost on failure, and a retry loop in a goroutine loses everything on deploy. Outbound webhooks are an at-least-once delivery problem, which is why they belong behind the outbox and the job queue rather than in a handler.

The system has to be able to:

  • Verify an inbound signature in constant time, against the exact bytes received.
  • Reject a request whose timestamp is outside a tolerance, so a captured request cannot be replayed.
  • Deduplicate by event identifier, so a retried delivery is processed once.
  • Respond quickly and do the work in a job, because senders time out and then retry.
  • Support the signature schemes real providers use, not one invented here.
  • Send outbound events from the outbox, so they commit with the data that caused them.
  • Retry an outbound delivery with backoff, and show what failed and why.

2.System requirements

Functional requirements

  • A verifier interface with implementations per provider.
  • An HMAC-SHA256 verifier reading a hex signature from a named header.
  • A Stripe-style verifier parsing a timestamped signature header with a five minute tolerance.
  • A GitHub-style verifier for its header and prefix convention.
  • Raw body capture before parsing, so the verified bytes are the received bytes.
  • A dedup table keyed on provider and event identifier.
  • Inbound handlers that acknowledge and enqueue rather than processing inline.
  • Outbound events enqueued through the outbox, signed on the way out.
  • Redelivery from the admin, for a webhook the receiver lost.

Non-functional requirements

  • Constant-time comparison. a signature compared with string equality leaks its prefix through timing. The comparison is the one line in the whole system where that matters, so it is the one line that must not be ordinary.
  • Verify the bytes, not the object. parsing and re-serialising produces different bytes and therefore a different signature. The raw body is captured before anything touches it.
  • Bounded replay window. five minutes. A captured request outside it is refused, so an intercepted webhook has a short life rather than an unlimited one.
  • Fast acknowledgement. senders time out in seconds and then retry. A handler that does the work inline turns one event into several deliveries of the same event.
  • Idempotent by construction. the dedup check is in the receive path, not left to each handler, because every handler would otherwise need to remember.
  • At least once outbound. the same guarantee the outbox gives. Receivers are told to expect duplicates and given an identifier to deduplicate on.

3.Capacity estimation

Numbers for a mid-sized deployment. They are here to size the thing, not to predict your traffic: change an assumption and the sums below move with it.

Assumptions

ParameterValue
Inbound webhooks per second~15 average, 150 peak during a provider retry burst
Outbound events per second~30
Subscribers per outbound event1 to 5
Payload size~3 KB
Signature verification costmicroseconds
Receiver timeout10 seconds

Inbound acknowledgement

verify: microseconds
dedup check: one indexed read, ~1 ms
enqueue: ~1 ms
total response time: single-digit milliseconds

The number that matters, because it is what keeps the sender from timing out and retrying. Doing the work inline would make it seconds.

Dedup table

15/s x 86,400 = ~1.3M rows/day
at ~100 bytes = ~130 MB/day
a 7 day window is ~900 MB and needs pruning

Seven days is far longer than any sender retries for. The window only needs to outlast the longest retry schedule you integrate with.

Outbound fan-out

30 events/s x 3 subscribers = 90 deliveries/s
at ~200 ms per delivery, 20 in parallel = 100/s capacity
one relay keeps up; a slow receiver is what breaks it

A receiver that is down

90 deliveries/s accumulating while a receiver is out for 1 hour
if they are one of three subscribers: 30/s x 3,600 = 108,000 queued
at 3 KB each: ~320 MB of pending messages

This is why outbound delivery needs a per-receiver circuit breaker. Otherwise one dead endpoint fills the queue and delays everybody else.

4.High level design

Inbound: capture, verify, deduplicate, acknowledge, enqueue. Outbound: outbox, sign, deliver, retry.

Core components

  • Raw body capture. middleware storing the request bytes before anything parses them. Without this, verification checks a re-serialisation and will disagree with the sender.
  • Verifier. a function per provider. HMAC in a header, Stripe’s timestamped scheme, GitHub’s prefixed one. Per provider because providers differ and inventing a scheme is not an option when the sender is not yours.
  • Dedup store. provider plus event identifier, with a window. Checked before the handler runs, so idempotence is a property of the pipeline.
  • Inbound handler. acknowledges and enqueues. The processing is a job, which gives it retries and a dashboard and keeps the response fast.
  • Outbound via outbox. the event is written in the transaction that caused it, so a webhook never announces something that was rolled back.
  • Signer. signs outgoing payloads with the subscriber’s secret, in the same scheme we ask of others.
  • Redelivery. the admin view of what was sent, and a button to send it again for the receiver who lost it.

Request flow

An inbound webhook, verified and deduplicated

12345678ProviderCapture Raw BodyVerify SignatureTimestamp WindowDedup Check200 ImmediatelyJob QueueHandler Job401 Rejected
  1. 1A provider posts an event. The request carries no session and no user, only a signature, so the signature is the entire basis for believing any of it.
  2. 2The raw bytes are captured before parsing. Verifying a parsed and re-serialised body verifies something the sender never signed, and the mismatch is intermittent and maddening.
  3. 3The signature is recomputed over the received bytes with the shared secret and compared in constant time. Where the scheme includes a timestamp, it must be inside a five minute tolerance, so a captured request cannot be replayed indefinitely.
  4. 4A failure is a 401 and nothing else. No processing, no logging of the payload as if it were real, no partial handling.
  5. 5A verified request is checked against the dedup store on provider and event identifier. Every sender worth integrating with retries, so the same event will arrive twice, and this is where that stops being a problem.
  6. 6A 200 goes back straight away. Senders time out in seconds and then retry, so a slow acknowledgement manufactures duplicate deliveries of an event that was received correctly.
  7. 7The real work is enqueued, which gives it retries, a timeout, a dead letter queue and a dashboard, instead of happening in a request that has already been answered.
  8. 8The job runs and does the work. It can fail and retry without the provider knowing or caring, which is the separation the whole arrangement exists to create.

Data flow

  • The signature is over the exact bytes. This is the single most common integration bug and the raw body capture is the only fix.
  • The dedup identifier comes from the sender’s event id. Hashing the body instead would treat a legitimately repeated event as a duplicate.
  • Outbound events go through the outbox, so a webhook announcing an order cannot be sent for a transaction that rolled back.
  • Delivered outbound webhooks are retained with their response status, because that record is what every integration dispute is settled with.

5.Technology stack

ComponentWhat it is
Inbound verificationHMAC-SHA256, constant-time comparison
Schemesgeneric HMAC header, Stripe timestamped, GitHub prefixed
Replay tolerance5 minutes
Deduplicationprovider plus event id, in a table with a window
Inbound processingacknowledge then enqueue
Outboundtransactional outbox, signed, retried with backoff
SpecStandard Webhooks for outbound, via grit plugin add webhooks

6.API design

Inbound

MethodEndpointWhat it does
POST/api/v1/webhooks/stripeVerified with the timestamped scheme
POST/api/v1/webhooks/githubVerified with the prefixed header scheme
POST/api/v1/webhooks/:providerGeneric HMAC header verification

Outbound administration

MethodEndpointWhat it does
GET/api/v1/admin/webhooksSubscriptions and their secrets
GET/api/v1/admin/webhooks/deliveriesWhat was sent, with response status
POST/api/v1/admin/webhooks/deliveries/:id/redeliverSend it again

Registering an inbound endpoint

webhooks.Register(r, webhooks.Endpoint{
Path: "/webhooks/stripe",
Verify: webhooks.StripeVerifier(cfg.StripeWebhookSecret),
// Acknowledge, then let a job do the work: the sender times out
// in seconds and a timeout becomes a duplicate delivery.
Handle: func(c *gin.Context, event webhooks.Event) error {
return jobs.EnqueueStripeEvent(c, event.ID, event.Raw)
},
})

7.Low level design

Core types

webhooks.Endpointinternal/webhooks/webhooks.go

A path, a verifier and a handler. The verifier is a field rather than a convention, because every provider signs differently.

webhooks.HMACVerifier

Validates a hex HMAC-SHA256 in a named header. The common case, and what most small providers do.

webhooks.StripeVerifier

Parses a header of the form t=<unix>,v1=<hex>, recomputes over timestamp and body, and enforces a five minute tolerance.

Dedup store

Provider plus event identifier with a window. In the receive path rather than in each handler, so no handler can forget.

Design principles applied

  • Verify first, parse second. nothing in an unverified payload is information. Parsing before verifying means processing attacker-controlled structure.
  • Constant time, always. the comparison is not a detail. An early-exit comparison leaks the signature a byte at a time to anybody willing to measure.
  • Acknowledge fast. the sender’s timeout is the real deadline. Missing it converts one event into several, which is a correctness problem and not a latency one.
  • Outbound belongs in the outbox. a webhook is an announcement about a committed fact. Sending it from a request that might roll back announces things that did not happen.

Patterns

PatternWhere it is used
Signature verificationHMAC over the exact received bytes
Replay windowa timestamp inside the signed material
Idempotent receiverdedup on the sender’s event id
Acknowledge and enqueuethe response decoupled from the work
Transactional outboxoutbound events committed with their cause

8.Scalability and performance

  • Inbound capacity is an acknowledgement rate, and acknowledgement is a verification plus two cheap operations, so it scales with ordinary API replicas.
  • A provider retry burst is the real inbound load spike: an outage on their side ends with everything arriving at once, and the dedup check is what makes that harmless.
  • The dedup table grows steadily and needs pruning. The window only has to outlast the longest retry schedule you integrate with.
  • Outbound delivery scales by running more relays, with the claim step keeping them from duplicating work.
  • One dead receiver is the thing that breaks outbound at scale, because its retries accumulate in the same queue as everybody else’s deliveries.
  • Retaining deliveries with response codes is what makes redelivery and dispute resolution possible, and it is also the table that grows fastest.

9.Bottlenecks and improvements

What breaks first

  • A dead receiver filling the queue. one endpoint that is down for an hour accumulates its share of every event, and its retries compete with deliveries to healthy receivers.
  • Slow inbound handlers. work done inline means the sender times out and retries, so a slow handler creates the duplicates that then make it slower.
  • Secret rotation. changing a signing secret invalidates every in-flight delivery, and a window where both are accepted has to exist or deliveries fail during the change.
  • Dedup table growth. millions of rows a day, all of them useless after the retry window has passed.
  • Verified but untrusted content. a valid signature proves who sent it, not that the payload is reasonable. The amounts and identifiers inside still need checking.

What to do about it

  • Circuit breaker per receiver. stop delivering to an endpoint that is failing consistently, retry on a long interval, and notify the subscriber. One dead endpoint stops being everybody’s problem.
  • Separate queue for inbound processing. so a burst of provider retries is a backlog on its own queue rather than a delay to password resets.
  • Accept two secrets during rotation. verify against the current and the previous for an overlap period, then drop the old one.
  • Prune the dedup table. a scheduled job deleting rows past the window. Cheap, and the table is otherwise unbounded.
  • Validate the payload after verifying it. a signature authenticates the sender. The contents still get the same schema and sanity checks as any other input.

Read next