Webhooks System Design
Receiving a call from someone else’s system and believing it, and making one to theirs that survives their outage.
internal/webhooksinternal/outbox1.Problem statement
Webhooks are two systems with the same name. Inbound, somebody else’s server posts to an endpoint of yours: a payment succeeded, a repository was pushed. Outbound, your server posts to theirs when something happens here.
The inbound problem is belief. The request arrives unauthenticated by any session, from an address you do not control, and it says a payment was captured. Acting on an unverified webhook means anybody who learns the URL can tell you they paid. Verification is a signature, and getting it wrong is subtle: comparing with string equality leaks timing, verifying a parsed and re-serialised body verifies something other than what was sent, and accepting any timestamp lets a captured request be replayed forever.
The inbound problem also includes duplicates. Every sender worth integrating with retries, so the same event will arrive twice, and a handler that is not idempotent will act twice.
The outbound problem is that the receiver will be down. A send inside the request that created the event either delays the response or is lost on failure, and a retry loop in a goroutine loses everything on deploy. Outbound webhooks are an at-least-once delivery problem, which is why they belong behind the outbox and the job queue rather than in a handler.
The system has to be able to:
- Verify an inbound signature in constant time, against the exact bytes received.
- Reject a request whose timestamp is outside a tolerance, so a captured request cannot be replayed.
- Deduplicate by event identifier, so a retried delivery is processed once.
- Respond quickly and do the work in a job, because senders time out and then retry.
- Support the signature schemes real providers use, not one invented here.
- Send outbound events from the outbox, so they commit with the data that caused them.
- Retry an outbound delivery with backoff, and show what failed and why.
2.System requirements
Functional requirements
- A verifier interface with implementations per provider.
- An HMAC-SHA256 verifier reading a hex signature from a named header.
- A Stripe-style verifier parsing a timestamped signature header with a five minute tolerance.
- A GitHub-style verifier for its header and prefix convention.
- Raw body capture before parsing, so the verified bytes are the received bytes.
- A dedup table keyed on provider and event identifier.
- Inbound handlers that acknowledge and enqueue rather than processing inline.
- Outbound events enqueued through the outbox, signed on the way out.
- Redelivery from the admin, for a webhook the receiver lost.
Non-functional requirements
- Constant-time comparison. a signature compared with string equality leaks its prefix through timing. The comparison is the one line in the whole system where that matters, so it is the one line that must not be ordinary.
- Verify the bytes, not the object. parsing and re-serialising produces different bytes and therefore a different signature. The raw body is captured before anything touches it.
- Bounded replay window. five minutes. A captured request outside it is refused, so an intercepted webhook has a short life rather than an unlimited one.
- Fast acknowledgement. senders time out in seconds and then retry. A handler that does the work inline turns one event into several deliveries of the same event.
- Idempotent by construction. the dedup check is in the receive path, not left to each handler, because every handler would otherwise need to remember.
- At least once outbound. the same guarantee the outbox gives. Receivers are told to expect duplicates and given an identifier to deduplicate on.
3.Capacity estimation
Numbers for a mid-sized deployment. They are here to size the thing, not to predict your traffic: change an assumption and the sums below move with it.
Assumptions
| Parameter | Value |
|---|---|
| Inbound webhooks per second | ~15 average, 150 peak during a provider retry burst |
| Outbound events per second | ~30 |
| Subscribers per outbound event | 1 to 5 |
| Payload size | ~3 KB |
| Signature verification cost | microseconds |
| Receiver timeout | 10 seconds |
Inbound acknowledgement
The number that matters, because it is what keeps the sender from timing out and retrying. Doing the work inline would make it seconds.
Dedup table
Seven days is far longer than any sender retries for. The window only needs to outlast the longest retry schedule you integrate with.
Outbound fan-out
A receiver that is down
This is why outbound delivery needs a per-receiver circuit breaker. Otherwise one dead endpoint fills the queue and delays everybody else.
4.High level design
Inbound: capture, verify, deduplicate, acknowledge, enqueue. Outbound: outbox, sign, deliver, retry.
Core components
- Raw body capture. middleware storing the request bytes before anything parses them. Without this, verification checks a re-serialisation and will disagree with the sender.
- Verifier. a function per provider. HMAC in a header, Stripe’s timestamped scheme, GitHub’s prefixed one. Per provider because providers differ and inventing a scheme is not an option when the sender is not yours.
- Dedup store. provider plus event identifier, with a window. Checked before the handler runs, so idempotence is a property of the pipeline.
- Inbound handler. acknowledges and enqueues. The processing is a job, which gives it retries and a dashboard and keeps the response fast.
- Outbound via outbox. the event is written in the transaction that caused it, so a webhook never announces something that was rolled back.
- Signer. signs outgoing payloads with the subscriber’s secret, in the same scheme we ask of others.
- Redelivery. the admin view of what was sent, and a button to send it again for the receiver who lost it.
Request flow
An inbound webhook, verified and deduplicated
- 1A provider posts an event. The request carries no session and no user, only a signature, so the signature is the entire basis for believing any of it.
- 2The raw bytes are captured before parsing. Verifying a parsed and re-serialised body verifies something the sender never signed, and the mismatch is intermittent and maddening.
- 3The signature is recomputed over the received bytes with the shared secret and compared in constant time. Where the scheme includes a timestamp, it must be inside a five minute tolerance, so a captured request cannot be replayed indefinitely.
- 4A failure is a 401 and nothing else. No processing, no logging of the payload as if it were real, no partial handling.
- 5A verified request is checked against the dedup store on provider and event identifier. Every sender worth integrating with retries, so the same event will arrive twice, and this is where that stops being a problem.
- 6A 200 goes back straight away. Senders time out in seconds and then retry, so a slow acknowledgement manufactures duplicate deliveries of an event that was received correctly.
- 7The real work is enqueued, which gives it retries, a timeout, a dead letter queue and a dashboard, instead of happening in a request that has already been answered.
- 8The job runs and does the work. It can fail and retry without the provider knowing or caring, which is the separation the whole arrangement exists to create.
Data flow
- The signature is over the exact bytes. This is the single most common integration bug and the raw body capture is the only fix.
- The dedup identifier comes from the sender’s event id. Hashing the body instead would treat a legitimately repeated event as a duplicate.
- Outbound events go through the outbox, so a webhook announcing an order cannot be sent for a transaction that rolled back.
- Delivered outbound webhooks are retained with their response status, because that record is what every integration dispute is settled with.
5.Technology stack
| Component | What it is |
|---|---|
| Inbound verification | HMAC-SHA256, constant-time comparison |
| Schemes | generic HMAC header, Stripe timestamped, GitHub prefixed |
| Replay tolerance | 5 minutes |
| Deduplication | provider plus event id, in a table with a window |
| Inbound processing | acknowledge then enqueue |
| Outbound | transactional outbox, signed, retried with backoff |
| Spec | Standard Webhooks for outbound, via grit plugin add webhooks |
6.API design
Inbound
| Method | Endpoint | What it does |
|---|---|---|
| POST | /api/v1/webhooks/stripe | Verified with the timestamped scheme |
| POST | /api/v1/webhooks/github | Verified with the prefixed header scheme |
| POST | /api/v1/webhooks/:provider | Generic HMAC header verification |
Outbound administration
| Method | Endpoint | What it does |
|---|---|---|
| GET | /api/v1/admin/webhooks | Subscriptions and their secrets |
| GET | /api/v1/admin/webhooks/deliveries | What was sent, with response status |
| POST | /api/v1/admin/webhooks/deliveries/:id/redeliver | Send it again |
Registering an inbound endpoint
webhooks.Register(r, webhooks.Endpoint{Path: "/webhooks/stripe",Verify: webhooks.StripeVerifier(cfg.StripeWebhookSecret),// Acknowledge, then let a job do the work: the sender times out// in seconds and a timeout becomes a duplicate delivery.Handle: func(c *gin.Context, event webhooks.Event) error {return jobs.EnqueueStripeEvent(c, event.ID, event.Raw)},})
7.Low level design
Core types
A path, a verifier and a handler. The verifier is a field rather than a convention, because every provider signs differently.
Validates a hex HMAC-SHA256 in a named header. The common case, and what most small providers do.
Parses a header of the form t=<unix>,v1=<hex>, recomputes over timestamp and body, and enforces a five minute tolerance.
Provider plus event identifier with a window. In the receive path rather than in each handler, so no handler can forget.
Design principles applied
- Verify first, parse second. nothing in an unverified payload is information. Parsing before verifying means processing attacker-controlled structure.
- Constant time, always. the comparison is not a detail. An early-exit comparison leaks the signature a byte at a time to anybody willing to measure.
- Acknowledge fast. the sender’s timeout is the real deadline. Missing it converts one event into several, which is a correctness problem and not a latency one.
- Outbound belongs in the outbox. a webhook is an announcement about a committed fact. Sending it from a request that might roll back announces things that did not happen.
Patterns
| Pattern | Where it is used |
|---|---|
| Signature verification | HMAC over the exact received bytes |
| Replay window | a timestamp inside the signed material |
| Idempotent receiver | dedup on the sender’s event id |
| Acknowledge and enqueue | the response decoupled from the work |
| Transactional outbox | outbound events committed with their cause |
8.Scalability and performance
- Inbound capacity is an acknowledgement rate, and acknowledgement is a verification plus two cheap operations, so it scales with ordinary API replicas.
- A provider retry burst is the real inbound load spike: an outage on their side ends with everything arriving at once, and the dedup check is what makes that harmless.
- The dedup table grows steadily and needs pruning. The window only has to outlast the longest retry schedule you integrate with.
- Outbound delivery scales by running more relays, with the claim step keeping them from duplicating work.
- One dead receiver is the thing that breaks outbound at scale, because its retries accumulate in the same queue as everybody else’s deliveries.
- Retaining deliveries with response codes is what makes redelivery and dispute resolution possible, and it is also the table that grows fastest.
9.Bottlenecks and improvements
What breaks first
- A dead receiver filling the queue. one endpoint that is down for an hour accumulates its share of every event, and its retries compete with deliveries to healthy receivers.
- Slow inbound handlers. work done inline means the sender times out and retries, so a slow handler creates the duplicates that then make it slower.
- Secret rotation. changing a signing secret invalidates every in-flight delivery, and a window where both are accepted has to exist or deliveries fail during the change.
- Dedup table growth. millions of rows a day, all of them useless after the retry window has passed.
- Verified but untrusted content. a valid signature proves who sent it, not that the payload is reasonable. The amounts and identifiers inside still need checking.
What to do about it
- Circuit breaker per receiver. stop delivering to an endpoint that is failing consistently, retry on a long interval, and notify the subscriber. One dead endpoint stops being everybody’s problem.
- Separate queue for inbound processing. so a burst of provider retries is a backlog on its own queue rather than a delay to password resets.
- Accept two secrets during rotation. verify against the current and the previous for an overlap period, then drop the old one.
- Prune the dedup table. a scheduled job deleting rows past the window. Cheap, and the table is otherwise unbounded.
- Validate the payload after verifying it. a signature authenticates the sender. The contents still get the same schema and sanity checks as any other input.
