Realtime Updates System Design
Pushing a change to the browsers that care about it, across replicas, with a fallback for the proxies that eat WebSockets.
internal/realtimeinternal/handlers1.Problem statement
An operator has a list of orders open. Another operator marks one as shipped. The first screen should change, and with plain request and response it does not: it changes when somebody reloads, or after whatever polling interval was guessed at.
Polling is the easy answer and it scales badly in a specific way. Every open tab asks every few seconds, so cost is proportional to viewers times frequency and almost every request returns nothing changed. Twenty tabs on a five second poll is 240 requests a minute to tell nobody anything.
A socket inverts that: the server speaks when there is something to say. Which introduces three problems that are invisible on one developer machine. A hub that lives in a process only knows the clients connected to that process, so with two replicas a user connected to A never hears an event published on B, and nothing errors because the push succeeded into a registry that does not contain them. A socket that authenticates with a cookie and accepts any origin lets a page on another domain open a socket as the signed-in user. And a platform or proxy that strips the upgrade header makes the handshake fail forever, with a client reconnect loop that never succeeds and never complains.
All three of those shipped here before they were found. The design is what remains after fixing them.
The system has to be able to:
- Hold a socket per browser tab, authenticated as the signed-in user.
- Subscribe to a named channel, authorised per channel rather than per user.
- Deliver an event to every subscriber, whichever replica they are connected to.
- Fall back to plain HTTP streaming where WebSockets do not survive the network.
- Report who is present in a channel, across replicas.
- Let clients send each other low-value events directly, without a round trip through the API.
- Refuse a socket from an origin that is not ours, and cap how many one account can hold.
2.System requirements
Functional requirements
- A WebSocket endpoint authenticated by the access cookie.
- A Server-Sent Events endpoint with the same delivery semantics, for networks that break upgrades.
- Subscribe and unsubscribe messages on the socket.
- A channel authoriser registry, matched by pattern, so a channel name maps to a permission check.
- A Redis backplane, so a publish on one replica reaches subscribers on all of them.
- Presence channels, listing who is subscribed, aggregated across replicas.
- Client events on private and presence channels, for typing and cursors.
- An origin check and a per-account socket cap.
- Hub counters in the health endpoint: delivered, dropped, publish failures.
Non-functional requirements
- Correct with more than one replica. this is the requirement that makes the backplane mandatory rather than an optimisation. An in-process hub is not a smaller version of the right thing, it is silently wrong the first time two replicas overlap during a rolling deploy.
- Authorised per channel. a socket proves who you are. A channel decides what you may hear. Conflating them means every subscriber of a record gets every event for that user.
- Degrades, never fails silently. where the upgrade is stripped, SSE carries the same events. A client retry loop that can never succeed is the worst outcome, because nobody is told.
- Bounded. a cap per account, because nothing otherwise stops one client opening sockets until the replica runs out of file descriptors.
- Best effort delivery. a slow client is dropped rather than buffered without limit, and the drop is counted. Realtime is a view accelerator, not the source of truth, and a reload is always the recovery.
3.Capacity estimation
Numbers for a mid-sized deployment. They are here to size the thing, not to predict your traffic: change an assumption and the sums below move with it.
Assumptions
| Parameter | Value |
|---|---|
| Concurrent sockets | 5,000 at peak |
| Events published per second | ~120 |
| Average subscribers per event | 8 |
| Event size | ~600 bytes |
| API replicas | 4 |
| Memory per idle socket | ~10 KB of buffers and bookkeeping |
Fan-out
Socket memory
Sockets are cheap until a slow client makes the write buffer grow. The cap and the drop policy are what keep this number honest.
Backplane traffic
The backplane broadcasts, so received traffic grows with replica count even when the subscriber is on one replica. That is the limit this design hits first at very large scale.
Against polling
Roughly the same number, except one of them is useful. The polling version also pays for 1,000 authentications and database reads a second.
4.High level design
A hub per replica holds the sockets. A backplane makes the hubs behave as one. Channels decide who hears what.
Core components
- Hub. the registry of connected clients in this replica, and the send path. Its connection field is allowed to be nil, which is what let SSE be added without changing it.
- Backplane. Redis pub/sub between replicas. Without it the hub is an in-process registry, and realtime stops working the moment there are two replicas, invisibly.
- Channels. named streams with subscribe and unsubscribe on the socket, and a registry of authorisers matched by pattern. This is how a page about one invoice hears about that invoice and nothing else.
- Presence. who is subscribed to a presence channel, aggregated across replicas rather than reported per process.
- Whispers. client events on private and presence channels. Typing indicators and cursors, which are not worth a REST round trip and are not worth persisting.
- Guard. the origin check and the per-account socket cap. Added after the handshake was found to accept every origin while authenticating from a cookie.
- SSE fallback. plain HTTP with a response that never ends. Survives proxies, gateways and platforms that strip the upgrade header before it reaches the application.
Request flow
An event reaching a subscriber on a different replica
- 1A write happens: an order is marked shipped by somebody on replica A.
- 2The service publishes an event naming a channel, not a user. The channel is the record, so whoever is watching that record hears about it and nobody else does.
- 3The publish goes to the backplane rather than only to the local hub. This is the step whose absence is invisible on one instance and wrong on two.
- 4Every replica receives it, including the one holding the subscriber. The writer has no idea where the subscriber is connected, and does not need to.
- 5Replica A delivers to its own subscribers too. It receives its own publish through the backplane rather than short-circuiting, so there is one delivery path and not two.
- 6A subscription was authorised when it was made, by an authoriser matched to the channel pattern. The check happens at subscribe, not at every event, so a permission change needs the subscription to be reconsidered rather than being caught per message.
- 7Delivery goes through one send function whether the client holds a WebSocket or an SSE stream. A slow client that cannot keep up is dropped and counted, rather than buffered until the replica suffers for it.
- 8The browser applies the change. In practice that means invalidating a query, so the data is refetched through the normal authorised path rather than trusted from the socket payload.
Data flow
- An event is a notification, not a record. The payload says what changed, and the client refetches through the ordinary API, which keeps authorisation in one place.
- Authorisation happens at subscribe. The socket says who you are, the authoriser says whether you may hear this channel.
- The origin check matters specifically because the handshake authenticates from a cookie. Accepting any origin means any page can open an authenticated socket.
- Counters for delivered, dropped and publish failures are in the health endpoint, because a realtime system that stops working does so without raising an error anywhere.
5.Technology stack
| Component | What it is |
|---|---|
| Transport | WebSocket, with Server-Sent Events as the fallback |
| Cross-replica | Redis pub/sub backplane |
| Addressing | named channels with per-pattern authorisers |
| Extras | presence, client events |
| Guards | origin allowlist, per-account socket cap |
| Observability | delivered, dropped and publish-failure counters in /api/health |
| Client | subscribe, then invalidate a React Query key on an event |
6.API design
Connection endpoints
| Method | Endpoint | What it does |
|---|---|---|
| GET | /api/ws | WebSocket upgrade, authenticated by cookie, origin checked |
| GET | /api/realtime/sse | Same events over plain HTTP streaming |
| GET | /api/health | Includes hub counters: sockets, delivered, dropped |
Subscribing, and what an event is for
const socket = useRealtime()useEffect(() => {socket.subscribe(`private-order.${orderId}`)const off = socket.on('order.updated', () => {// Refetch through the normal authorised path rather than// trusting the payload. The event says "look again", not "here it is".queryClient.invalidateQueries({ queryKey: ['order', orderId] })})return () => { off(); socket.unsubscribe(`private-order.${orderId}`) }}, [orderId])
7.Low level design
Core types
The per-replica registry and the send path. Client.Conn may be nil, which is what allowed SSE to reuse everything.
RegisterUnregisterSendPublishRedis pub/sub between replicas. Every publish goes through it, including to local subscribers, so there is one delivery path.
A pattern registry. A channel name is matched to an authoriser, which decides whether this user may subscribe.
The origin allowlist and the per-account cap. Both added after review: the handshake accepted every origin, and nothing bounded sockets per account.
The fallback. Plain HTTP, a response that never ends, the same hub.
Design principles applied
- Assume two replicas from the start. an in-process hub is correct on one instance and silently wrong on two, and the second instance arrives during a rolling deploy without anybody deciding.
- Identity on the socket, permission on the channel. two different questions. Answering only the first means every subscriber gets everything addressed to their user.
- One send path. nil was already allowed for the connection, so SSE needed no hub changes. A second delivery path would have been a second place for the drop policy to differ.
- Events say look again. refetching through the API keeps authorisation in the API. A payload trusted from a socket is a second authorisation surface.
Patterns
| Pattern | Where it is used |
|---|---|
| Pub/sub | named channels rather than direct addressing |
| Backplane | process-local hubs made to behave as one |
| Graceful degradation | SSE where upgrades do not survive |
| Presence set | subscriber membership aggregated across replicas |
| Bounded buffer | a slow client dropped and counted, not buffered |
8.Scalability and performance
- Sockets scale with replicas, and they are cheap: a few thousand per replica is memory measured in tens of megabytes.
- The backplane is the first real limit. Every replica receives every publish, so received traffic grows with replica count regardless of where the subscribers are. Sharding channels across Redis channels is the next step after that hurts.
- Channel addressing is what keeps fan-out small. Publishing to a user means every tab they have open; publishing to a record means the tabs looking at it.
- A sticky load balancer is not needed, because the backplane makes any replica able to serve any subscriber. That is worth more than it sounds: it means a rolling deploy does not need session affinity.
- SSE costs one held HTTP connection per client, which is the same order as a socket but interacts worse with some HTTP/1.1 connection limits. It is the fallback, not the default, for that reason.
- Dropping slow clients bounds memory by design. The alternative is a buffer that grows until one bad network connection degrades a whole replica.
9.Bottlenecks and improvements
What breaks first
- Backplane broadcast. every replica receives every event. At forty replicas that is forty times the publish traffic to deliver the same messages.
- Channels that are too broad. a channel per user rather than per record means a page about one invoice receives everything that happens to that user.
- Permission changes mid-subscription. authorisation happens at subscribe, so revoking access does not close an existing subscription until something re-evaluates it.
- Reconnect storms. a deploy disconnects every socket at once, and they all reconnect at once, each one authenticating and re-subscribing.
- Events treated as the source of truth. a client that applies the payload directly rather than refetching will diverge the moment one event is dropped, and drops are allowed by design.
What to do about it
- Shard the backplane by channel. a replica subscribes only to the Redis channels it has subscribers for, so received traffic follows interest rather than replica count.
- Address records, not users. narrower channels cut fan-out at the source, which is cheaper than any optimisation applied afterwards.
- Re-check subscriptions on a permission change. publish a revocation event and have the hub drop affected subscriptions, rather than waiting for the client to reconnect.
- Jitter the client reconnect. exponential backoff with randomness, so a deploy is a spread of reconnections rather than a spike.
- Treat the event as an invalidation. refetch, do not apply. It is correct under dropped events and it keeps authorisation in one place.
