Observability System Design
Being able to answer what happened, where the time went, and whether the thing is actually working, from outside the process.
internal/tracinginternal/healthinternal/middleware1.Problem statement
A user reports that saving an order took eleven seconds at about ten past four. There is no way to answer that from the code. The question needs evidence collected before anybody knew it would be asked, which is the whole premise of observability: the instrumentation has to already be there, because the interesting failures are the ones nobody predicted.
Three different questions get conflated into one word. What happened is logs, and a log line without a request identifier cannot be joined to the other forty lines from the same request. Where did the time go is tracing, and a trace crossing service boundaries is the one artifact that cannot be reconstructed from three sets of logs with three different request identifiers. And is it working is health, which sounds like the easy one and is where the subtlest mistake lives.
That mistake is a boolean. Every component here used to report one, and that boolean covered two opposite situations: Redis is down, and this deployment has no Redis. An operator cannot tell those apart, and neither could the dashboard, so it guessed, differently per component. Redis read one way, email another, and storage was not reported at all. A monitoring system that cannot distinguish broken from absent trains the people reading it to ignore it.
The system has to be able to:
- Give every request an identifier, in the response and in every log line it causes.
- Make that identifier the same string as the trace identifier, so one lookup serves both.
- Emit spans to a collector when one is configured, and cost nothing when it is not.
- Report component health in four states, not two.
- Let anything wired into the application register its own probe without editing the handler.
- Decide the overall status from the component states by one rule in one place.
- Record who did what, as a queryable log rather than as text.
- Expose counters for the things that fail silently: dropped realtime messages, queue depth.
2.System requirements
Functional requirements
- A request identifier middleware, generating one or honouring an inbound header.
- The identifier in the X-Request-ID response header and in structured log lines.
- OpenTelemetry spans per request, exported when an OTLP endpoint is configured and inert otherwise.
- The request identifier and the trace identifier being the same value.
- A health registry: a name and a probe function, registered from anywhere.
- Four component states: ok, degraded, off, unknown.
- An overall status derived from the components, where only degraded lowers it.
- An activity log of user actions, with the actor, the action, the target and the time.
- Hub and queue counters in the health response.
Non-functional requirements
- Four states, not a boolean. absent and broken are different facts and must be different values. Collapsing them means the dashboard has to guess, and it will guess differently in each place.
- Unknown is not degraded. not knowing is not the same as being broken. Treating a probe that could not run as a failure is how an on-call rota learns to ignore a page.
- One joining key. the request identifier and the trace identifier are the same string. Two identifiers for one request means every investigation starts with a translation step.
- Inert when unconfigured. tracing with no endpoint costs nothing and logs nothing. Instrumentation that requires a collector to exist is instrumentation that gets removed.
- Probes are cheap and bounded. health is polled by a load balancer every few seconds. A probe that runs an expensive query turns monitoring into load.
- The rule lives in one place. how component states become an overall status is written once, so a new component cannot quietly invent its own interpretation.
3.Capacity estimation
Numbers for a mid-sized deployment. They are here to size the thing, not to predict your traffic: change an assumption and the sums below move with it.
Assumptions
| Parameter | Value |
|---|---|
| Requests per second | 926 average, 4,630 peak |
| Log lines per request | 3 to 8 |
| Span bytes per request | ~1.5 KB at full sampling |
| Health poll interval | 5 seconds, from each load balancer |
| Activity log rows per day | ~200,000 |
| Log retention | 14 days |
Log volume
Log volume is the cost that surprises people, and it is why levels and sampling exist. Debug logging left on in production is this number times several.
Trace volume at full sampling
Which is why sampling exists. At 10% head sampling that is 12 GB/day, and tail sampling keeps the slow and failed requests that are the ones worth having.
Health probe load
Negligible, as long as each probe is a ping rather than a count. A probe doing SELECT COUNT(*) turns this into real database load forever.
Activity log growth
It is in the application database because it is queried alongside application data. That makes it the fastest growing table in most projects.
4.High level design
Three mechanisms with one joining key: an identifier per request, spans around the work, and a registry of probes.
Core components
- Request identifier. middleware generating one or accepting an inbound header. In the response and in every log line, so forty lines become one request.
- Tracing. a span per request, with child spans around database calls and outbound requests. Exported to an OTLP collector when one is configured, inert otherwise.
- Structured logs. key and value rather than a sentence, so a log store can filter by status, route and identifier instead of matching text.
- Health registry. a name and a function. Snapshot runs them all. A plugin or anything wired in main registers itself without the handler being edited.
- Four-state status. ok, degraded, off, unknown, with the derivation rule in one place: only degraded lowers the overall status.
- Activity log. who did what, as rows. A hash chain over them, so the record can be shown not to have been edited.
- Counters. what the realtime hub delivered, dropped and failed to publish, and what the queues hold, in the health response. These are the things that fail without raising an error.
Request flow
One request, instrumented three ways
- 1A request arrives. If it carries a request identifier from an upstream service, that one is kept, so a trace does not start over at every hop.
- 2A root span opens, and its trace identifier is the request identifier. One string joins the log store and the trace store, which removes the translation step from every investigation.
- 3The handler runs. It does not know it is being traced or logged, because instrumentation that each handler participates in is instrumentation that half of them forget.
- 4Every log line carries the identifier, the route, the status and the duration as fields rather than as prose, so the log store can filter rather than grep.
- 5Child spans wrap the database calls and the outbound requests. This is what answers where the eleven seconds went, and it answers it without anybody having predicted the question.
- 6Actions worth attributing are written to the activity log: who did what, to what, when. A hash chain over the rows means the record can be shown not to have been altered afterwards.
- 7Spans go to the collector when an OTLP endpoint is configured. With none, the whole path is inert: no buffering, no errors, no cost. Instrumentation that demands infrastructure gets deleted.
- 8Separately, the health endpoint runs every registered probe and reports four states per component. Only degraded lowers the overall status, because a component that is off was never asked for and one that is unknown has not accused anybody of anything.
Data flow
- The request identifier propagates outbound as a header, so a downstream service continues the same trace rather than starting one.
- Health probes are cheap by rule. A ping, a connection check, a cached count. A probe doing real work becomes load, forever, every five seconds.
- The activity log is in the application database because it is queried with application data. That is also why it grows faster than anything else.
- Counters are in the health response rather than only in logs, because a realtime hub dropping messages or a queue backing up produces no error anywhere.
5.Technology stack
| Component | What it is |
|---|---|
| Tracing | OpenTelemetry, exported on OTEL_EXPORTER_OTLP_ENDPOINT |
| Correlation | X-Request-ID, identical to the trace id |
| Logging | structured, levelled |
| Health | a probe registry and four states |
| Profiling | Pulse, per-request |
| Audit | an activity log with a SHA-256 hash chain |
| Unconfigured | tracing inert, health still answers |
6.API design
Operational endpoints
| Method | Endpoint | What it does |
|---|---|---|
| GET | /api/health | Component states, overall status, hub and queue counters |
| GET | /api/v1/admin/activity | The activity log, filtered and paginated |
| GET | /api/v1/admin/activity/integrity | Verify the hash chain |
A health response that distinguishes absent from broken
{"status": "ok","components": {"database": { "state": "ok", "detail": "postgres 16.2" },"cache": { "state": "off", "detail": "REDIS_URL not set" },"storage": { "state": "ok", "detail": "local disk at ./storage" },"mail": { "state": "off", "detail": "MAIL_MAILER=log" },"queue": { "state": "unknown", "detail": "no poll taken yet" }},"realtime": { "sockets": 412, "delivered": 918244, "dropped": 3 }}
7.Low level design
Core types
Takes a name and a probe function. The registry is what lets a plugin report itself without the handler knowing it exists.
Runs every probe and returns the component states. The handler merges them into the response and adds nothing of its own.
Derives the overall status. Only degraded lowers it; off and unknown do not. One function, so a new component cannot invent its own rule.
The per-request span, which also makes the request identifier and the trace identifier the same string.
Sets up the exporter when an endpoint is configured and returns a no-op provider otherwise, so nothing downstream branches on whether tracing is on.
Design principles applied
- Four states beat a boolean. the two-state version made the dashboard guess, and it guessed differently per component. Putting the distinction in the response left the page nothing to infer.
- Not knowing is not failing. unknown exists so that a probe which could not run does not page anybody. Alert fatigue is a monitoring failure, not an operator failure.
- One key joins everything. the request identifier is the trace identifier. Two identifiers for one request doubles the work of every investigation.
- A registry, not a switch statement. components register themselves. The alternative is a handler that has to be edited for every new component, which is where storage got left out.
- Inert by default. tracing with no collector costs nothing. The moment instrumentation requires infrastructure, a project without that infrastructure removes the instrumentation.
Patterns
| Pattern | Where it is used |
|---|---|
| Correlation identifier | one key joining logs, traces and responses |
| Registry | self-registering probes instead of a central list |
| Four-state health | absent, broken, unknown and fine as distinct facts |
| No-op provider | instrumentation that is free when unconfigured |
| Hash chain | an audit log that can be shown to be unedited |
8.Scalability and performance
- Logging is the dominant cost and it is linear in traffic. Levels and sampling are the only controls that matter, and debug left on in production is the single most common way to pay for it.
- Traces need sampling above modest traffic. Head sampling is cheap and loses the interesting requests; tail sampling keeps the slow and failed ones and costs a collector that buffers.
- Health probes are polled forever, so a probe that does real work is a permanent load increase. Caching a probe result for a few seconds is almost always correct.
- The activity log grows faster than any other table and eventually needs partitioning or archival to cold storage.
- The hash chain means verification walks the rows, which gets slower as the log grows. Verifying the recent tail is the practical answer, with full verification as an offline job.
- With nothing configured the whole system still works: health answers, logs go to stdout, tracing is inert. A single-binary deployment is not asked to run a collector.
9.Bottlenecks and improvements
What breaks first
- Log volume and cost. at scale the log bill can exceed the compute bill, and the usual cause is a debug level or a line per database query.
- Expensive health probes. a probe that counts rows runs several times a second forever, and it is load that was added to detect load.
- Cardinality explosion. a metric labelled with a user identifier or a path containing one creates a time series per value, and metric stores fall over on cardinality rather than volume.
- Sampling losing the evidence. head sampling at 10% means the eleven second request is 90% likely not to have been recorded, which is exactly the request that was worth recording.
- Activity log growth. the fastest growing table, in the same database as the application data it is competing with.
What to do about it
- Sample logs, not just traces. log every error and a fraction of successes. The successes are almost all identical, and the errors are all different.
- Cache probe results. a few seconds of staleness in a health response is not a problem. Permanent query load to produce it is.
- Template the labels. label by route pattern rather than by path. The pattern is a fixed set; the path is unbounded.
- Tail sample on duration and status. buffer the spans, keep the slow and the failed. It costs a collector and it keeps precisely the traces anybody will ever look at.
- Partition and archive the activity log. by month, with old partitions moved to object storage. The hash chain is per partition so verification stays bounded.
