All systems
Embedded

Pulse System Design

Request tracing, query profiling, runtime metrics and alerting that run inside the process, with no collector, no Redis and no second deployment.

github.com/MUKE-coder/pulse/pulsev1.2.0

1.Problem statement

Something is slow. Answering why needs per-request latency, the queries each request ran, how long each took, where in the code they came from, what the heap and the goroutine count were doing at the time, and which of those is unusual. That is a known problem with known answers, and all of the standard ones are a second deployment: a collector, a time series database, a dashboard, and a budget line.

For a large team that is correct. For the application that is one container and a database, it is a monitoring stack that is bigger and harder to run than the thing it monitors, so it does not get installed, and the answer to "why was it slow" stays "nobody knows".

An in-process profiler changes the trade. It costs a dependency instead of an architecture. What it must not cost is the thing it is measuring: a monitor that adds latency to the request path, or that grows without bound until it is the reason for the outage, has made the problem it was bought to solve.

There is a subtler requirement too. The moment you sample to control cost, every total computed from the samples is wrong, and an error rate that is wrong in an unknown direction is worse than no error rate. So the counts and the traces have to be separated: sample what you store, count everything.

The system has to be able to:

  • Trace every request with a latency, a trace identifier and a route.
  • Capture every database query with its duration and the file and line that issued it.
  • Detect N plus one patterns within a single request.
  • Sample continuously: heap, goroutines, garbage collection pauses.
  • Capture panics and errors with a stack trace, fingerprinted so duplicates collapse.
  • Run health checks on a schedule and answer Kubernetes probes.
  • Alert on a threshold without firing on a single spike.
  • Compute service level objectives from exact counts, not from samples.
  • Never block the request path on storage.
  • Say when its own buffers, rather than the retention setting, are what limits history.

2.System requirements

Functional requirements

  • Middleware for tracing and for error and panic capture.
  • A GORM plugin capturing queries, callers, N plus one patterns and pool statistics.
  • A runtime sampler on a five second interval, with goroutine leak detection.
  • Memory ring buffer storage, or SQLite for survival across restarts.
  • Per-minute rollups of every request, independent of the sample rate.
  • A threshold alert engine with two-phase firing and a cooldown.
  • Multiwindow burn-rate alerting on service level objectives.
  • Secret redaction over captured bodies, errors and outbound URLs.
  • Notification channels: Slack, Discord, email, and webhooks signed with HMAC.
  • A Prometheus exposition endpoint, and JSON or CSV export.
  • A WebSocket feed with per-client channel subscriptions.
  • An embedded React dashboard of twelve pages.

Non-functional requirements

  • Never in the request path. storing a metric is a ring buffer push or a queue append. The request never waits for storage, because a profiler that adds latency is measuring itself.
  • Bounded by construction. fixed-size ring buffers with an atomic head index. Memory does not grow with traffic, which is what lets it run in a container with a limit.
  • Exact counts despite sampling. per-minute rollups count every request. Totals, error rates, objectives and load-test comparisons are therefore independent of the sample rate and of how much the buffer held.
  • Honest about its own limits. it alerts when buffer capacity rather than the retention setting is what bounds history, when writes were dropped, and after an unclean restart. A monitor quietly losing data is worse than no monitor.
  • Redacts before storing. passwords, tokens, card numbers, JWTs and keys are stripped from bodies, error messages and outbound URLs whatever the field is called. A profiler is otherwise a credential store nobody meant to build.
  • No CGo, no external service. pure Go SQLite and an embedded frontend. The reason it is installed is that installing it is one line.

3.Capacity estimation

Numbers for a mid-sized deployment. They are here to size the thing, not to predict your traffic: change an assumption and the sums below move with it.

Assumptions

ParameterValue
Requests per second926 average, 4,630 peak
Queries per request4 to 12
Request record~500 bytes
Query record~300 bytes
Runtime sample interval5 seconds
Default retention24 hours

Per-request overhead

a trace identifier, a timer, and a ring buffer push
the push is an atomic index plus a short write lock
microseconds against a request measured in milliseconds

The design requirement is that this stays invisible. The moment it is a percentage of request time, every latency number it reports includes itself.

Memory at full sampling

926 req/s x 500 bytes = ~460 KB/s of request records
plus 926 x 8 queries x 300 bytes = ~2.2 MB/s of query records
a ring buffer sized for 100,000 requests holds ~50 MB and about 2 minutes

Which is exactly why the capacity warning exists. At this rate the buffer, not the 24 hour retention setting, is what decides how much history there is, and without being told nobody would know.

Sampling, and what it does not touch

at SampleRate 0.1: 93 traces/s stored instead of 926
memory and query volume fall by ten
the per-minute rollups still count all 926

This is the central design decision. Sampling controls storage; the rollups keep totals and error rates exact. Computing an error rate from a 10% sample would give an answer with no stated error bar, which people would then act on.

SQLite throughput

appends go to a queue, committed in batches by one writer
one writer, so writes serialise, which SQLite needs anyway
survives restarts, at a lower ceiling than the ring buffer

Runtime sampling

one sample every 5 seconds = 17,280/day
at ~200 bytes = ~3.5 MB/day

4.High level design

Middleware and a database plugin feed a storage layer that never blocks, and everything else reads from it.

Core components

  • Tracing middleware. a trace identifier, a timer and a route label per request. Sampled for storage, counted always.
  • Error middleware. recovers panics, captures stack traces and request context, fingerprints the error so repeats collapse into one entry with a count.
  • GORM plugin. every query with its duration and the file and line that issued it, plus per-request tallies of repeated patterns for N plus one detection, plus pool statistics.
  • Runtime sampler. heap, goroutines and garbage collection on a five second tick, with a leak threshold.
  • Storage. a generic ring buffer in memory, or a batched SQLite writer. Both are append-and-return; neither makes a request wait.
  • Aggregator. per-minute rollups of every request, which is what makes totals and objectives exact while traces are sampled.
  • Alert engine. threshold rules with two-phase firing so a single spike does not page anybody, a cooldown, and signed webhook delivery.
  • Capacity watch. notices when the buffer rather than the retention setting bounds history, when writes were dropped, and when the last shutdown was unclean.

Request flow

A request measured, and what it costs the request

12345678RequestTracing MWHandlerGORM PluginN+1 TallyRedactRing BufferPer-Minute RollupDashboard + Alerts
  1. 1A request arrives. The tracing middleware starts a timer and attaches a trace identifier, which also goes out in the response so a complaint can be tied to a record.
  2. 2The handler runs. Nothing about it knows it is being measured, which is the point: instrumentation each handler participates in is instrumentation half of them forget.
  3. 3Every query the handler issues passes through the GORM plugin, which times it and records the file and line it came from. The caller is what turns "this query is slow" into "this line is slow".
  4. 4Repeated query patterns inside one request are tallied. Five or more of the same shape is an N plus one, which is the single most common cause of a slow page and the hardest to see from a log of individual queries, because each one is fast.
  5. 5Captured context is redacted before it is stored: passwords, tokens, card numbers, JWTs, keys, whatever the field is named. Without this a profiler becomes a store of credentials that nobody decided to build and nobody is guarding.
  6. 6The records are pushed into a fixed-size ring buffer, or appended to a queue that one writer commits in batches. Either way the call returns immediately; the request does not wait on storage, because a profiler that adds latency is measuring itself.
  7. 7Separately, and regardless of the sample rate, the request is counted into a per-minute rollup. This is the decision that makes the numbers trustworthy: traces are sampled, counts are complete, so an error rate is a fact rather than an estimate with no error bar.
  8. 8The dashboard, the alert engine and the objective burn rates all read from the rollups for anything that is a total, and from the buffers for anything that is an example. When the buffer rather than the retention setting is what limits history, it says so.

Data flow

  • Sampling decides what is stored, never what is counted. Every total in the product comes from the rollups.
  • Redaction happens before storage, not before display, so a leaked secret is never written down in the first place.
  • The ring buffer is a fixed slice with an atomic head. Bounded memory is not a tuning option, it is the data structure.
  • The SQLite path trades peak write throughput for surviving a restart, which is the right trade for an application that is restarted more often than it is busy.

5.Technology stack

ComponentWhat it is
LanguageGo, no CGo
HostGin and GORM
Storagein-memory ring buffers, or pure Go SQLite
Default retention24 hours
Defaultssample rate 1.0, slow request 1s, slow query 200ms, N+1 at 5
Runtime samplingevery 5 seconds, leak threshold 100 goroutines
ExportPrometheus, JSON, CSV, OpenTelemetry
Dashboard12 pages embedded in the binary

6.API design

Mounted under the configured prefix

MethodEndpointWhat it does
GET/pulse/uiThe embedded dashboard
GET/pulse/api/requestsTraces, filtered by route, status and latency
GET/pulse/api/db/n1N plus one findings grouped by route and pattern
GET/pulse/api/slosCompliance, error budget and burn rates
GET/pulse/metricsPrometheus exposition
GET/pulse/liveKubernetes liveness
GET/pulse/readyKubernetes readiness, composite over health checks

How Grit mounts it, including the bug the comment is there to prevent

pulseOpts := []pulse.Option{
pulse.WithAppName(cfg.AppName),
pulse.WithCredentials(cfg.PulseUsername, cfg.PulsePassword),
pulse.WithExcludePaths("/studio/*", "/sentinel/*", "/docs/*", "/pulse/*"),
pulse.WithPrometheus(),
// Request-body capture needs v1.0.1 or later. Before that the error
// middleware read the body and put back only the first 4 KB, so every
// request carrying a Content-Length reached the handler truncated:
// uploads from mobile and curl failed while browsers, which send
// chunked, did not. Pin below v1.0.1 and you want
// pulse.WithRequestBodyCaptureDisabled() back.
}
if cfg.PulseStorage == "sqlite" && cfg.PulseStorageDSN != "" {
pulseOpts = append(pulseOpts, pulse.WithSQLite(cfg.PulseStorageDSN))
}
p := pulse.Mount(context.Background(), r, db, pulseOpts...)

7.Low level design

Core types

RingBuffer[T]

O(1) append into a fixed-size slice behind an atomic head index and a short write lock. The reason memory is bounded regardless of traffic.

GORM pluginpulse/gorm_plugin.go

Times every query, records the caller, and keeps per-request tallies of repeated patterns. Tallies have an idle expiry so an abandoned request cannot leak one.

Aggregatorpulse/aggregator.go

Per-minute rollups of every request. The separation between what is sampled and what is counted lives here.

Capacity watchpulse/capacity.go

Reports when a ring buffer rather than the retention setting is what bounds history. A monitor that silently holds two minutes of data while its settings claim a day is actively misleading.

Pointer config fields

*bool and *float64 so "not set" and "explicitly zero" are different. A sample rate of 0 and an unset sample rate must not mean the same thing.

Design principles applied

  • Measuring must not be measurable. push and return. Every design choice in the storage layer exists so the request path never waits.
  • Sample the examples, count the totals. sampling is necessary and making totals approximate is not. Rollups keep the numbers exact at a cost that does not depend on the sample rate.
  • Bounded by the data structure. a ring buffer cannot grow. That is a stronger guarantee than a cleanup goroutine that is supposed to keep up.
  • Report your own failures. dropped writes, truncated history and unclean restarts are surfaced. A monitoring tool that hides its own gaps is the one tool that must not.
  • Redact on the way in. before storage, not before display. The alternative is a database of captured secrets with a filter in front of it.

Patterns

PatternWhere it is used
Ring bufferbounded retention as a data structure
Batched single writerdurability without blocking the request
Fingerprintingduplicate errors collapsed into one entry with a count
Two-phase alertinga threshold must hold, not just be touched
Multiwindow burn ratethe Google SRE workbook objective alerting

8.Scalability and performance

  • Overhead per request is microseconds and does not grow with traffic, because the work is a timer and a buffer push.
  • Memory is fixed by the buffer sizes. The question is never how much memory it will use, it is how much history that memory buys, which is what the capacity warning answers.
  • Sampling is the lever for storage volume, and it costs nothing in accuracy because the rollups count everything.
  • The SQLite backend serialises on one writer, which is the right shape for SQLite and a lower ceiling than memory. Choose it for surviving restarts, not for throughput.
  • Everything is per process. Behind several replicas each has its own dashboard and its own view, which is the main structural limit of an in-process design and the point at which a real collector earns its cost.
  • The OpenTelemetry exporter is the bridge out: keep the in-process dashboard for the single-instance case and ship spans to a collector when there are enough instances for that to be worth running.

9.Bottlenecks and improvements

What breaks first

  • Per-process view. with four replicas there are four dashboards and no combined picture, and the one you reach through the load balancer is whichever replica answered.
  • Buffer rather than retention. at high traffic a memory buffer holds minutes while the configured retention says a day. It is surfaced, and it still surprises people.
  • Capture that changes behaviour. reading a request body to record it is a real risk to the request. In v1.0.0 it truncated every body with a Content-Length to 4 KB, so mobile and curl uploads failed while browsers did not.
  • Dashboard exposure. it holds request bodies, errors and query text. Redaction reduces what it contains; it does not make the dashboard safe to leave open.
  • Alert noise. threshold rules on a noisy metric page people for nothing, and the second time that happens the alert stops being read.

What to do about it

  • Export to a collector once there are replicas. the OpenTelemetry path gives one combined view. Keep Pulse mounted for the per-process detail it is good at.
  • Size buffers from the actual rate. requests per second times the history you want. Or move to SQLite, which trades write ceiling for a retention setting that means what it says.
  • Pin the version when capture is on. v1.2.0 or later wraps the body and passes it through whole. Below v1.0.1, turn request body capture off.
  • Authenticate the dashboard and exclude it from itself. credentials always, and exclude the monitoring routes from tracing so the tool does not fill its own buffers with views of itself.
  • Prefer burn-rate alerts to thresholds. multiwindow burn rate fires on budget actually being consumed rather than on an instant being bad, which is most of the difference between an alert people act on and one they mute.

Read next