Pulse System Design
Request tracing, query profiling, runtime metrics and alerting that run inside the process, with no collector, no Redis and no second deployment.
github.com/MUKE-coder/pulse/pulsev1.2.01.Problem statement
Something is slow. Answering why needs per-request latency, the queries each request ran, how long each took, where in the code they came from, what the heap and the goroutine count were doing at the time, and which of those is unusual. That is a known problem with known answers, and all of the standard ones are a second deployment: a collector, a time series database, a dashboard, and a budget line.
For a large team that is correct. For the application that is one container and a database, it is a monitoring stack that is bigger and harder to run than the thing it monitors, so it does not get installed, and the answer to "why was it slow" stays "nobody knows".
An in-process profiler changes the trade. It costs a dependency instead of an architecture. What it must not cost is the thing it is measuring: a monitor that adds latency to the request path, or that grows without bound until it is the reason for the outage, has made the problem it was bought to solve.
There is a subtler requirement too. The moment you sample to control cost, every total computed from the samples is wrong, and an error rate that is wrong in an unknown direction is worse than no error rate. So the counts and the traces have to be separated: sample what you store, count everything.
The system has to be able to:
- Trace every request with a latency, a trace identifier and a route.
- Capture every database query with its duration and the file and line that issued it.
- Detect N plus one patterns within a single request.
- Sample continuously: heap, goroutines, garbage collection pauses.
- Capture panics and errors with a stack trace, fingerprinted so duplicates collapse.
- Run health checks on a schedule and answer Kubernetes probes.
- Alert on a threshold without firing on a single spike.
- Compute service level objectives from exact counts, not from samples.
- Never block the request path on storage.
- Say when its own buffers, rather than the retention setting, are what limits history.
2.System requirements
Functional requirements
- Middleware for tracing and for error and panic capture.
- A GORM plugin capturing queries, callers, N plus one patterns and pool statistics.
- A runtime sampler on a five second interval, with goroutine leak detection.
- Memory ring buffer storage, or SQLite for survival across restarts.
- Per-minute rollups of every request, independent of the sample rate.
- A threshold alert engine with two-phase firing and a cooldown.
- Multiwindow burn-rate alerting on service level objectives.
- Secret redaction over captured bodies, errors and outbound URLs.
- Notification channels: Slack, Discord, email, and webhooks signed with HMAC.
- A Prometheus exposition endpoint, and JSON or CSV export.
- A WebSocket feed with per-client channel subscriptions.
- An embedded React dashboard of twelve pages.
Non-functional requirements
- Never in the request path. storing a metric is a ring buffer push or a queue append. The request never waits for storage, because a profiler that adds latency is measuring itself.
- Bounded by construction. fixed-size ring buffers with an atomic head index. Memory does not grow with traffic, which is what lets it run in a container with a limit.
- Exact counts despite sampling. per-minute rollups count every request. Totals, error rates, objectives and load-test comparisons are therefore independent of the sample rate and of how much the buffer held.
- Honest about its own limits. it alerts when buffer capacity rather than the retention setting is what bounds history, when writes were dropped, and after an unclean restart. A monitor quietly losing data is worse than no monitor.
- Redacts before storing. passwords, tokens, card numbers, JWTs and keys are stripped from bodies, error messages and outbound URLs whatever the field is called. A profiler is otherwise a credential store nobody meant to build.
- No CGo, no external service. pure Go SQLite and an embedded frontend. The reason it is installed is that installing it is one line.
3.Capacity estimation
Numbers for a mid-sized deployment. They are here to size the thing, not to predict your traffic: change an assumption and the sums below move with it.
Assumptions
| Parameter | Value |
|---|---|
| Requests per second | 926 average, 4,630 peak |
| Queries per request | 4 to 12 |
| Request record | ~500 bytes |
| Query record | ~300 bytes |
| Runtime sample interval | 5 seconds |
| Default retention | 24 hours |
Per-request overhead
The design requirement is that this stays invisible. The moment it is a percentage of request time, every latency number it reports includes itself.
Memory at full sampling
Which is exactly why the capacity warning exists. At this rate the buffer, not the 24 hour retention setting, is what decides how much history there is, and without being told nobody would know.
Sampling, and what it does not touch
This is the central design decision. Sampling controls storage; the rollups keep totals and error rates exact. Computing an error rate from a 10% sample would give an answer with no stated error bar, which people would then act on.
SQLite throughput
Runtime sampling
4.High level design
Middleware and a database plugin feed a storage layer that never blocks, and everything else reads from it.
Core components
- Tracing middleware. a trace identifier, a timer and a route label per request. Sampled for storage, counted always.
- Error middleware. recovers panics, captures stack traces and request context, fingerprints the error so repeats collapse into one entry with a count.
- GORM plugin. every query with its duration and the file and line that issued it, plus per-request tallies of repeated patterns for N plus one detection, plus pool statistics.
- Runtime sampler. heap, goroutines and garbage collection on a five second tick, with a leak threshold.
- Storage. a generic ring buffer in memory, or a batched SQLite writer. Both are append-and-return; neither makes a request wait.
- Aggregator. per-minute rollups of every request, which is what makes totals and objectives exact while traces are sampled.
- Alert engine. threshold rules with two-phase firing so a single spike does not page anybody, a cooldown, and signed webhook delivery.
- Capacity watch. notices when the buffer rather than the retention setting bounds history, when writes were dropped, and when the last shutdown was unclean.
Request flow
A request measured, and what it costs the request
- 1A request arrives. The tracing middleware starts a timer and attaches a trace identifier, which also goes out in the response so a complaint can be tied to a record.
- 2The handler runs. Nothing about it knows it is being measured, which is the point: instrumentation each handler participates in is instrumentation half of them forget.
- 3Every query the handler issues passes through the GORM plugin, which times it and records the file and line it came from. The caller is what turns "this query is slow" into "this line is slow".
- 4Repeated query patterns inside one request are tallied. Five or more of the same shape is an N plus one, which is the single most common cause of a slow page and the hardest to see from a log of individual queries, because each one is fast.
- 5Captured context is redacted before it is stored: passwords, tokens, card numbers, JWTs, keys, whatever the field is named. Without this a profiler becomes a store of credentials that nobody decided to build and nobody is guarding.
- 6The records are pushed into a fixed-size ring buffer, or appended to a queue that one writer commits in batches. Either way the call returns immediately; the request does not wait on storage, because a profiler that adds latency is measuring itself.
- 7Separately, and regardless of the sample rate, the request is counted into a per-minute rollup. This is the decision that makes the numbers trustworthy: traces are sampled, counts are complete, so an error rate is a fact rather than an estimate with no error bar.
- 8The dashboard, the alert engine and the objective burn rates all read from the rollups for anything that is a total, and from the buffers for anything that is an example. When the buffer rather than the retention setting is what limits history, it says so.
Data flow
- Sampling decides what is stored, never what is counted. Every total in the product comes from the rollups.
- Redaction happens before storage, not before display, so a leaked secret is never written down in the first place.
- The ring buffer is a fixed slice with an atomic head. Bounded memory is not a tuning option, it is the data structure.
- The SQLite path trades peak write throughput for surviving a restart, which is the right trade for an application that is restarted more often than it is busy.
5.Technology stack
| Component | What it is |
|---|---|
| Language | Go, no CGo |
| Host | Gin and GORM |
| Storage | in-memory ring buffers, or pure Go SQLite |
| Default retention | 24 hours |
| Defaults | sample rate 1.0, slow request 1s, slow query 200ms, N+1 at 5 |
| Runtime sampling | every 5 seconds, leak threshold 100 goroutines |
| Export | Prometheus, JSON, CSV, OpenTelemetry |
| Dashboard | 12 pages embedded in the binary |
6.API design
Mounted under the configured prefix
| Method | Endpoint | What it does |
|---|---|---|
| GET | /pulse/ui | The embedded dashboard |
| GET | /pulse/api/requests | Traces, filtered by route, status and latency |
| GET | /pulse/api/db/n1 | N plus one findings grouped by route and pattern |
| GET | /pulse/api/slos | Compliance, error budget and burn rates |
| GET | /pulse/metrics | Prometheus exposition |
| GET | /pulse/live | Kubernetes liveness |
| GET | /pulse/ready | Kubernetes readiness, composite over health checks |
How Grit mounts it, including the bug the comment is there to prevent
pulseOpts := []pulse.Option{pulse.WithAppName(cfg.AppName),pulse.WithCredentials(cfg.PulseUsername, cfg.PulsePassword),pulse.WithExcludePaths("/studio/*", "/sentinel/*", "/docs/*", "/pulse/*"),pulse.WithPrometheus(),// Request-body capture needs v1.0.1 or later. Before that the error// middleware read the body and put back only the first 4 KB, so every// request carrying a Content-Length reached the handler truncated:// uploads from mobile and curl failed while browsers, which send// chunked, did not. Pin below v1.0.1 and you want// pulse.WithRequestBodyCaptureDisabled() back.}if cfg.PulseStorage == "sqlite" && cfg.PulseStorageDSN != "" {pulseOpts = append(pulseOpts, pulse.WithSQLite(cfg.PulseStorageDSN))}p := pulse.Mount(context.Background(), r, db, pulseOpts...)
7.Low level design
Core types
O(1) append into a fixed-size slice behind an atomic head index and a short write lock. The reason memory is bounded regardless of traffic.
Times every query, records the caller, and keeps per-request tallies of repeated patterns. Tallies have an idle expiry so an abandoned request cannot leak one.
Per-minute rollups of every request. The separation between what is sampled and what is counted lives here.
Reports when a ring buffer rather than the retention setting is what bounds history. A monitor that silently holds two minutes of data while its settings claim a day is actively misleading.
*bool and *float64 so "not set" and "explicitly zero" are different. A sample rate of 0 and an unset sample rate must not mean the same thing.
Design principles applied
- Measuring must not be measurable. push and return. Every design choice in the storage layer exists so the request path never waits.
- Sample the examples, count the totals. sampling is necessary and making totals approximate is not. Rollups keep the numbers exact at a cost that does not depend on the sample rate.
- Bounded by the data structure. a ring buffer cannot grow. That is a stronger guarantee than a cleanup goroutine that is supposed to keep up.
- Report your own failures. dropped writes, truncated history and unclean restarts are surfaced. A monitoring tool that hides its own gaps is the one tool that must not.
- Redact on the way in. before storage, not before display. The alternative is a database of captured secrets with a filter in front of it.
Patterns
| Pattern | Where it is used |
|---|---|
| Ring buffer | bounded retention as a data structure |
| Batched single writer | durability without blocking the request |
| Fingerprinting | duplicate errors collapsed into one entry with a count |
| Two-phase alerting | a threshold must hold, not just be touched |
| Multiwindow burn rate | the Google SRE workbook objective alerting |
8.Scalability and performance
- Overhead per request is microseconds and does not grow with traffic, because the work is a timer and a buffer push.
- Memory is fixed by the buffer sizes. The question is never how much memory it will use, it is how much history that memory buys, which is what the capacity warning answers.
- Sampling is the lever for storage volume, and it costs nothing in accuracy because the rollups count everything.
- The SQLite backend serialises on one writer, which is the right shape for SQLite and a lower ceiling than memory. Choose it for surviving restarts, not for throughput.
- Everything is per process. Behind several replicas each has its own dashboard and its own view, which is the main structural limit of an in-process design and the point at which a real collector earns its cost.
- The OpenTelemetry exporter is the bridge out: keep the in-process dashboard for the single-instance case and ship spans to a collector when there are enough instances for that to be worth running.
9.Bottlenecks and improvements
What breaks first
- Per-process view. with four replicas there are four dashboards and no combined picture, and the one you reach through the load balancer is whichever replica answered.
- Buffer rather than retention. at high traffic a memory buffer holds minutes while the configured retention says a day. It is surfaced, and it still surprises people.
- Capture that changes behaviour. reading a request body to record it is a real risk to the request. In v1.0.0 it truncated every body with a Content-Length to 4 KB, so mobile and curl uploads failed while browsers did not.
- Dashboard exposure. it holds request bodies, errors and query text. Redaction reduces what it contains; it does not make the dashboard safe to leave open.
- Alert noise. threshold rules on a noisy metric page people for nothing, and the second time that happens the alert stops being read.
What to do about it
- Export to a collector once there are replicas. the OpenTelemetry path gives one combined view. Keep Pulse mounted for the per-process detail it is good at.
- Size buffers from the actual rate. requests per second times the history you want. Or move to SQLite, which trades write ceiling for a retention setting that means what it says.
- Pin the version when capture is on. v1.2.0 or later wraps the body and passes it through whole. Below v1.0.1, turn request body capture off.
- Authenticate the dashboard and exclude it from itself. credentials always, and exclude the monitoring routes from tracing so the tool does not fill its own buffers with views of itself.
- Prefer burn-rate alerts to thresholds. multiwindow burn rate fires on budget actually being consumed rather than on an instant being bad, which is most of the difference between an alert people act on and one they mute.
