Sentinel System Design
A web application firewall, threat intelligence and audit layer that mounts in one call, and the accuracy numbers to say how often it is wrong.
github.com/MUKE-coder/sentinel/v2v2.6.01.Problem statement
An application on the public internet is scanned within hours of getting a DNS record. Most of it is automated and most of it is pointless, and the small remainder is the reason the rest cannot be ignored: SQL injection against a parameter somebody forgot to bind, a path traversal in a file endpoint, credential stuffing against the login, an open redirect in a callback.
The platform answers are a managed firewall in front of the application, which costs money and sees only what the edge sees, or a library, which sees the parsed request and the authenticated user and runs for the price of a dependency.
A firewall inside the application has one dominant failure mode, and it is not missing an attack. It is blocking legitimate traffic. A pattern that matches real requests takes the site down for the people it was protecting, and it does it quietly, as a fraction of users who cannot complete something. Sentinel shipped exactly that twice: an SSRF pattern matched `0.0.0.0` inside browser version strings like `Chrome/140.0.0.0`, which blocked every Chrome user, and a bare `--` matched inside base64url cookies, which rejected about one session in ten at random.
That is why the interesting number here is not detection but false positives, and why the accuracy is measured against a corpus and pinned in CI rather than asserted.
The system has to be able to:
- Inspect the path, query, body and headers for known attack shapes.
- Run in log mode or block mode, with configurable strictness per rule family.
- Decode layered encoding before matching, so one more percent sign is not a bypass.
- Limit requests per address, per user, per route and globally, shared across replicas.
- Lock out brute force against a login, with an optional CAPTCHA tier.
- Profile threat actors and check addresses against reputation feeds.
- Keep an audit log that can be shown not to have been edited.
- Report accuracy against a fixed corpus, so a rule change has a measurable cost.
- Refuse configuration that compiles and silently does nothing.
- Do all of it without the request waiting on any of the analysis.
2.System requirements
Functional requirements
- Mount on a Gin router, with MountE returning an error rather than killing the host.
- Detection for SQL injection, XSS, path traversal, command injection, SSRF and XXE.
- Four sensitivity levels per rule family, and custom rules with a block or log action.
- Recursive decoding of layered encoding across path, parameters and form bodies.
- Route exclusions with wildcard and globstar patterns.
- A body inspection cap, with oversized bodies rejected rather than partly scanned.
- Rate limiting with fixed window, sliding window and token bucket strategies.
- A shared counter store interface, with a Redis implementation and a conformance suite.
- Auth shield: failed login budget, lockout, optional CAPTCHA.
- Trusted proxy handling that walks the forwarded chain from the right.
- A GORM plugin recording database changes, panic-guarded so auditing cannot break a write.
- A hash-chained audit log with optional HMAC and a verify endpoint.
- An asynchronous event pipeline with a ring buffer and visible drop counts.
- Live config: dashboard changes that survive a restart and reach every replica.
- Compliance reports carrying provenance and a warning when the data cannot support them.
- An embedded dashboard of fifteen pages, plus alerting to Slack, email, webhook and PagerDuty.
Non-functional requirements
- False positives are the headline number. with the default rules, 10% false positives and 100% detection against a corpus of 89 legitimate-but-suspicious requests and 56 attacks, pinned in CI. It was 38% and 92%. A firewall whose accuracy is not measured is a firewall whose accuracy is unknown.
- Analysis never blocks the request. events go to a ring buffer and are processed by background workers. The decision to block is synchronous and cheap; everything after it is not.
- Drops are visible. the pipeline counts what it could not buffer, and the count is on the dashboard. An attack large enough to overwhelm analysis must not also hide itself.
- Dead config is refused. config that compiles, reads as correct and does nothing was the common thread behind four separate issues. Mount validates and logs, and the validator can be called to fail a deploy.
- Shared state across replicas. counters and lockouts in Redis, because counted per process behind N instances a client gets N times every limit and N times the failed-login budget.
- Fails open on infrastructure. if Redis is down, requests are allowed. A rate limiter that takes the site down when its counter store is unavailable has become the outage.
- No published secret. an unset dashboard key is random per process, the default password works from localhost only, and WebSocket handshakes must be same-origin.
3.Capacity estimation
Numbers for a mid-sized deployment. They are here to size the thing, not to predict your traffic: change an assumption and the sums below move with it.
Assumptions
| Parameter | Value |
|---|---|
| Requests per second | 926 average, 4,630 peak |
| Attack traffic | 1% to 20%, bursty |
| Body inspection cap | 1 MB as Grit configures it |
| Pipeline ring buffer | 10,000 events |
| Audit retention | 365 days |
Inspection cost
The cap is the whole story on worst case. Without one, an attacker chooses how much CPU each request costs by choosing how large a body to send.
What a 10% false positive rate means
Which is why block mode belongs behind route exclusions for the endpoints that carry rich text, and why Grit runs log mode in development. The number is small and it is not zero, and the people it hits cannot tell you why.
Pipeline headroom
An attack is exactly when the event rate spikes and exactly when losing the record is worst, so the drop count is on the dashboard rather than in a log.
Rate limit cost
Audit volume
4.High level design
A synchronous decision on the request path, and everything expensive behind a buffer.
Core components
- Middleware chain. client address resolution, rate limit, WAF, auth shield. The only parts that can block, and all of them cheap.
- Decoder. unwraps layered encoding before matching. Values were decoded once, so a double-encoded payload passed straight through.
- Detector. the pattern set, scoped per rule to where it may match. Patterns that scan headers as well as parameters are how browser version strings came to look like SSRF.
- Scorer and classifier. severity, a CVSS vector per threat type, and sensitivity levels that change what counts as a match.
- Counter store. the interface behind rate limiting and the failed-login budget. In process by default, Redis for shared, with a conformance suite for other implementations.
- Event pipeline. a buffered channel and background workers. Everything that is not the block decision happens here, with drops counted.
- Intelligence. threat actor profiling, reputation lookups and anomaly detection, all fed by the pipeline rather than by the request.
- Live config. settings changed from the dashboard, stored and propagated to every replica within the sync interval.
- Audit chain. hash-chained entries with optional HMAC, and an endpoint that verifies them.
Request flow
A request inspected, blocked, and analysed afterwards
- 1A request arrives. Before anything else the client address is resolved by walking the forwarded chain from the right past the configured trusted proxies. Taking the leftmost entry lets any client write its own address and walk past rate limits, address blocks and the lockout, which was the v2.2.2 fix.
- 2The rate limiter counts the request against its keys: address, user, route, global. The strategy is a real sliding window by default, approximated from two fixed windows, with token bucket and fixed window also available.
- 3Counting happens in the shared store when one is configured. Per process, behind four replicas, a client got four times every limit. If the store is unreachable the request is allowed, because a limiter that fails closed is an outage with a security rationale.
- 4The firewall decodes before it matches. Leftover encoding in the path, parameters and form bodies is unwrapped, because a single decode pass meant a double-encoded payload went through unexamined.
- 5The pattern set runs, each rule scoped to where it is allowed to match. That scoping is the lesson from two production incidents: an unscoped host pattern matched browser version strings, and a bare statement terminator matched inside session cookies.
- 6A match blocks or logs, depending on mode. Grit logs in development and blocks in production, and excludes the authenticated rich-text admin routes, whose bodies are legitimately full of markup that the XSS rules would otherwise flag on every save.
- 7The event goes to a ring buffer and the request continues. Nothing after this point is on the request path, which is what lets the analysis be expensive.
- 8Background workers score it, profile the actor, check reputation feeds, run anomaly detection, write the audit entry into the hash chain, and push it to the dashboard. If the buffer was full the event is dropped and the drop is counted, because an attack large enough to overwhelm the analysis must not also erase the evidence of itself.
Data flow
- The block decision is synchronous and everything else is not. That split is what makes a firewall affordable in the request path.
- Every pattern declares where it may match. A pattern with no declared scope will eventually match something in a header that nobody considered.
- Counters and lockouts are shared state when a store is configured, and the semantics are pinned by a conformance suite so another implementation behaves the same way.
- Data sent to an AI provider is redacted by default: query values, bodies, full addresses and personal data are stripped first.
5.Technology stack
| Component | What it is |
|---|---|
| Language | Go 1.24+, no CGo |
| Host | Gin, mounted with one call |
| Storage | SQLite, PostgreSQL |
| Shared counters | Redis, single, sentinel or cluster |
| Limiter strategies | sliding window (default), fixed window, token bucket |
| Pipeline | a 10,000 event ring buffer with background workers |
| Audit | hash chain, optional HMAC, 365 day retention |
| Accuracy | 10% false positive, 100% detection, pinned in CI |
| Dashboard | 15 pages embedded, 50+ API endpoints |
6.API design
Mounted under /sentinel
| Method | Endpoint | What it does |
|---|---|---|
| GET | /sentinel/ui | The embedded dashboard |
| GET | /sentinel/api/threats | Threat events, with severity and CVSS |
| GET | /sentinel/api/audit-logs/verify | Verify the hash chain |
| GET | /sentinel/api/performance/overview | Latency, error rates and pipeline drops |
| POST | /sentinel/api/blocks | Block an address, 24 hours by default |
| DELETE | /sentinel/api/settings/live | Discard dashboard overrides everywhere |
| POST | /sentinel/csp-report | CSP violations into the same dashboard |
How Grit mounts it, and why each line is there
// Counted per process, N replicas gave a client N times every limit.var counters sentinel.CounterStoreif svc.Cache != nil {counters = redisstore.New(svc.Cache.Client())}// MountE, so a misconfiguration in dev does not kill the host.err := sentinel.MountE(r, db, sentinel.Config{Counters: counters,WAF: sentinel.WAFConfig{Enabled: true,// Log in development, block in production.Mode: mode,// Empty means ignore X-Forwarded-For entirely, which is the// safe default. Populate it only behind a known proxy.TrustedProxies: cfg.SentinelTrustedProxies,MaxBodyBytes: 1 * 1024 * 1024,RejectOversizedBody: true,},})
Asserting your own paths against your exclusions
// Exclusions are matched against the real request path, not gin's route// template, so "/api/blogs/:id" matches the literal ":id" and nothing else.// That was silent dead config. Use a subtree match, and test it.m := sentinel.NewRouteMatcher(cfg.WAF.ExcludeRoutes)if !m.Matches("/v1/payments/collect") {t.Fatal("payments endpoint is not excluded from the WAF")}
7.Low level design
Core types
Mounts and returns an error. The earlier Mount killed the host process on a library failure, which is the wrong trade for a security add-on.
Typed issues for every silent-dead-config trap: unmatchable route patterns, unknown storage drivers that fall back to memory, broken regexes, alert sinks with no credentials, unreachable CAPTCHA tiers. Run by Mount, callable to fail a deploy.
The interface behind rate limiting and the lockout. The caller passes the clock so one request’s operations share a reading, and package countertest checks an implementation against the in-memory semantics.
A buffered channel, background workers, and atomic counters for emitted and dropped. Default capacity 10,000.
Decoding, patterns, scoring and sensitivity, with a corpus test that pins the false positive and detection rates in CI.
Exported, because the dashboard login limiter used gin’s version, which trusts the forwarded header from anyone, and could therefore be reset by rotating a header.
Design principles applied
- Measure the false positives. a corpus of legitimate-but-suspicious requests alongside the attacks, and both rates pinned in CI. Without it a rule change is a guess with production as the test.
- Scope every pattern. each rule declares where it may match. Both of the incidents that blocked real users were unscoped patterns matching somewhere nobody considered.
- Config that does nothing must be an error. four separate issues shared one cause: configuration that compiled, read as correct and was never consulted. A validator is the only defence against that class.
- Fail open on infrastructure, closed on attack. Redis unreachable allows the request; a matched attack pattern does not. The distinction is between our dependency failing and the caller being hostile.
- Keep the request path cheap. decide synchronously, analyse asynchronously, and count what the buffer could not take.
Patterns
| Pattern | Where it is used |
|---|---|
| Chain of responsibility | middleware each able to refuse |
| Async pipeline | a bounded buffer between decision and analysis |
| Strategy | three limiter algorithms behind one interface |
| Conformance suite | a shared test proving another counter store behaves the same |
| Hash chain | tamper-evident audit entries |
| Corpus testing | accuracy as a pinned number rather than a claim |
8.Scalability and performance
- Inspection is per request and cheap, so it scales with the application. The body cap is what bounds the worst case, and without it the attacker picks the cost.
- The asynchronous pipeline is what keeps analysis off the request path, and its buffer is the thing that fills under attack, which is when it matters.
- Shared counters cost one Redis round trip per limited request, halved in v2.6.0. That is the price of limits that mean the same thing behind every replica.
- Live config propagates within the sync interval, five seconds, so a block made on one replica applies on the others without a deploy.
- Audit and threat data grow steadily and have a retention setting, and at a year of retention the audit table is the largest thing the library owns.
- The dashboard is per process like any embedded UI, so the combined picture across replicas comes from shared storage rather than from the page.
9.Bottlenecks and improvements
What breaks first
- False positives. the dominant risk, and the one that presents as a fraction of users unable to do something rather than as an error anybody sees.
- Exclusions that match nothing. patterns are matched against the real path, not the route template, so an exclusion written as "/api/blogs/:id" matches the literal string and is dead config that looks correct.
- Pipeline saturation under attack. the event rate spikes exactly when the records matter most, and past the buffer they are dropped.
- Proxy configuration. trusting the forwarded header wrongly either lets clients spoof their address or makes every request appear to come from the load balancer, and the two mistakes look identical from inside.
- An embedded dashboard holding attack data. it contains request payloads and actor profiles, so it is itself a target and must not be reachable with default credentials.
What to do about it
- Run log mode first, then block. a week in log mode on real traffic shows what would have been refused. Grit does this by default in development for the same reason.
- Test the exclusions against real paths. the route matcher is exported so a test can assert that a concrete production path is excluded. That turns dead config into a failing test.
- Alert on the drop counter. it is on the performance overview. Watching it is how you learn that an attack was large enough that the record is incomplete.
- Be explicit about proxies. leave the trusted list empty to ignore the forwarded header entirely, or populate it with the actual proxy ranges. Never leave it to a default.
- Set the dashboard secret and password. an unset key is random per process, so sessions end on every restart, and the default password works only from localhost. Both are deliberate and both need a real value in production.
