Rate Limiting and Abuse Control System Design
Keeping one caller from consuming the capacity of all of them, and making a password-guessing attack cost something.
Sentinelinternal/middleware1.Problem statement
An API with no limit serves whoever asks loudest. A script with a loop will take as much capacity as the hardware can give it, and the users it starves are indistinguishable from the attacker until somebody reads a log.
There are three different problems wearing the same name. Protecting capacity means limiting how often anybody can call anything. Protecting an account means limiting how often somebody can guess its password, which has to be counted per account rather than per address or a botnet defeats it. Protecting an expensive endpoint means limiting the specific operations that cost real money, like sending an email or generating a PDF.
They need different keys, different windows and different responses, which is why one global limit is never enough.
The system has to be able to:
- Limit requests per address, with a window short enough to matter and long enough not to punish a burst.
- Limit sign-in attempts per account, independently of where they come from.
- Limit specific expensive operations more tightly than ordinary reads.
- Tell a rejected caller when to come back, rather than simply refusing.
- Keep working when the shared counter store is unavailable, without either opening up or shutting down.
- Make the limits visible, so an operator can see who is being limited and why.
2.System requirements
Functional requirements
- Per-address limiting on the whole API, configurable per route group.
- Account lockout after a number of failed sign-ins within a window, counted per account.
- Tighter limits on sign-in, password reset, invitation and other endpoints that send mail.
- Standard rate limit headers on every response, and Retry-After on a refusal.
- A dashboard showing current limits, recent refusals and the addresses involved.
- Limits that survive a restart, because an attacker can simply wait for a deploy otherwise.
Non-functional requirements
- Cheap on the happy path. the overwhelming majority of requests are not limited. The check must cost microseconds, not a database round trip.
- Counted before the expensive work. the limiter runs before password hashing and before the handler, or the attack still costs the server what it was meant to save.
- Honest refusals. a 429 with Retry-After lets a well-behaved client back off correctly instead of retrying immediately and making it worse.
- Degrades sensibly. if the shared counter is unreachable, the limiter falls back to a local one rather than failing open to unlimited or closed to nothing.
- Not a membership oracle. account lockout applies to real accounts. An unknown address must not behave differently, or the endpoint reveals who is registered.
3.Capacity estimation
Numbers for a mid-sized deployment. They are here to size the thing, not to predict your traffic: change an assumption and the sums below move with it.
Assumptions
| Parameter | Value |
|---|---|
| Requests per second | 926 average, 4,630 peak |
| Distinct client addresses per minute | ~30,000 |
| Default limit | 100 requests per minute per address |
| Sign-in limit | 10 failures per account per 15 minutes |
| Counter entry | ~64 bytes |
Counter memory
The state is tiny. Rate limiting is not a storage problem; it is a coordination problem between replicas.
Check cost
The difference is the whole design question. A per-process counter is free and wrong by a factor of the replica count; a shared one is correct and adds a hop to every request.
Effective limit across replicas
This is why a shared store matters once there is more than one replica: otherwise the configured number is not the enforced number.
4.High level design
Three limiters with different keys, all running before the handler. The first protects capacity, the second protects accounts, the third protects specific costs.
Core components
- Request limiter. per address, applied across the API. The broad one, sized so that normal use never meets it.
- Account lockout. per account, counting failed sign-ins. Independent of address, which is what makes it work against a distributed attack.
- Endpoint limiter. tighter rules on the routes that cost money or send mail, keyed by whichever identifier fits that route.
- Counter store. shared when available so the configured limit is the enforced limit, with a local fallback when it is not.
- Dashboard. what is being limited, how often, and from where. A limiter nobody can observe is a limiter nobody can tune.
Request flow
Where the limiters sit relative to the work
- 1The request arrives. Nothing has been parsed beyond the route and the client address.
- 2The request limiter increments the counter for that address and reads the current count.
- 3If the shared store is unreachable the limiter uses a per-process counter instead. The limit becomes approximate rather than absent, which is the right trade in an outage.
- 4Over the limit, the request ends here with 429 and a Retry-After naming when the window resets. No handler ran and nothing was parsed.
- 5On a sign-in route the account lockout check runs next, keyed by the submitted address. An account already barred is refused before anything expensive happens.
- 6Only now does the handler run.
- 7Password hashing happens last, which is the point of the ordering: an attack that is going to be rejected is rejected before it costs sixty milliseconds of CPU.
- 8Counters feed a dashboard, so the limits can be tuned against what is actually happening rather than guessed at.
Data flow
- Counters are keyed by address for the broad limit and by account for lockout. Mixing the two produces a limiter that a botnet defeats and that punishes an office behind one address.
- An unknown email address never increments an account counter, or anybody could lock out an address they can guess.
- Rate limit headers are set on successful responses too, so a client can slow down before it is refused.
- The counter store is the only shared state. Everything else is per-request.
5.Technology stack
| Component | What it is |
|---|---|
| Request limiting | Sentinel, a sliding window per key |
| Counter store | Redis when configured, in-process otherwise |
| Account lockout | database-backed, so it survives a restart |
| Headers | X-RateLimit-Limit, X-RateLimit-Remaining, Retry-After |
| Status | 429 Too Many Requests |
| Observability | the Sentinel dashboard, mounted in the API |
6.Data model
login_attempts
Durable on purpose. An in-memory counter is reset by a deploy, which is a window an attacker can simply wait for.
| Column | Holds |
|---|---|
| the account being guessed at, indexed | |
| failed_count | attempts inside the current window |
| locked_until | when it becomes usable again |
| last_attempt_at | drives the window |
Where it lives
- Request counters are ephemeral and belong in a store with expiry. Losing them on a restart costs one window of accuracy.
- Account lockout is durable, because the whole point is that it cannot be cleared by waiting for a deploy.
7.API design
What a limited caller sees
| Method | Endpoint | What it does |
|---|---|---|
| GET | any route | Rate limit headers on every response, not only refusals |
| POST | /api/v1/auth/login | 429 when limited, 423 when the account is locked |
Operations
| Method | Endpoint | What it does |
|---|---|---|
| GET | /sentinel/ui | The dashboard: current limits and recent refusals |
| GET | /api/v1/admin/security/lockouts | Accounts currently barred |
| DELETE | /api/v1/admin/security/lockouts/:email | Clear one, for support |
A refusal that tells the client what to do
HTTP/1.1 429 Too Many RequestsX-RateLimit-Limit: 100X-RateLimit-Remaining: 0Retry-After: 37{"error": {"code": "RATE_LIMITED","message": "Too many requests. Try again in 37 seconds."}}
8.Low level design
Core types
The request limiter. Runs before everything, keyed by address, with per-route-group configuration.
The account counter. Increments only on a failed attempt against a real account, and reads before any hashing is paid for.
IsLockedRecordFailureClearAn interface with two implementations, shared and local. The fallback is a decision the limiter makes, not an outage the request sees.
Design principles applied
- Reject before you pay. the ordering of the middleware chain is the security property. A limiter after the expensive work limits nothing that matters.
- Key by the thing you are protecting. capacity is per address, accounts are per account, cost is per operation. One key cannot serve all three.
- Degrade, do not fail. an unreachable counter store makes the limit approximate. Failing open removes protection; failing closed turns a cache outage into an outage.
- Make it observable. a limit nobody can see is tuned by guesswork and discovered when a customer complains.
Patterns
| Pattern | Where it is used |
|---|---|
| Sliding window | smoother than a fixed window at the boundary |
| Token bucket | for endpoints that should tolerate a burst and then slow down |
| Circuit fallback | local counting when the shared store is unavailable |
| Middleware ordering | cheap checks first, expensive work last |
9.Scalability and performance
- A shared counter store makes the configured limit the enforced limit across replicas. Without it, the real limit is the configured one multiplied by the replica count.
- The counter check adds one round trip to the shared store. On a local network that is a fraction of a millisecond, which is the price of the limit being true.
- Counters expire with their window, so the store does not grow with traffic over time.
- Per-tenant keys matter in a multitenant deployment, or one customer can consume the capacity that everybody shares.
- Limits are configured per route group, so an expensive export can be a hundred times tighter than a list without changing the global number.
10.Bottlenecks and improvements
What breaks first
- Shared addresses. an office or a mobile carrier presents one address for thousands of people, who then share one limit and all get refused together.
- The counter store as a dependency. putting a network call in front of every request makes the limiter a new thing that can be down.
- Lockout as a denial of service. if an attacker can lock any account by guessing wrong ten times, lockout is a weapon pointed at your users.
- Window boundary bursts. a fixed window lets a caller spend the whole allowance at the end of one and the start of the next, for double the intended rate.
What to do about it
- Key authenticated traffic by user. once a caller is identified, limiting by user rather than by address removes the shared-address problem entirely.
- Local fallback with a short leash. the limiter keeps working on a per-process counter when the store is away, and says so in the dashboard rather than silently.
- Delay rather than lock. an increasing delay after each failure costs an attacker far more than it costs a user who mistyped, without handing anybody a lockout weapon.
- Sliding windows. counting over a moving interval removes the boundary burst without the bookkeeping of a full leaky bucket.
