All systems
Access

Rate Limiting and Abuse Control System Design

Keeping one caller from consuming the capacity of all of them, and making a password-guessing attack cost something.

Sentinelinternal/middleware

1.Problem statement

An API with no limit serves whoever asks loudest. A script with a loop will take as much capacity as the hardware can give it, and the users it starves are indistinguishable from the attacker until somebody reads a log.

There are three different problems wearing the same name. Protecting capacity means limiting how often anybody can call anything. Protecting an account means limiting how often somebody can guess its password, which has to be counted per account rather than per address or a botnet defeats it. Protecting an expensive endpoint means limiting the specific operations that cost real money, like sending an email or generating a PDF.

They need different keys, different windows and different responses, which is why one global limit is never enough.

The system has to be able to:

  • Limit requests per address, with a window short enough to matter and long enough not to punish a burst.
  • Limit sign-in attempts per account, independently of where they come from.
  • Limit specific expensive operations more tightly than ordinary reads.
  • Tell a rejected caller when to come back, rather than simply refusing.
  • Keep working when the shared counter store is unavailable, without either opening up or shutting down.
  • Make the limits visible, so an operator can see who is being limited and why.

2.System requirements

Functional requirements

  • Per-address limiting on the whole API, configurable per route group.
  • Account lockout after a number of failed sign-ins within a window, counted per account.
  • Tighter limits on sign-in, password reset, invitation and other endpoints that send mail.
  • Standard rate limit headers on every response, and Retry-After on a refusal.
  • A dashboard showing current limits, recent refusals and the addresses involved.
  • Limits that survive a restart, because an attacker can simply wait for a deploy otherwise.

Non-functional requirements

  • Cheap on the happy path. the overwhelming majority of requests are not limited. The check must cost microseconds, not a database round trip.
  • Counted before the expensive work. the limiter runs before password hashing and before the handler, or the attack still costs the server what it was meant to save.
  • Honest refusals. a 429 with Retry-After lets a well-behaved client back off correctly instead of retrying immediately and making it worse.
  • Degrades sensibly. if the shared counter is unreachable, the limiter falls back to a local one rather than failing open to unlimited or closed to nothing.
  • Not a membership oracle. account lockout applies to real accounts. An unknown address must not behave differently, or the endpoint reveals who is registered.

3.Capacity estimation

Numbers for a mid-sized deployment. They are here to size the thing, not to predict your traffic: change an assumption and the sums below move with it.

Assumptions

ParameterValue
Requests per second926 average, 4,630 peak
Distinct client addresses per minute~30,000
Default limit100 requests per minute per address
Sign-in limit10 failures per account per 15 minutes
Counter entry~64 bytes

Counter memory

30,000 active addresses x 64 bytes = ~2 MB
plus sign-in counters for accounts under attack, small

The state is tiny. Rate limiting is not a storage problem; it is a coordination problem between replicas.

Check cost

in-process counter: one map lookup and an increment, ~50 nanoseconds
shared counter: one round trip, ~0.3 ms

The difference is the whole design question. A per-process counter is free and wrong by a factor of the replica count; a shared one is correct and adds a hop to every request.

Effective limit across replicas

100/minute per address, enforced per process
x 6 replicas = up to 600/minute actually allowed

This is why a shared store matters once there is more than one replica: otherwise the configured number is not the enforced number.

4.High level design

Three limiters with different keys, all running before the handler. The first protects capacity, the second protects accounts, the third protects specific costs.

Core components

  • Request limiter. per address, applied across the API. The broad one, sized so that normal use never meets it.
  • Account lockout. per account, counting failed sign-ins. Independent of address, which is what makes it work against a distributed attack.
  • Endpoint limiter. tighter rules on the routes that cost money or send mail, keyed by whichever identifier fits that route.
  • Counter store. shared when available so the configured limit is the enforced limit, with a local fallback when it is not.
  • Dashboard. what is being limited, how often, and from where. A limiter nobody can observe is a limiter nobody can tune.

Request flow

Where the limiters sit relative to the work

12345678ClientRequest LimiterCounter StoreLocal Fallback429 + Retry-AfterAccount LockoutHandlerPassword HashDashboard
  1. 1The request arrives. Nothing has been parsed beyond the route and the client address.
  2. 2The request limiter increments the counter for that address and reads the current count.
  3. 3If the shared store is unreachable the limiter uses a per-process counter instead. The limit becomes approximate rather than absent, which is the right trade in an outage.
  4. 4Over the limit, the request ends here with 429 and a Retry-After naming when the window resets. No handler ran and nothing was parsed.
  5. 5On a sign-in route the account lockout check runs next, keyed by the submitted address. An account already barred is refused before anything expensive happens.
  6. 6Only now does the handler run.
  7. 7Password hashing happens last, which is the point of the ordering: an attack that is going to be rejected is rejected before it costs sixty milliseconds of CPU.
  8. 8Counters feed a dashboard, so the limits can be tuned against what is actually happening rather than guessed at.

Data flow

  • Counters are keyed by address for the broad limit and by account for lockout. Mixing the two produces a limiter that a botnet defeats and that punishes an office behind one address.
  • An unknown email address never increments an account counter, or anybody could lock out an address they can guess.
  • Rate limit headers are set on successful responses too, so a client can slow down before it is refused.
  • The counter store is the only shared state. Everything else is per-request.

5.Technology stack

ComponentWhat it is
Request limitingSentinel, a sliding window per key
Counter storeRedis when configured, in-process otherwise
Account lockoutdatabase-backed, so it survives a restart
HeadersX-RateLimit-Limit, X-RateLimit-Remaining, Retry-After
Status429 Too Many Requests
Observabilitythe Sentinel dashboard, mounted in the API

6.Data model

login_attempts

Durable on purpose. An in-memory counter is reset by a deploy, which is a window an attacker can simply wait for.

ColumnHolds
emailthe account being guessed at, indexed
failed_countattempts inside the current window
locked_untilwhen it becomes usable again
last_attempt_atdrives the window

Where it lives

  • Request counters are ephemeral and belong in a store with expiry. Losing them on a restart costs one window of accuracy.
  • Account lockout is durable, because the whole point is that it cannot be cleared by waiting for a deploy.

7.API design

What a limited caller sees

MethodEndpointWhat it does
GETany routeRate limit headers on every response, not only refusals
POST/api/v1/auth/login429 when limited, 423 when the account is locked

Operations

MethodEndpointWhat it does
GET/sentinel/uiThe dashboard: current limits and recent refusals
GET/api/v1/admin/security/lockoutsAccounts currently barred
DELETE/api/v1/admin/security/lockouts/:emailClear one, for support

A refusal that tells the client what to do

HTTP/1.1 429 Too Many Requests
X-RateLimit-Limit: 100
X-RateLimit-Remaining: 0
Retry-After: 37
{
"error": {
"code": "RATE_LIMITED",
"message": "Too many requests. Try again in 37 seconds."
}
}

8.Low level design

Core types

Sentinel middleware

The request limiter. Runs before everything, keyed by address, with per-route-group configuration.

LoginAttemptService

The account counter. Increments only on a failed attempt against a real account, and reads before any hashing is paid for.

IsLockedRecordFailureClear
CounterStore

An interface with two implementations, shared and local. The fallback is a decision the limiter makes, not an outage the request sees.

Design principles applied

  • Reject before you pay. the ordering of the middleware chain is the security property. A limiter after the expensive work limits nothing that matters.
  • Key by the thing you are protecting. capacity is per address, accounts are per account, cost is per operation. One key cannot serve all three.
  • Degrade, do not fail. an unreachable counter store makes the limit approximate. Failing open removes protection; failing closed turns a cache outage into an outage.
  • Make it observable. a limit nobody can see is tuned by guesswork and discovered when a customer complains.

Patterns

PatternWhere it is used
Sliding windowsmoother than a fixed window at the boundary
Token bucketfor endpoints that should tolerate a burst and then slow down
Circuit fallbacklocal counting when the shared store is unavailable
Middleware orderingcheap checks first, expensive work last

9.Scalability and performance

  • A shared counter store makes the configured limit the enforced limit across replicas. Without it, the real limit is the configured one multiplied by the replica count.
  • The counter check adds one round trip to the shared store. On a local network that is a fraction of a millisecond, which is the price of the limit being true.
  • Counters expire with their window, so the store does not grow with traffic over time.
  • Per-tenant keys matter in a multitenant deployment, or one customer can consume the capacity that everybody shares.
  • Limits are configured per route group, so an expensive export can be a hundred times tighter than a list without changing the global number.

10.Bottlenecks and improvements

What breaks first

  • Shared addresses. an office or a mobile carrier presents one address for thousands of people, who then share one limit and all get refused together.
  • The counter store as a dependency. putting a network call in front of every request makes the limiter a new thing that can be down.
  • Lockout as a denial of service. if an attacker can lock any account by guessing wrong ten times, lockout is a weapon pointed at your users.
  • Window boundary bursts. a fixed window lets a caller spend the whole allowance at the end of one and the start of the next, for double the intended rate.

What to do about it

  • Key authenticated traffic by user. once a caller is identified, limiting by user rather than by address removes the shared-address problem entirely.
  • Local fallback with a short leash. the limiter keeps working on a per-process counter when the store is away, and says so in the dashboard rather than silently.
  • Delay rather than lock. an increasing delay after each failure costs an attacker far more than it costs a user who mistyped, without handing anybody a lockout weapon.
  • Sliding windows. counting over a moving interval removes the boundary burst without the bookkeeping of a full leaky bucket.

Read next