Session Management System Design
Turning a stateless token into something you can revoke, list by device, and expire two different ways.
internal/sessioninternal/models1.Problem statement
A signed token is self-contained by design: the server can verify it without looking anything up. That property is also the problem. A token that needs no lookup cannot be cancelled, so "sign out this device", "sign out everywhere", and "that laptop was stolen" have no implementation.
The usual answer is a denylist of revoked tokens, which puts a lookup back in front of every request and grows without bound. The better answer is to keep the access token stateless and short, and to make the long-lived refresh token the thing that is tracked. A session row then represents a logged-in device, and revoking it means the next refresh fails.
That leaves one subtle attack. If a refresh token is copied, both the real client and the attacker can use it. Whoever refreshes second presents a token that has already been rotated away, and that is a signal: the pair has been cloned.
The system has to be able to:
- Record one row per logged-in device without ever storing a token that could be replayed from it.
- Rotate the refresh token on every use, so a captured one has a short window.
- Notice when a token that was already rotated is presented again, and treat the whole session as compromised.
- Expire a session two ways: after inactivity, and after an absolute age regardless of activity.
- Show a user their devices with enough detail to recognise one, and let them end any of them.
- End every session for a user when their password changes.
2.System requirements
Functional requirements
- Create a session on sign-in, carrying the user agent and address the request arrived with.
- Rotate on refresh: write the new token hash, keep the previous one.
- Reject and revoke on replay: a presented token matching the previous hash kills the session.
- Revoke one session by id, every session for a user, or every session except the current one.
- List a user their live sessions, newest activity first, with the current one marked.
- Expire idle sessions after seven days and all sessions after thirty, both configurable.
- Sweep dead rows on a schedule so the table stays proportional to live devices.
Non-functional requirements
- Unreplayable at rest. the table stores SHA-256 digests. Somebody with a dump can recognise a token you show them but cannot produce one.
- One lookup. validating a refresh is a single hit on a unique index, not a scan. The cost does not grow with the number of sessions.
- Bounded size. rows represent live devices, not history. An audit trail of sign-ins belongs in the audit log, which is append-only and has its own retention.
- Recognisable. a device list is useless if every row says "Mozilla/5.0". It carries the parsed agent, the address and the last-seen time.
- Fail closed. unknown, revoked, idle, expired and replayed all produce the same refusal. There is no state in which an ambiguous session is honoured.
3.Capacity estimation
Numbers for a mid-sized deployment. They are here to size the thing, not to predict your traffic: change an assumption and the sums below move with it.
Assumptions
| Parameter | Value |
|---|---|
| Daily active users | 200,000 |
| Devices per active user | 3 |
| Refresh interval | 15 minutes, while a client is open |
| Session row with indexes | ~650 bytes |
| Idle timeout | 7 days |
| Absolute timeout | 30 days |
Live sessions
Comfortably inside the primary database. The number is driven by devices, not by time, which is the point of sweeping.
Rotation write rate
This is the single heaviest write in the authentication path. It is an update to one row by unique key, which Postgres handles at this rate on modest hardware, but it is the number that decides whether the access token lifetime can be shortened.
Sweep volume
Small enough for one batched delete. Deleting in chunks rather than one statement keeps the lock short.
4.High level design
The session store sits between the token layer and the database. Nothing else writes to the table, which is what keeps the rotation invariant true.
Core components
- Session store. creates, rotates, revokes and lists. The only writer, so the rule that a rotation always records the previous hash lives in one function.
- Rotation check. given a presented refresh token, decides whether it belongs to a live session, a replayed one, or nothing at all.
- Expiry policy. two clocks per row. last_seen_at drives idle expiry and moves on every refresh; expires_at is fixed at creation and does not.
- Device list. the read side. Returns live sessions with the parsed agent and marks whichever row the current request is using.
- Sweeper. a scheduled job that deletes rows past their absolute expiry and revoked rows past the retention window.
Request flow
Refresh, rotation, and what a replay looks like
- 1The client posts its refresh token, either in the body or from the HttpOnly cookie scoped to /api/auth.
- 2The handler hands the raw token to the session store. No other component sees it.
- 3The store hashes it and looks for a live session by token_hash: not revoked, not past either deadline.
- 4If nothing matches by token_hash, the store looks for a match on prev_token_hash. A hit there means this token was already rotated away, so two clients hold the pair.
- 5That is treated as theft, not as a stale client. The session is revoked outright, which signs out both the attacker and the real user, who then signs in again.
- 6On a clean match the store writes the new hash, moves the old one to prev_token_hash, and advances last_seen_at. The JWT service mints the new pair.
- 7The rotation is recorded outside the response path, and a detected replay is recorded as a security event rather than an ordinary one.
Data flow
- The raw refresh token exists in memory for the length of one request and is never written anywhere.
- prev_token_hash is kept for exactly one generation. Keeping more would widen the window in which an old token is merely rejected rather than recognised as a replay.
- last_seen_at is an update on every refresh, which is why the row is designed to be narrow: the write is frequent.
- Revoked rows are kept rather than deleted so a device list can show recent sign-outs and the sweeper can age them out on its own schedule.
5.Technology stack
| Component | What it is |
|---|---|
| Storage | the primary relational database |
| Token digest | SHA-256, stored hex |
| Primary key | UUIDv7, time-ordered |
| Lookup index | unique on token_hash, secondary on prev_token_hash |
| Cookie | HttpOnly, SameSite=Lax, Path=/api/auth |
| Sweeper | a cron entry in the scheduling system |
6.Data model
sessions
One row per logged-in device. Narrow on purpose: last_seen_at is written on every refresh.
| Column | Holds |
|---|---|
| id | UUIDv7 |
| user_id | owner, indexed for the device list |
| token_hash | SHA-256 of the current refresh token, unique |
| prev_token_hash | the previous generation, indexed, for replay detection |
| user_agent | as sent, parsed for display |
| ip | the address the session was last refreshed from |
| created_at | when the device signed in |
| last_seen_at | moved on every refresh; drives idle expiry |
| expires_at | fixed at creation; the absolute deadline |
| revoked_at | set on sign-out, not deleted |
Where it lives
- One table, no cache in front of it. A refresh is already a write, so caching the read would save nothing and risk serving a revoked session.
- The unique index on token_hash is the hot path. The index on prev_token_hash is only touched when the first lookup misses, which is the rare case.
- Both timeouts are enforced in the query, not by a background job, so a session is dead the moment it qualifies rather than when the sweeper next runs.
7.API design
Sessions
| Method | Endpoint | What it does |
|---|---|---|
| POST | /api/v1/auth/refresh | Rotate and reissue |
| GET | /api/v1/auth/sessions | Live devices, newest activity first |
| DELETE | /api/v1/auth/sessions/:id | End one device |
| POST | /api/v1/auth/sessions/revoke-all | End every device |
| POST | /api/v1/auth/logout | End the current device only |
The device list
{"data": [{"id": "01a11e83-a9b0-75e4-9f8b-0ea9507f6994","user_agent": "Chrome 141 on Windows","ip": "102.203.209.219","created_at": "2026-10-02T08:14:09Z","last_seen_at": "2026-10-09T17:02:30Z","current": true},{"id": "01a0f2b1-5c3d-7a21-8e44-1b9cc2f40a77","user_agent": "Safari on iPhone","ip": "41.210.8.3","last_seen_at": "2026-10-07T21:40:11Z","current": false}]}
8.Low level design
Core types
The only writer. Every mutation goes through it, which is what keeps the rule that a rotation records the previous hash from being forgotten at a new call site.
CreateSessionRotateSessionRevokeRevokeAllForUserListForUserSweepThe three things a session needs from the request: the agent, the address and the context. Taking a value rather than a *gin.Context is what lets a background job create a session.
One error for every way a session can fail to be usable. Callers cannot accidentally distinguish "revoked" from "expired" in a response and tell an attacker which it was.
Design principles applied
- Single writer. the invariant is "a rotation always records the previous hash". It holds because there is exactly one function that can rotate.
- Interface segregation. the store takes a clock and a database handle, not a web framework. The expiry tests move the clock instead of sleeping.
- Tell, do not ask. callers ask for a rotation and get a result. They do not read the row, decide it is valid, and write it back, which would be a race.
Patterns
| Pattern | Where it is used |
|---|---|
| Token rotation | a new refresh token on every use, with one generation of memory |
| Sliding plus absolute expiry | two clocks, so neither activity nor inactivity alone decides |
| Tombstone | revoked rows persist until swept, so the device list can explain itself |
9.Scalability and performance
- The rotation write is a single-row update by unique key. It scales with the primary database and is the reason to think twice before shortening the access token lifetime.
- Reads for the device list are per-user and rare. They do not need an index beyond the one on user_id.
- The sweeper deletes in batches on a schedule rather than in one statement, so a large backlog does not hold a lock for minutes.
- Because validity is computed in the query, a replica lagging behind the primary can briefly honour a session that was just revoked. Revocation therefore reads and writes the primary.
- Nothing here is cached. A cache would turn revocation from immediate into eventually consistent, which is the one property this system exists to provide.
10.Bottlenecks and improvements
What breaks first
- Write amplification from short tokens. the rotation rate is inversely proportional to the access token lifetime. A 5 minute token triples the write rate on this table.
- Replay false positives. a client that retries a refresh after a network timeout can legitimately present a token that already rotated, and gets its session killed for it.
- Unbounded growth without a sweeper. if the scheduled sweep stops running, nothing else deletes rows, and the unique index grows with every sign-in ever made.
- Device lists that nobody can read. a raw user agent string tells a user nothing, which makes the "end this session" button unusable in practice.
What to do about it
- Jittered refresh. clients refresh at a random point in the last third of the lifetime, which spreads the write rate instead of synchronising it.
- A grace window on rotation. accepting the previous token for a few seconds after a rotation, once, distinguishes a retry from a replay without weakening the detection.
- Sweep monitoring. the sweeper reports rows deleted. A sudden zero is the signal that it stopped, long before the table size is a problem.
- Parse the agent on write. storing both the raw string and a readable summary means the list is useful without parsing on every read.
