All systems
Data

Caching System Design

Serving a repeated answer without asking again, while keeping a cache failure a slowdown rather than an outage.

internal/cacheinternal/middleware

1.Problem statement

Most read traffic asks the same questions. A dashboard count, a settings lookup, a permission set, a list that changes hourly and is read thousands of times an hour. Recomputing each of them from the database is work the database did not need to do.

A cache is easy to add and easy to get wrong in four specific ways. It can serve a stale answer after a write. It can become a dependency, so that losing it takes the application down rather than slowing it. It can stampede, where a hot key expires and every concurrent request rebuilds it at once. And a deploy can write ten thousand keys in the same second, which then all expire in the same second and hand the database the whole working set back at once.

Each of those has a known remedy, and the value of a cache layer in a framework is that the remedies are already in it rather than being rediscovered per project.

The system has to be able to:

  • Read through: ask for a value, get it from the cache or compute and store it, in one call.
  • Fall through to the loader when the cache is unavailable, so an outage is a slowdown.
  • Spread expiry times so a synchronised write does not become a synchronised miss.
  • Let one caller rebuild a hot key while the others wait briefly rather than all rebuilding.
  • Invalidate by key and by prefix when the underlying data changes.
  • Run with no cache at all, because a small deployment should not need Redis.

2.System requirements

Functional requirements

  • A Remember call taking a key, a time to live, a destination and a loader.
  • Get, Set, Delete and DeleteByPrefix for the cases that are not read-through.
  • A response cache for whole GET responses on routes that opt in.
  • Automatic invalidation on write for generated resources, by prefix.
  • Jitter of up to ten per cent added to every time to live.
  • Single-flight rebuilding of an expired key.
  • Complete absence as a supported configuration: no Redis URL means no cache and no errors.

Non-functional requirements

  • Never load-bearing. every read path works with the cache removed. That is what makes it acceptable for a cache to be a network dependency at all.
  • Correct before fast. a write invalidates before it returns. Serving a value the caller just changed is worse than not caching.
  • Keyed completely. a key includes everything that varies the answer, including the owner and the tenant. A key that omits them is a data leak, not a performance bug.
  • Bounded. everything has a time to live. A cache with no expiry is a second database with no migrations.
  • Observable. hit rate and eviction are visible. A cache nobody measures is a cache nobody can tune.

3.Capacity estimation

Numbers for a mid-sized deployment. They are here to size the thing, not to predict your traffic: change an assumption and the sums below move with it.

Assumptions

ParameterValue
Read requests per second926 average, 4,630 peak
Cacheable share of reads~60%
Target hit rate85%
Average cached value4 KB
Distinct hot keys~200,000
Default time to live5 minutes

Database load avoided

926 reads/s x 60% cacheable = 556 cacheable reads/s
at 85% hit rate = 473 reads/s served from cache
database sees 453 reads/s instead of 926

Roughly half the read load. The number that matters is not the hit rate on its own but the hit rate times the cacheable share.

Memory

200,000 keys x 4 KB = 800 MB
plus Redis overhead of ~100 bytes/key = ~820 MB
a 1 GB instance holds the working set with room to spare

Stampede without protection

a hot key read 200 times/second, time to live expires
without single-flight: 200 concurrent rebuilds of the same query
with it: 1 rebuild, 199 short waits

The failure is worst exactly when the key is most valuable, which is why this is not an optimisation but a correctness feature of the cache layer.

Synchronised expiry

a deploy warms 10,000 keys in one second with a 5 minute life
without jitter: 10,000 misses in the same second, five minutes later
with up to 10% jitter: spread over 30 seconds

4.High level design

One read-through call covers most uses. The response cache and the invalidation hooks are layered on top of the same store.

Core components

  • Cache client. the store interface. Backed by Redis when configured and by nothing when not, so the absence case needs no branches at the call sites.
  • Remember. the read-through call. Cache-aside with the three things a hand-written version usually misses: fall-through on failure, jittered expiry and single-flight rebuilds.
  • Response cache. middleware that stores a whole GET response for routes that opt in, keyed by path, query and identity.
  • Invalidation hooks. generated resources delete their prefix on create, update and delete, so a write cannot leave a stale list behind.
  • Key builder. the convention that makes keys include the owner and the tenant, so no key can be shared across a boundary.

Request flow

A read-through cache with a write invalidating it

12345678ClientHandlerRememberCache StoreSingle FlightLoaderDatabaseInvalidate PrefixWrite Path
  1. 1A read request arrives on a route whose answer is worth caching.
  2. 2The handler calls Remember with a key, a lifetime, a destination and a function that can compute the value.
  3. 3The cache is asked. A hit returns immediately. A cache that is unreachable is treated as a miss rather than an error, which is what keeps it from being load-bearing.
  4. 4On a miss, single-flight ensures one caller runs the loader while the others wait for its result instead of all running the same query.
  5. 5The loader computes the value, usually a query the handler would have run anyway.
  6. 6The result is stored with the requested lifetime plus up to ten per cent of jitter, so keys written together do not expire together.
  7. 7Separately, a write to the underlying resource deletes the cache prefix for it.
  8. 8Invalidation happens before the write returns, so a client that reads immediately after writing cannot see the value it just replaced.

Data flow

  • Keys carry everything that varies the answer. For an owned resource that includes the owner, and in a multitenant deployment the organisation.
  • Nothing is cached without a lifetime. An entry that never expires is state that no migration will ever fix.
  • Invalidation is by prefix rather than by enumerating keys, because the set of list keys for a resource is unbounded in query parameters.
  • A cache miss and a cache outage take the same path, so there is no error handling at the call site to get wrong.

5.Technology stack

ComponentWhat it is
StoreRedis, optional
Patterncache-aside, read-through
Default lifetime5 minutes, jittered by up to 10%
Stampede controlsingle-flight per key
Invalidationby key and by prefix, on write
Absenceno Redis URL disables cache, jobs and cron, with no errors

6.Data model

There is no schema. The design question is the key, and the convention is what keeps it correct.

key convention

Everything that changes the answer appears in the key. The commonest cache bug in any system is a key that omits the owner.

ColumnHolds
resource:ida single record
resource:list:<hash of query>a list, hashed because the parameter space is unbounded
user:<id>:...anything that varies per user
org:<id>:...anything that varies per tenant
perm:role:<id>resolved permission sets, invalidated when a role changes

7.Low level design

Core types

cache.Clientinternal/cache/cache.go

The store. Get, Set, Delete, DeleteByPrefix. Every method tolerates the backend being absent or unreachable.

GetSetDeleteDeleteByPrefix
cache.Rememberinternal/cache/remember.go

The call that should be used by default. Generic over the value type, so the destination is typed rather than an interface the caller asserts.

Response cache middleware

Caches whole GET responses for opted-in routes. Opt-in rather than automatic, because a response that varies by identity is easy to cache wrongly.

Design principles applied

  • Make the correct call the shortest one. Remember is one line and handles the four classic mistakes. Get and Set are available and longer to write, which is the right incentive.
  • A dependency that may be absent. the client treats unreachable as empty. There is no error to handle at the call site and therefore no error handling to get wrong.
  • Invalidate at the write, not on a timer. relying on a short lifetime to hide staleness means choosing between stale data and no caching. Invalidating removes the choice.

Patterns

PatternWhere it is used
Cache asideread through, write around, invalidate on change
Single flightone rebuild per key, the rest wait
Jittered expiryspreading synchronised writes so they do not become synchronised misses
Null objectthe no-op cache when none is configured

8.Scalability and performance

  • A cache hit removes a database round trip, which is the cheapest capacity available until the hit rate stops improving.
  • Hit rate times cacheable share is the real number. A 99% hit rate on 5% of traffic is worth less than 80% on 60%.
  • Memory grows with the working set, not with total data. Sizing follows from distinct hot keys, not from table sizes.
  • Redis is single-threaded per instance. Very hot single keys are the limit, which single-flight helps with on the application side.
  • Prefix invalidation is a scan on some backends. Keeping prefixes narrow keeps that cost bounded.
  • With no cache configured the application is slower and entirely correct, which is the right default for a small deployment.

9.Bottlenecks and improvements

What breaks first

  • Keys that omit the owner. the worst cache bug there is, because it does not look like a bug. It looks like one user occasionally seeing another user’s data.
  • The cache becoming load-bearing. a path that errors when the cache is unreachable has turned an optional component into a required one, usually without anybody deciding to.
  • Invalidation gaps. a write path that does not go through the service does not invalidate, and the stale value lives for a full lifetime.
  • Unbounded key growth. list keys hashed from arbitrary query parameters can be generated without limit by a caller who varies them.

What to do about it

  • Build keys, do not write them. a helper that takes the context and appends the owner and tenant makes the complete key the easy one.
  • Test with the cache off. running the suite with no cache configured proves no path depends on it, which is the property that is otherwise assumed.
  • Invalidate in the service. putting the hook where every writer passes, rather than in the handler, closes the gap for jobs and commands.
  • Whitelist the query parameters. hashing only the parameters that actually vary the answer bounds the key space regardless of what a caller sends.

Read next