Feature Flags System Design
Turning a feature on for ten per cent of users without a deploy, and having the same user stay in the same group.
internal/flagsinternal/models1.Problem statement
A redesigned checkout is ready. Shipping it to everybody at once means that if it is wrong, the fix is a deploy, and the damage is everybody. Shipping it to ten per cent means the damage is ten per cent and the fix is a toggle.
That requires the decision to be made at runtime rather than at build time, which rules out a constant and a compile-time flag. It also requires the decision to be stable per user: somebody who sees the new checkout on Monday must see it on Tuesday, or they are not a user of either version, they are a user of a flickering application.
The naive implementation reads a row per check. A page rendering a dozen flagged components does a dozen queries, and the flag system becomes the slowest part of the request it was meant to make safe. The naive exposure tracking is worse: a goroutine and an insert per check turned a busy page into thousands of writes a second.
And a flag is only half useful without knowing who saw what. A ten per cent rollout with no record of which users were in it cannot be measured, so the experiment produces a feeling rather than a result.
The system has to be able to:
- Turn a feature on or off at runtime, without a deploy.
- Roll out to a percentage of users, stably per user.
- Serve several variants, for comparing two new versions rather than one.
- Check a flag without touching the database.
- Propagate a change in seconds, not on the next deploy.
- Record exposures, so a rollout can be measured.
- Target specific users, roles or tenants, not only percentages.
- Tell clients a flag changed, so they refetch rather than waiting.
2.System requirements
Functional requirements
- A flags table with a key, a state, rules and variants.
- An engine holding the flags in memory, one per process.
- A background refresh on an interval, 30 seconds by default.
- An immediate refresh and a realtime broadcast on an administrative change.
- Percentage rollout bucketed per user and flag.
- Variants, for more than two arms.
- Exposure records, written by one writer rather than per check.
- An admin page to create, edit, roll out and delete.
- A client-side hook reading the evaluated flags for the current user.
Non-functional requirements
- A flag check never touches the database. all flags are in memory. A page with a dozen checks must not be a page with a dozen queries, or the mechanism costs more than the feature it guards.
- Sticky per user and flag. the bucket is a hash of the user identifier and the flag name, so a user always lands in the same bucket for a given flag and their assignment does not flicker between sessions.
- Independent per flag. including the flag name in the hash means the ten per cent for one flag is not the same ten per cent as another. Otherwise one unlucky cohort gets every experiment.
- Propagates in seconds. 30 seconds from the refresh, immediately on an administrative change. A flag that needs a deploy to take effect is a constant.
- Exposure tracking never blocks. fire and forget to a single writer. A goroutine and an insert per check was thousands of writes a second on a busy page.
- Fails to a known state. an unknown flag is off. The default has to be the safe one, because the failure case is a flag that was never created or was just deleted.
3.Capacity estimation
Numbers for a mid-sized deployment. They are here to size the thing, not to predict your traffic: change an assumption and the sums below move with it.
Assumptions
| Parameter | Value |
|---|---|
| Flags in a mature project | 20 to 100 |
| Flag checks per request | 1 to 12 |
| Requests per second | 926 average, 4,630 peak |
| Refresh interval | 30 seconds |
| Flag row size | ~400 bytes |
Check cost
Against a database read per check at ~1 ms, 11,000 checks a second would be eleven seconds of database time per second. That is the difference the in-memory cache makes.
Refresh load
Memory
Which is why holding all of them is obviously right rather than a trade-off worth analysing.
Exposure writes
The naive version is the thing that broke. A single writer with batching is three orders of magnitude fewer statements for the same information.
4.High level design
One engine per process holding every flag in memory, refreshed on a timer and on change, with bucketing by hash.
Core components
- Flags table. the definition: a key, whether it is on, its rules, its rollout percentage and its variants.
- Engine. the in-memory map, one per process, pre-warmed at boot behind a read-write lock. A flag check reads the map and never the database.
- Refresher. a background loop pulling fresh state every 30 seconds, which is quick enough for an administrative change to feel immediate without polling hard.
- Bucketer. SHA-256 of the user identifier and the flag name, modulo 100. Sticky per pair, and independent across flags because the name is in the hash.
- Variant selector. the same bucket mapped onto several arms, for comparing two new versions rather than one against the old.
- Exposure writer. one writer fed by a channel. Each check used to start a goroutine and an insert of its own, which a busy page turned into thousands a second.
- Change broadcast. an administrative change refreshes the engine immediately and publishes a realtime event, so subscribed clients refetch instead of waiting out the interval.
Request flow
A flag check, and a change propagating
- 1A request renders a page with a dozen flagged components, each asking whether its feature is on for this user.
- 2Each check reads the in-memory map under a read lock. No database round trip, which is the property that lets a page have twelve checks instead of one.
- 3For a percentage rollout, the bucket is SHA-256 of the user identifier and the flag name, modulo 100. The user identifier makes it sticky, so Monday’s assignment is Tuesday’s. The flag name makes flags independent, so the unlucky ten per cent of one experiment is not the unlucky ten per cent of all of them.
- 4With variants, the same bucket maps onto several arms rather than on and off, which is how two new versions get compared against each other instead of only against the old one.
- 5The exposure is recorded, fire and forget, through one writer fed by a channel. Every check starting its own goroutine and its own insert is what made a busy page thousands of writes a second.
- 6Separately, somebody changes a flag in the admin: enables it, moves it to twenty-five per cent, adds a variant.
- 7The engine refreshes immediately rather than waiting out the interval, and the timer continues as the fallback that keeps every other process in step within thirty seconds.
- 8A realtime event goes out, so clients holding evaluated flags refetch rather than showing a stale answer until their next navigation. An anonymous user buckets on a random per-request value, which is effectively random, so anything that needs stickiness without a login passes a session or device identifier instead.
Data flow
- The user identifier and the flag name are both in the hash. Omitting the first makes assignment flicker; omitting the second makes one cohort receive every experiment.
- An unknown flag is off. That covers the window after a deploy that references a flag nobody created yet, and the window after a deletion.
- Exposures are a stream to one writer, not a write per check. The information is the same and the statement count is three orders of magnitude lower.
- An anonymous check is random per request unless the caller supplies a stable identifier, and the package says so rather than pretending otherwise.
5.Technology stack
| Component | What it is |
|---|---|
| Storage | a flags table and an exposures table |
| Evaluation | in memory, one engine per process |
| Refresh | 30 seconds, plus immediate on change |
| Bucketing | SHA-256 of user id and flag name, modulo 100 |
| Variants | several arms over the same bucket |
| Exposure | fire and forget to a single batched writer |
| Propagation | a realtime flag.updated event |
6.Data model
feature_flags
| Column | Holds |
|---|---|
| key | the name used in code |
| enabled | the master switch |
| rules | JSON: rollout percentage, targeted users, roles, tenants |
| variants | JSON: the arms and their weights |
| description | what it is for, because a flag outlives whoever added it |
flag_exposures
This is what turns a rollout into a measurement. Without it, a ten per cent experiment produces a feeling rather than a result.
| Column | Holds |
|---|---|
| flag_key | which flag |
| user_id | who saw it |
| variant | which arm they got |
| created_at | when |
7.API design
Flags
| Method | Endpoint | What it does |
|---|---|---|
| GET | /api/v1/flags | Evaluated flags for the current caller |
| GET | /api/v1/admin/flags | All flags with their rules |
| POST | /api/v1/admin/flags | Create one |
| PUT | /api/v1/admin/flags/:key | Change state, rollout or variants |
| DELETE | /api/v1/admin/flags/:key | Remove it |
Two arms, or several
if flags.IsEnabled(c, "new_checkout") {return h.newCheckout(c)}switch flags.Variant(c, "checkout_redesign") {case "control": return h.oldFlow(c)case "variant_a": return h.newFlow(c)case "variant_b": return h.alternateFlow(c)}
8.Low level design
Core types
Owns the in-memory cache, one per process, pre-warmed at construction. Takes the realtime hub optionally, so broadcasts are available without being required.
IsEnabledVariantRefresh30 seconds. Quick enough that an administrative change feels immediate, slow enough not to poll the database for no reason.
SHA-256 of the user identifier, a separator and the flag name, modulo 100. Both inputs matter: the first for stickiness, the second for independence between flags.
Feeds one writer. Each check used to start a goroutine and an insert of its own, which a busy page turned into thousands a second.
Design principles applied
- Evaluate in memory. a flag check has to be cheaper than the branch it guards, or nobody will put one on a hot path and the mechanism goes unused where it matters.
- Hash both inputs. the user for stickiness, the flag for independence. Each omission produces a different and confusing failure.
- Unknown means off. the default has to be the safe one, because the uncertain cases are a flag not yet created and a flag just deleted.
- One writer for a high-frequency side effect. the exposure write happens as often as the check. A goroutine per check is not a smaller version of a writer, it is a different system with no backpressure.
- A flag needs a description and an owner. flags outlive the people who add them, and a flag nobody can explain is never removed.
Patterns
| Pattern | Where it is used |
|---|---|
| Consistent hashing | stable bucket assignment per user and flag |
| Read-through cache with refresh | in-memory state on an interval |
| Fire and forget | exposures that never block a check |
| Single writer | high-frequency inserts batched through one path |
| Pub/sub invalidation | a change broadcast so clients refetch |
9.Scalability and performance
- Checks are free: a map read and a hash, so flag count and check frequency both scale without cost.
- Refresh load is one query per process per interval, which is independent of traffic entirely.
- Memory is tens of kilobytes, which is why holding every flag is obviously right rather than a decision.
- Exposure writes are the only part that scales with traffic, and batching through one writer is what keeps them from being the dominant write load in the application.
- The thirty second window means processes can briefly disagree. For a rollout that is harmless; for a kill switch it is the reason the immediate refresh and the broadcast exist.
- The real scaling problem is organisational: flags accumulate, and a project with two hundred of them has two hundred untested combinations.
10.Bottlenecks and improvements
What breaks first
- Flag debt. a flag left in after its rollout finished is a permanent branch, and a hundred of them is a codebase with no single defined behaviour.
- Propagation window. thirty seconds of disagreement between processes, which matters when the flag is being used to turn something off in a hurry.
- Exposure volume. the exposure table grows with checks rather than with actions, which makes it the fastest growing table in the system.
- Anonymous stickiness. with no user identifier the bucket is random per request, so an anonymous user sees a different arm on every page unless a stable identifier is passed.
- Untested combinations. twenty independent flags is a million combinations, and the suite tests one of them.
What to do about it
- Expire flags deliberately. a review date on every flag and a report of the overdue ones. Removing a flag is part of shipping the feature, not a separate project.
- Broadcast every change. immediate refresh plus a realtime event, with the interval as the fallback. The window then only applies when the broadcast is missed.
- Sample or aggregate exposures. one row per user per flag per day rather than per check. The measurement is the same and the volume is bounded by users rather than by traffic.
- Bucket anonymous users on a device identifier. a session or device value gives stickiness without a login, and the package should take it rather than silently being random.
- Test both arms of what matters. not every combination, but both sides of any flag guarding a path with its own tests. The combinatorial problem is a reason to have fewer flags, not more tests.
