All systems
Platform

Feature Flags System Design

Turning a feature on for ten per cent of users without a deploy, and having the same user stay in the same group.

internal/flagsinternal/models

1.Problem statement

A redesigned checkout is ready. Shipping it to everybody at once means that if it is wrong, the fix is a deploy, and the damage is everybody. Shipping it to ten per cent means the damage is ten per cent and the fix is a toggle.

That requires the decision to be made at runtime rather than at build time, which rules out a constant and a compile-time flag. It also requires the decision to be stable per user: somebody who sees the new checkout on Monday must see it on Tuesday, or they are not a user of either version, they are a user of a flickering application.

The naive implementation reads a row per check. A page rendering a dozen flagged components does a dozen queries, and the flag system becomes the slowest part of the request it was meant to make safe. The naive exposure tracking is worse: a goroutine and an insert per check turned a busy page into thousands of writes a second.

And a flag is only half useful without knowing who saw what. A ten per cent rollout with no record of which users were in it cannot be measured, so the experiment produces a feeling rather than a result.

The system has to be able to:

  • Turn a feature on or off at runtime, without a deploy.
  • Roll out to a percentage of users, stably per user.
  • Serve several variants, for comparing two new versions rather than one.
  • Check a flag without touching the database.
  • Propagate a change in seconds, not on the next deploy.
  • Record exposures, so a rollout can be measured.
  • Target specific users, roles or tenants, not only percentages.
  • Tell clients a flag changed, so they refetch rather than waiting.

2.System requirements

Functional requirements

  • A flags table with a key, a state, rules and variants.
  • An engine holding the flags in memory, one per process.
  • A background refresh on an interval, 30 seconds by default.
  • An immediate refresh and a realtime broadcast on an administrative change.
  • Percentage rollout bucketed per user and flag.
  • Variants, for more than two arms.
  • Exposure records, written by one writer rather than per check.
  • An admin page to create, edit, roll out and delete.
  • A client-side hook reading the evaluated flags for the current user.

Non-functional requirements

  • A flag check never touches the database. all flags are in memory. A page with a dozen checks must not be a page with a dozen queries, or the mechanism costs more than the feature it guards.
  • Sticky per user and flag. the bucket is a hash of the user identifier and the flag name, so a user always lands in the same bucket for a given flag and their assignment does not flicker between sessions.
  • Independent per flag. including the flag name in the hash means the ten per cent for one flag is not the same ten per cent as another. Otherwise one unlucky cohort gets every experiment.
  • Propagates in seconds. 30 seconds from the refresh, immediately on an administrative change. A flag that needs a deploy to take effect is a constant.
  • Exposure tracking never blocks. fire and forget to a single writer. A goroutine and an insert per check was thousands of writes a second on a busy page.
  • Fails to a known state. an unknown flag is off. The default has to be the safe one, because the failure case is a flag that was never created or was just deleted.

3.Capacity estimation

Numbers for a mid-sized deployment. They are here to size the thing, not to predict your traffic: change an assumption and the sums below move with it.

Assumptions

ParameterValue
Flags in a mature project20 to 100
Flag checks per request1 to 12
Requests per second926 average, 4,630 peak
Refresh interval30 seconds
Flag row size~400 bytes

Check cost

a read lock on a map, plus one SHA-256 over ~40 bytes
= sub-microsecond
926 req/s x 12 checks = ~11,000 checks/s, unmeasurable

Against a database read per check at ~1 ms, 11,000 checks a second would be eleven seconds of database time per second. That is the difference the in-memory cache makes.

Refresh load

1 query every 30 s per process
x 12 processes = 0.4 queries/s
returning ~100 rows of ~400 bytes

Memory

100 flags x ~400 bytes = 40 KB per process

Which is why holding all of them is obviously right rather than a trade-off worth analysing.

Exposure writes

naive: one goroutine and one INSERT per check = 11,000/s
batched through one writer: ~10 inserts/s of 1,000 rows

The naive version is the thing that broke. A single writer with batching is three orders of magnitude fewer statements for the same information.

4.High level design

One engine per process holding every flag in memory, refreshed on a timer and on change, with bucketing by hash.

Core components

  • Flags table. the definition: a key, whether it is on, its rules, its rollout percentage and its variants.
  • Engine. the in-memory map, one per process, pre-warmed at boot behind a read-write lock. A flag check reads the map and never the database.
  • Refresher. a background loop pulling fresh state every 30 seconds, which is quick enough for an administrative change to feel immediate without polling hard.
  • Bucketer. SHA-256 of the user identifier and the flag name, modulo 100. Sticky per pair, and independent across flags because the name is in the hash.
  • Variant selector. the same bucket mapped onto several arms, for comparing two new versions rather than one against the old.
  • Exposure writer. one writer fed by a channel. Each check used to start a goroutine and an insert of its own, which a busy page turned into thousands a second.
  • Change broadcast. an administrative change refreshes the engine immediately and publishes a realtime event, so subscribed clients refetch instead of waiting out the interval.

Request flow

A flag check, and a change propagating

12345678Requestflags.IsEnabledIn-Memory MapHash BucketVariantExposure WriterAdmin ChangeRefresh + BroadcastClients Refetch
  1. 1A request renders a page with a dozen flagged components, each asking whether its feature is on for this user.
  2. 2Each check reads the in-memory map under a read lock. No database round trip, which is the property that lets a page have twelve checks instead of one.
  3. 3For a percentage rollout, the bucket is SHA-256 of the user identifier and the flag name, modulo 100. The user identifier makes it sticky, so Monday’s assignment is Tuesday’s. The flag name makes flags independent, so the unlucky ten per cent of one experiment is not the unlucky ten per cent of all of them.
  4. 4With variants, the same bucket maps onto several arms rather than on and off, which is how two new versions get compared against each other instead of only against the old one.
  5. 5The exposure is recorded, fire and forget, through one writer fed by a channel. Every check starting its own goroutine and its own insert is what made a busy page thousands of writes a second.
  6. 6Separately, somebody changes a flag in the admin: enables it, moves it to twenty-five per cent, adds a variant.
  7. 7The engine refreshes immediately rather than waiting out the interval, and the timer continues as the fallback that keeps every other process in step within thirty seconds.
  8. 8A realtime event goes out, so clients holding evaluated flags refetch rather than showing a stale answer until their next navigation. An anonymous user buckets on a random per-request value, which is effectively random, so anything that needs stickiness without a login passes a session or device identifier instead.

Data flow

  • The user identifier and the flag name are both in the hash. Omitting the first makes assignment flicker; omitting the second makes one cohort receive every experiment.
  • An unknown flag is off. That covers the window after a deploy that references a flag nobody created yet, and the window after a deletion.
  • Exposures are a stream to one writer, not a write per check. The information is the same and the statement count is three orders of magnitude lower.
  • An anonymous check is random per request unless the caller supplies a stable identifier, and the package says so rather than pretending otherwise.

5.Technology stack

ComponentWhat it is
Storagea flags table and an exposures table
Evaluationin memory, one engine per process
Refresh30 seconds, plus immediate on change
BucketingSHA-256 of user id and flag name, modulo 100
Variantsseveral arms over the same bucket
Exposurefire and forget to a single batched writer
Propagationa realtime flag.updated event

6.Data model

feature_flags

ColumnHolds
keythe name used in code
enabledthe master switch
rulesJSON: rollout percentage, targeted users, roles, tenants
variantsJSON: the arms and their weights
descriptionwhat it is for, because a flag outlives whoever added it

flag_exposures

This is what turns a rollout into a measurement. Without it, a ten per cent experiment produces a feeling rather than a result.

ColumnHolds
flag_keywhich flag
user_idwho saw it
variantwhich arm they got
created_atwhen

7.API design

Flags

MethodEndpointWhat it does
GET/api/v1/flagsEvaluated flags for the current caller
GET/api/v1/admin/flagsAll flags with their rules
POST/api/v1/admin/flagsCreate one
PUT/api/v1/admin/flags/:keyChange state, rollout or variants
DELETE/api/v1/admin/flags/:keyRemove it

Two arms, or several

if flags.IsEnabled(c, "new_checkout") {
return h.newCheckout(c)
}
switch flags.Variant(c, "checkout_redesign") {
case "control": return h.oldFlow(c)
case "variant_a": return h.newFlow(c)
case "variant_b": return h.alternateFlow(c)
}

8.Low level design

Core types

flags.Engineinternal/flags/flags.go

Owns the in-memory cache, one per process, pre-warmed at construction. Takes the realtime hub optionally, so broadcasts are available without being required.

IsEnabledVariantRefresh
DefaultRefreshInterval

30 seconds. Quick enough that an administrative change feels immediate, slow enough not to poll the database for no reason.

Bucketing

SHA-256 of the user identifier, a separator and the flag name, modulo 100. Both inputs matter: the first for stickiness, the second for independence between flags.

Exposure channel

Feeds one writer. Each check used to start a goroutine and an insert of its own, which a busy page turned into thousands a second.

Design principles applied

  • Evaluate in memory. a flag check has to be cheaper than the branch it guards, or nobody will put one on a hot path and the mechanism goes unused where it matters.
  • Hash both inputs. the user for stickiness, the flag for independence. Each omission produces a different and confusing failure.
  • Unknown means off. the default has to be the safe one, because the uncertain cases are a flag not yet created and a flag just deleted.
  • One writer for a high-frequency side effect. the exposure write happens as often as the check. A goroutine per check is not a smaller version of a writer, it is a different system with no backpressure.
  • A flag needs a description and an owner. flags outlive the people who add them, and a flag nobody can explain is never removed.

Patterns

PatternWhere it is used
Consistent hashingstable bucket assignment per user and flag
Read-through cache with refreshin-memory state on an interval
Fire and forgetexposures that never block a check
Single writerhigh-frequency inserts batched through one path
Pub/sub invalidationa change broadcast so clients refetch

9.Scalability and performance

  • Checks are free: a map read and a hash, so flag count and check frequency both scale without cost.
  • Refresh load is one query per process per interval, which is independent of traffic entirely.
  • Memory is tens of kilobytes, which is why holding every flag is obviously right rather than a decision.
  • Exposure writes are the only part that scales with traffic, and batching through one writer is what keeps them from being the dominant write load in the application.
  • The thirty second window means processes can briefly disagree. For a rollout that is harmless; for a kill switch it is the reason the immediate refresh and the broadcast exist.
  • The real scaling problem is organisational: flags accumulate, and a project with two hundred of them has two hundred untested combinations.

10.Bottlenecks and improvements

What breaks first

  • Flag debt. a flag left in after its rollout finished is a permanent branch, and a hundred of them is a codebase with no single defined behaviour.
  • Propagation window. thirty seconds of disagreement between processes, which matters when the flag is being used to turn something off in a hurry.
  • Exposure volume. the exposure table grows with checks rather than with actions, which makes it the fastest growing table in the system.
  • Anonymous stickiness. with no user identifier the bucket is random per request, so an anonymous user sees a different arm on every page unless a stable identifier is passed.
  • Untested combinations. twenty independent flags is a million combinations, and the suite tests one of them.

What to do about it

  • Expire flags deliberately. a review date on every flag and a report of the overdue ones. Removing a flag is part of shipping the feature, not a separate project.
  • Broadcast every change. immediate refresh plus a realtime event, with the interval as the fallback. The window then only applies when the broadcast is missed.
  • Sample or aggregate exposures. one row per user per flag per day rather than per check. The measurement is the same and the volume is bounded by users rather than by traffic.
  • Bucket anonymous users on a device identifier. a session or device value gives stickiness without a login, and the package should take it rather than silently being random.
  • Test both arms of what matters. not every combination, but both sides of any flag guarding a path with its own tests. The combinatorial problem is a reason to have fewer flags, not more tests.

Read next