All systems
Access

Data Isolation System Design

Making "your invoices" mean yours, in the query rather than in the handler, so a forgotten check is not a breach.

internal/authzinternal/services

1.Problem statement

Permissions decide what kind of thing a caller may touch. They cannot decide which rows. A permission model alone gives every user with invoices.view the ability to read every invoice in the system, which is the single most common serious bug in web applications and has its own name: broken object level authorization.

The reason it is so common is that the correct check is easy to write and easy to forget. Every list endpoint must narrow, every read by id must verify, every update and delete must verify, and the failure is silent: the endpoint works, the tests pass, and the data is wrong only from somebody else’s point of view.

The fix is to move the check off the handler and into the layer that builds the query, so that a developer has to go out of their way to write an unscoped read rather than having to remember to scope one.

The system has to be able to:

  • Narrow a list query to the caller’s rows without the handler saying anything.
  • Refuse a read, update or delete of a row the caller does not own.
  • Answer 404 rather than 403 for a row they may not see, so the response does not confirm it exists.
  • Let an administrator see everything, because the operator console reads the same endpoints.
  • Match nothing at all when the caller is unknown, rather than opening up.
  • Apply in a background job and a command, not only in a request.

2.System requirements

Functional requirements

  • A --owned-by flag on resource generation that adds the owner column and the scoping.
  • Scoping applied inside the service, in the function that builds the query.
  • The owner stamped from the caller on create, never read from the request body.
  • The owner column not writable through update or patch.
  • An administrator exempt from the narrowing, on the same endpoints.
  • A guessed id answering 404, not 403.

Non-functional requirements

  • Default deny. a context with no actor matches no rows. A scope that opened up when it could not tell who was asking would read as protection and provide none.
  • In the service, not the handler. the handler is one caller of the service. A job, a command and a webhook are others, and all of them get the same narrowing.
  • Not leaky in errors. the difference between "does not exist" and "not yours" is not observable, because 403 on the second is itself a disclosure.
  • Not forgeable. the owner comes from the authenticated identity. A body field with the same name is ignored.
  • Testable as a breach. the property is "Ada cannot see Basil’s row", which is a test that fails loudly when the scoping is removed.

3.Capacity estimation

Numbers for a mid-sized deployment. They are here to size the thing, not to predict your traffic: change an assumption and the sums below move with it.

Assumptions

ParameterValue
Owned resources in a typical project3 to 10
Rows per owned tableup to tens of millions
Rows per usertens to thousands
List queriesthe majority of authenticated reads

Query selectivity

10,000,000 invoices across 200,000 users
= ~50 rows per user
an index on (user_id, created_at) turns a full scan into ~50 rows read

The scoping is not only correctness, it is the difference between a sorted scan of ten million rows and an index read of fifty. The owner column is indexed for that reason.

Cost of the check on a read by id

one row loaded, one comparison against the actor
no extra query: the ownership column is already on the row

4.High level design

Two primitives do the whole job. One narrows a query, the other judges a loaded row. Everything else is making sure they are the only way through.

Core components

  • ScopeOwned. takes a query and a column and returns a query narrowed to the actor. An administrator gets it back untouched; an unknown caller gets a query that matches nothing.
  • Owns. takes a loaded row and says whether the actor may see it. Used on read by id, update and delete.
  • Generated service. where both are called. The generator emits the calls, so an owned resource is scoped the moment it exists rather than when somebody remembers.
  • Create-time stamping. the owner is set from the context in the create path and ignored in the request, so it cannot be forged.
  • Isolation tests. generated alongside, asserting that one user cannot reach another’s rows and that the operator still can.

Request flow

A list and a read by id, both narrowed

1234567ClientHandlerServiceScopeOwnedDatabaseOwns404 Not FoundRows Returned
  1. 1The request arrives already authenticated and already past the permission check. Permission said "may read invoices"; nothing yet has said which.
  2. 2The handler calls the service. It passes the context and no ownership argument, because it does not know about ownership.
  3. 3The service builds the query and passes it through ScopeOwned, which adds the predicate. An administrator gets the query unchanged; a context with no actor gets one that matches nothing.
  4. 4The narrowed query reaches the database. The owner column is indexed, so this is also the difference between reading fifty rows and scanning ten million.
  5. 5On a read by id the row comes back first and is judged by Owns, because the primary key lookup does not carry the predicate.
  6. 6A row the actor does not own is reported as not found. Answering 403 would confirm that an invoice with that id exists, which is the disclosure the check exists to prevent.
  7. 7Otherwise the rows are returned. The handler never saw the decision, which is what keeps it from being forgotten in a new endpoint.

Data flow

  • The owner column is written once, from the identity, at creation. It is excluded from the update and patch paths, so it cannot be reassigned by a request.
  • The column name reaching the query builder is a fixed string from the generator, never anything derived from input.
  • Administrators are exempt by the same primitive, so the operator console uses the same endpoints as the customer and there is no second code path to keep in step.
  • A background job that acts for a user carries the same actor on its context, so the narrowing applies there too.

5.Technology stack

ComponentWhat it is
Scoping primitivea query modifier applied in the service
Ownership columna foreign key to users, indexed
Identity sourcethe actor on context.Context
Exemptionthe administrator flag on the actor
Failure modematch nothing when the caller is unknown
Verificationgenerated isolation tests over HTTP

6.Data model

There is no table for this system. It is a column on the resources that have an owner, plus the discipline about where it is read and written.

any owned resource

The index matters as much as the column. Without (user_id, created_at) a scoped list is a filtered sort of the whole table.

ColumnHolds
user_idthe owner, indexed, set at creation from the identity
...the resource’s own columns

7.Low level design

Core types

ScopeOwnedinternal/authz/actor.go

The query narrowing. Three outcomes: untouched for an administrator, impossible for an unknown caller, narrowed for everybody else. The order of those cases is the security property.

Ownsinternal/authz/actor.go

The row judgement, for the paths where a primary key lookup has already happened.

Ownable

A one-method interface: GetOwnerID. A model satisfies it by having an owner, which is how the generator knows the resource is scoped.

OwnsOr404internal/security/...

The handler-side convenience that converts a failed ownership check into a not-found rather than a forbidden.

Design principles applied

  • Make the safe path the easy path. the generator writes the scoping. A developer has to delete a line to produce an unscoped read, which is a visible act in review rather than an omission.
  • Fail closed, loudly. an unknown actor produces a query that matches nothing. The endpoint returns an empty list, which is noticed, rather than everything, which is not.
  • Enforce where every caller passes. the service is the narrowest point that all callers share. The handler is not, because jobs and commands do not go through it.
  • Do not leak through status codes. the distinction between forbidden and missing is itself information, so it is not offered.

Patterns

PatternWhere it is used
Query scopea predicate added by a shared function rather than by each call site
Guard clauseownership judged before the row is returned or mutated
Ambient contextthe identity travels on the context, so background work inherits it

8.Scalability and performance

  • The predicate makes queries more selective, not less. On a large table the scoped version is dramatically faster than the unscoped one it replaces.
  • The composite index on the owner column and the sort column is what keeps a scoped list from degrading into a sort of the whole table.
  • Nothing is cached, because a cache keyed without the owner is exactly the bug this system prevents.
  • The check costs no additional round trip. The ownership column arrives with the row that was being read anyway.
  • As a table grows, scoping becomes more valuable rather than less: the work per request stays proportional to one user’s data.

9.Bottlenecks and improvements

What breaks first

  • A hand-written endpoint that forgets. the generator scopes what it generates. A custom endpoint added later is where this breaks, and it breaks silently.
  • Aggregates and reports. a count or a sum written as raw SQL bypasses the service and the scope with it, and the number it returns is the whole table.
  • Caching by id. a cache key that does not include the owner will serve one user’s row to another, turning a correct system into an incorrect one.
  • Relationships that cross the fence. an owned invoice pointing at a customer that is not owned means a user can enumerate customers through the relationship.
  • Missing index. the scope is still correct without an index on the owner column, and still slow. Correctness hides the performance problem until the table is large.

What to do about it

  • Isolation tests that would catch a breach. tests asserting that one user cannot see another’s rows, verified to fail when the scoping line is deleted. A test that passes either way proves nothing.
  • Route the aggregates through the service. counts and sums take the same scoped query builder, so a report cannot see further than a list.
  • Owner in the cache key. every cache entry for an owned resource is keyed by owner as well as id, which makes the cross-user case impossible rather than unlikely.
  • Index with the sort. a composite index on owner and the default sort column, created with the resource rather than added after the first slow query.

Read next