All systems
Operations

File Storage System Design

Accepting an upload, keeping it somewhere that survives a redeploy, and serving it back to the people allowed to see it.

internal/storageinternal/filesinternal/media

1.Problem statement

A user uploads an avatar. Writing it to the container’s filesystem works until the container restarts, and then the avatar is gone and so is every other upload. Writing it to the database as bytes works until the table is sixty gigabytes and every backup takes an hour. Neither is a storage design; both are what happens when storage is not designed.

Object storage is the answer, and it brings its own set of decisions that are easy to get wrong in ways that only show up later. A browser that uploads through the API means every file passes through a request, holding a connection for its duration, bounded by the request body limit and the proxy timeout. A file served through the API means every read costs an application request. A file served directly from a public bucket means access control has been given up entirely, which is fine for an avatar and not for an invoice.

And there is the local case. A small deployment should not need a bucket. It should be able to write files to a directory and serve them, with the same code, so that a single binary and SQLite remains a real option rather than a demo.

That last part was broken here for a long time in an interesting way: the local driver worked, and the fallback that was supposed to select it when no bucket was configured never fired, so a project with no MinIO got errors instead of a working directory.

The system has to be able to:

  • Store a file through one interface, whichever backend is configured.
  • Support S3-compatible buckets and a local directory with the same code.
  • Accept an upload through the API, with a size limit and a type check.
  • Offer a pre-signed URL so a browser can upload straight to the bucket.
  • Serve a private file through an authorised route, and a public one directly.
  • Derive thumbnails and sized variants without blocking the upload.
  • Record every file as a row, so it can be listed, attributed and deleted.
  • Delete the object when the row goes, and not before.

2.System requirements

Functional requirements

  • A Disk interface: put, get, delete, exists, URL.
  • Drivers for S3, MinIO, R2 and B2, and a local directory driver.
  • Driver selection from STORAGE_DRIVER, with local as a working default.
  • Multipart upload handling with a configured size limit.
  • Content type validation against an allowlist, by sniffed type and not only by extension.
  • Pre-signed PUT URLs valid for an hour, with a clear error from drivers that cannot.
  • An uploads table: key, size, type, owner, created time.
  • Image processing in a background job, producing sized variants.
  • An admin files page listing, previewing and deleting.

Non-functional requirements

  • Survives a redeploy. the container filesystem is scratch space. Anything that must outlive a deploy goes to a bucket or to a mounted volume, and that is a decision made once rather than per feature.
  • Local must actually work. a project with no bucket configured writes to a directory and serves from the API. Not as a degraded mode, as a supported deployment.
  • Private by default. a bucket is not public unless a file is meant to be. The default must be the safe one, because the unsafe one is invisible until it is indexed.
  • Bounded uploads. a size limit and a type allowlist, enforced in the handler. Without them an upload endpoint is a way to fill a disk.
  • Derived work is asynchronous. thumbnailing happens in a job. An upload that waits for image processing is an upload that times out on a large file.
  • The row is the record. a file is a row that references an object. Orphaned objects and orphaned rows are both bugs, and the row is what makes either detectable.

3.Capacity estimation

Numbers for a mid-sized deployment. They are here to size the thing, not to predict your traffic: change an assumption and the sums below move with it.

Assumptions

ParameterValue
Uploads per day~12,000
Average upload size1.8 MB
Reads per upload over its life~40
Thumbnail variants per image3
Retentionindefinite

Storage growth

12,000 x 1.8 MB = ~21 GB/day of originals
variants add ~15%: ~24 GB/day
x 365 = ~8.8 TB/year

Which is the argument for lifecycle rules rather than for deleting things. Moving year-old objects to infrequent access costs nothing in code and most of the bill.

Upload bandwidth through the API

12,000/day is ~0.14/s average, but bursts to ~5/s
5 x 1.8 MB = 9 MB/s through the API
each one holding a request for its duration

This is what pre-signed URLs remove. The bytes go browser to bucket, and the API handles only the authorisation and the row.

Read bandwidth

12,000 x 40 reads = 480,000 reads/day
x 1.8 MB = ~860 GB/day egress

Egress is usually the largest line on an object storage bill, and a CDN in front of it is the only thing that changes the number materially.

Thumbnail work

12,000 images x 3 variants = 36,000 operations/day
at ~400 ms each = 4 hours of CPU
spread across workers, a fraction of one core

4.High level design

One interface, several drivers, a row per file, and the derived work in a job.

Core components

  • Disk interface. what every driver implements. New code takes a Disk rather than a concrete store, which is what makes the local case first class instead of special.
  • Drivers. S3-compatible for s3, minio, r2 and b2, and local for a directory on this machine served from the API.
  • Storage service. holds one Disk, chosen by configuration, and reports which driver it is. The report matters: "local disk at ./storage" in a health response answers a question that otherwise takes an hour.
  • Upload handler. multipart, with a size limit and a type allowlist. Writes the object and the row in that order, so a failed write never leaves a row pointing at nothing.
  • Pre-signed URLs. a PUT URL valid for an hour so the browser uploads directly. Drivers that cannot say so with a named error rather than failing obscurely.
  • Image processing. a job producing sized variants. Asynchronous because a large image takes seconds and an upload should not.
  • Uploads table. the row per file: key, size, type, owner. What makes a file listable, attributable and deletable.

Request flow

An upload, two ways, and the read path

12345678BrowserUpload HandlerSize + Type CheckDiskPresigned PUTUpload RowThumbnail JobAuthorised ReadCDN or API
  1. 1A browser uploads. For a small file that goes through the API, which is simple and costs a held request for the duration of the transfer.
  2. 2The handler checks the size against the configured limit and the content type against an allowlist, by sniffing the bytes rather than trusting the extension, because the extension is supplied by the uploader.
  3. 3The object is written to the configured Disk: a bucket, or a directory on this machine. The same code either way, which is what makes a single-binary deployment a real option.
  4. 4The alternative path, for large files: the API returns a pre-signed PUT valid for an hour and the browser uploads straight to the bucket. The API authorises and records; it never carries the bytes. Drivers that cannot pre-sign return a named error rather than failing in some other way.
  5. 5A row is written after the object exists, never before. A row pointing at an object that was never written is a broken link; an object with no row is garbage that a sweep can find.
  6. 6Derived work is queued. Three thumbnail sizes take a couple of seconds, which is fine in a worker and not in a request.
  7. 7A read is authorised against the row. This is where private and public diverge: a private file is checked and streamed or given a short-lived signed URL, a public one is served directly.
  8. 8Public files go through a CDN, because egress is the largest part of an object storage bill and caching is the only thing that changes it. Private files are served by the API or by signed URLs with a short life.

Data flow

  • Object first, row second. The reverse ordering produces rows pointing at nothing, which is the harder of the two failure modes to detect.
  • Deleting removes the row and the object, in that order, so a failure leaves an orphaned object rather than a broken reference.
  • Keys are generated, not taken from the upload. A user-supplied filename is a path traversal attempt waiting to happen, and it is kept as a display name only.
  • The driver name is reported in health, because "which store is this project using" is otherwise answered by reading configuration on a machine nobody has access to.

5.Technology stack

ComponentWhat it is
Interfaceone Disk, several drivers
Object storesS3, MinIO, Cloudflare R2, Backblaze B2
Locala directory, served from the API, a supported deployment
SelectionSTORAGE_DRIVER
Direct uploadpre-signed PUT, one hour
Processingsized variants in a background job
Recordan uploads table

6.Data model

uploads

original_name is display only and is never part of a path. Using it as a key is how a path traversal gets in.

ColumnHolds
idprimary key
keythe object key, generated, never user supplied
original_namewhat the uploader called it, for display only
content_typeas sniffed, not as claimed
sizebytes
user_idwho uploaded it, for ownership scoping
variantsJSON: the derived sizes and their keys
created_atwhen

Where it lives

  • Objects live in the bucket or the directory; only metadata is in the database. Bytes in a database make every backup the size of every upload.
  • Lifecycle rules on the bucket move old objects to cheaper classes, which is a configuration change rather than a code one.
  • The local driver’s directory must be a mounted volume in a container, or it is scratch space with a longer name.

7.API design

Uploads

MethodEndpointWhat it does
POST/api/v1/uploadsMultipart upload through the API
POST/api/v1/uploads/presignA pre-signed PUT for a direct browser upload
GET/api/v1/uploadsList, scoped to the caller
GET/api/v1/uploads/:idAuthorised read of a private file
DELETE/api/v1/uploads/:idRemove the row and the object

8.Low level design

Core types

storage.Diskinternal/storage/disk.go

The interface every driver implements. New code takes a Disk, which is what keeps the local case from being a special case.

PutGetDeleteExistsURL
storage.NewLocalinternal/storage/local.go

Keeps files in a directory and serves them from the API. Always worked; the fallback that selects it when no bucket is configured did not fire until v3.370.0.

storage.Storage

Holds one Disk chosen by STORAGE_DRIVER and reports which. The report is in the health response, which is where somebody finds out what a deployment is actually doing.

DriverPresignPutURL
ErrPresignUnsupported

Returned by drivers that cannot pre-sign, such as local. A named error so the caller can fall back to uploading through the API rather than guessing.

Design principles applied

  • One interface, no special cases. the local driver implements the same interface as S3. The moment local is a branch rather than a driver, it stops being tested.
  • Generate the key. a user-supplied filename in a path is a traversal. Keep it as a display name and generate the key.
  • Object before row. the ordering decides which failure you get. An orphaned object is findable; a row pointing at nothing is a broken page.
  • Report the driver. the configuration is on a machine nobody can read. The health response is where the answer belongs.
  • Name what a driver cannot do. a specific error for an unsupported pre-sign lets the caller fall back. A generic failure makes it look like a bug.

Patterns

PatternWhere it is used
Strategythe Disk interface with a driver per store
Pre-signed URLthe bytes bypass the application entirely
Metadata and blob splitrows in the database, objects in the store
Asynchronous derivationvariants built in a job after the upload

9.Scalability and performance

  • Object storage scales without the application participating, which is most of why it is the answer.
  • Pre-signed uploads remove the API from the data path, so upload capacity stops being an application concern at all.
  • Egress is the dominant cost at volume, and a CDN in front of public files is the only change that moves it significantly.
  • Private files served through the API cost a request per read. Short-lived signed URLs give most of the benefit of direct serving while keeping authorisation.
  • Thumbnail work scales with the worker pool and is cheap. Generating variants on demand and caching them is the alternative, and it trades storage for latency on first view.
  • The local driver scales to one machine and its disk, which is the correct limit for the deployment it exists to serve.

10.Bottlenecks and improvements

What breaks first

  • Uploads through the API. every upload holds a request, bounded by the body limit and the proxy timeout. A large file fails in a way that looks like a network problem.
  • Egress cost. at hundreds of gigabytes a day, bandwidth is the bill, and nothing in the application code is where it is decided.
  • Orphans in both directions. objects with no row accumulate silently from failed requests; rows with no object break pages.
  • Public buckets. a bucket made public to make serving simple is a bucket whose entire contents are enumerable by anybody who guesses the naming scheme.
  • Unbounded uploads. no size limit and no type check turns an upload endpoint into a way to fill a disk and host arbitrary content.

What to do about it

  • Pre-sign everything large. the browser uploads to the bucket, the API authorises and records. Removes the body limit, the timeout and the bandwidth from the application.
  • CDN the public prefix. cache at the edge. It is a configuration change and it is the only thing that materially changes the egress bill.
  • Sweep for orphans. a scheduled job comparing keys to rows in both directions, reporting rather than deleting until the counts are understood.
  • Keep the bucket private, sign the reads. short-lived signed URLs for private files. Authorisation stays in the application and the bytes still come from the store.
  • Limit and sniff. a size cap and an allowlist checked against the sniffed type. The extension is supplied by the uploader and means nothing.

Read next