File Storage System Design
Accepting an upload, keeping it somewhere that survives a redeploy, and serving it back to the people allowed to see it.
internal/storageinternal/filesinternal/media1.Problem statement
A user uploads an avatar. Writing it to the container’s filesystem works until the container restarts, and then the avatar is gone and so is every other upload. Writing it to the database as bytes works until the table is sixty gigabytes and every backup takes an hour. Neither is a storage design; both are what happens when storage is not designed.
Object storage is the answer, and it brings its own set of decisions that are easy to get wrong in ways that only show up later. A browser that uploads through the API means every file passes through a request, holding a connection for its duration, bounded by the request body limit and the proxy timeout. A file served through the API means every read costs an application request. A file served directly from a public bucket means access control has been given up entirely, which is fine for an avatar and not for an invoice.
And there is the local case. A small deployment should not need a bucket. It should be able to write files to a directory and serve them, with the same code, so that a single binary and SQLite remains a real option rather than a demo.
That last part was broken here for a long time in an interesting way: the local driver worked, and the fallback that was supposed to select it when no bucket was configured never fired, so a project with no MinIO got errors instead of a working directory.
The system has to be able to:
- Store a file through one interface, whichever backend is configured.
- Support S3-compatible buckets and a local directory with the same code.
- Accept an upload through the API, with a size limit and a type check.
- Offer a pre-signed URL so a browser can upload straight to the bucket.
- Serve a private file through an authorised route, and a public one directly.
- Derive thumbnails and sized variants without blocking the upload.
- Record every file as a row, so it can be listed, attributed and deleted.
- Delete the object when the row goes, and not before.
2.System requirements
Functional requirements
- A Disk interface: put, get, delete, exists, URL.
- Drivers for S3, MinIO, R2 and B2, and a local directory driver.
- Driver selection from STORAGE_DRIVER, with local as a working default.
- Multipart upload handling with a configured size limit.
- Content type validation against an allowlist, by sniffed type and not only by extension.
- Pre-signed PUT URLs valid for an hour, with a clear error from drivers that cannot.
- An uploads table: key, size, type, owner, created time.
- Image processing in a background job, producing sized variants.
- An admin files page listing, previewing and deleting.
Non-functional requirements
- Survives a redeploy. the container filesystem is scratch space. Anything that must outlive a deploy goes to a bucket or to a mounted volume, and that is a decision made once rather than per feature.
- Local must actually work. a project with no bucket configured writes to a directory and serves from the API. Not as a degraded mode, as a supported deployment.
- Private by default. a bucket is not public unless a file is meant to be. The default must be the safe one, because the unsafe one is invisible until it is indexed.
- Bounded uploads. a size limit and a type allowlist, enforced in the handler. Without them an upload endpoint is a way to fill a disk.
- Derived work is asynchronous. thumbnailing happens in a job. An upload that waits for image processing is an upload that times out on a large file.
- The row is the record. a file is a row that references an object. Orphaned objects and orphaned rows are both bugs, and the row is what makes either detectable.
3.Capacity estimation
Numbers for a mid-sized deployment. They are here to size the thing, not to predict your traffic: change an assumption and the sums below move with it.
Assumptions
| Parameter | Value |
|---|---|
| Uploads per day | ~12,000 |
| Average upload size | 1.8 MB |
| Reads per upload over its life | ~40 |
| Thumbnail variants per image | 3 |
| Retention | indefinite |
Storage growth
Which is the argument for lifecycle rules rather than for deleting things. Moving year-old objects to infrequent access costs nothing in code and most of the bill.
Upload bandwidth through the API
This is what pre-signed URLs remove. The bytes go browser to bucket, and the API handles only the authorisation and the row.
Read bandwidth
Egress is usually the largest line on an object storage bill, and a CDN in front of it is the only thing that changes the number materially.
Thumbnail work
4.High level design
One interface, several drivers, a row per file, and the derived work in a job.
Core components
- Disk interface. what every driver implements. New code takes a Disk rather than a concrete store, which is what makes the local case first class instead of special.
- Drivers. S3-compatible for s3, minio, r2 and b2, and local for a directory on this machine served from the API.
- Storage service. holds one Disk, chosen by configuration, and reports which driver it is. The report matters: "local disk at ./storage" in a health response answers a question that otherwise takes an hour.
- Upload handler. multipart, with a size limit and a type allowlist. Writes the object and the row in that order, so a failed write never leaves a row pointing at nothing.
- Pre-signed URLs. a PUT URL valid for an hour so the browser uploads directly. Drivers that cannot say so with a named error rather than failing obscurely.
- Image processing. a job producing sized variants. Asynchronous because a large image takes seconds and an upload should not.
- Uploads table. the row per file: key, size, type, owner. What makes a file listable, attributable and deletable.
Request flow
An upload, two ways, and the read path
- 1A browser uploads. For a small file that goes through the API, which is simple and costs a held request for the duration of the transfer.
- 2The handler checks the size against the configured limit and the content type against an allowlist, by sniffing the bytes rather than trusting the extension, because the extension is supplied by the uploader.
- 3The object is written to the configured Disk: a bucket, or a directory on this machine. The same code either way, which is what makes a single-binary deployment a real option.
- 4The alternative path, for large files: the API returns a pre-signed PUT valid for an hour and the browser uploads straight to the bucket. The API authorises and records; it never carries the bytes. Drivers that cannot pre-sign return a named error rather than failing in some other way.
- 5A row is written after the object exists, never before. A row pointing at an object that was never written is a broken link; an object with no row is garbage that a sweep can find.
- 6Derived work is queued. Three thumbnail sizes take a couple of seconds, which is fine in a worker and not in a request.
- 7A read is authorised against the row. This is where private and public diverge: a private file is checked and streamed or given a short-lived signed URL, a public one is served directly.
- 8Public files go through a CDN, because egress is the largest part of an object storage bill and caching is the only thing that changes it. Private files are served by the API or by signed URLs with a short life.
Data flow
- Object first, row second. The reverse ordering produces rows pointing at nothing, which is the harder of the two failure modes to detect.
- Deleting removes the row and the object, in that order, so a failure leaves an orphaned object rather than a broken reference.
- Keys are generated, not taken from the upload. A user-supplied filename is a path traversal attempt waiting to happen, and it is kept as a display name only.
- The driver name is reported in health, because "which store is this project using" is otherwise answered by reading configuration on a machine nobody has access to.
5.Technology stack
| Component | What it is |
|---|---|
| Interface | one Disk, several drivers |
| Object stores | S3, MinIO, Cloudflare R2, Backblaze B2 |
| Local | a directory, served from the API, a supported deployment |
| Selection | STORAGE_DRIVER |
| Direct upload | pre-signed PUT, one hour |
| Processing | sized variants in a background job |
| Record | an uploads table |
6.Data model
uploads
original_name is display only and is never part of a path. Using it as a key is how a path traversal gets in.
| Column | Holds |
|---|---|
| id | primary key |
| key | the object key, generated, never user supplied |
| original_name | what the uploader called it, for display only |
| content_type | as sniffed, not as claimed |
| size | bytes |
| user_id | who uploaded it, for ownership scoping |
| variants | JSON: the derived sizes and their keys |
| created_at | when |
Where it lives
- Objects live in the bucket or the directory; only metadata is in the database. Bytes in a database make every backup the size of every upload.
- Lifecycle rules on the bucket move old objects to cheaper classes, which is a configuration change rather than a code one.
- The local driver’s directory must be a mounted volume in a container, or it is scratch space with a longer name.
7.API design
Uploads
| Method | Endpoint | What it does |
|---|---|---|
| POST | /api/v1/uploads | Multipart upload through the API |
| POST | /api/v1/uploads/presign | A pre-signed PUT for a direct browser upload |
| GET | /api/v1/uploads | List, scoped to the caller |
| GET | /api/v1/uploads/:id | Authorised read of a private file |
| DELETE | /api/v1/uploads/:id | Remove the row and the object |
8.Low level design
Core types
The interface every driver implements. New code takes a Disk, which is what keeps the local case from being a special case.
PutGetDeleteExistsURLKeeps files in a directory and serves them from the API. Always worked; the fallback that selects it when no bucket is configured did not fire until v3.370.0.
Holds one Disk chosen by STORAGE_DRIVER and reports which. The report is in the health response, which is where somebody finds out what a deployment is actually doing.
DriverPresignPutURLReturned by drivers that cannot pre-sign, such as local. A named error so the caller can fall back to uploading through the API rather than guessing.
Design principles applied
- One interface, no special cases. the local driver implements the same interface as S3. The moment local is a branch rather than a driver, it stops being tested.
- Generate the key. a user-supplied filename in a path is a traversal. Keep it as a display name and generate the key.
- Object before row. the ordering decides which failure you get. An orphaned object is findable; a row pointing at nothing is a broken page.
- Report the driver. the configuration is on a machine nobody can read. The health response is where the answer belongs.
- Name what a driver cannot do. a specific error for an unsupported pre-sign lets the caller fall back. A generic failure makes it look like a bug.
Patterns
| Pattern | Where it is used |
|---|---|
| Strategy | the Disk interface with a driver per store |
| Pre-signed URL | the bytes bypass the application entirely |
| Metadata and blob split | rows in the database, objects in the store |
| Asynchronous derivation | variants built in a job after the upload |
9.Scalability and performance
- Object storage scales without the application participating, which is most of why it is the answer.
- Pre-signed uploads remove the API from the data path, so upload capacity stops being an application concern at all.
- Egress is the dominant cost at volume, and a CDN in front of public files is the only change that moves it significantly.
- Private files served through the API cost a request per read. Short-lived signed URLs give most of the benefit of direct serving while keeping authorisation.
- Thumbnail work scales with the worker pool and is cheap. Generating variants on demand and caching them is the alternative, and it trades storage for latency on first view.
- The local driver scales to one machine and its disk, which is the correct limit for the deployment it exists to serve.
10.Bottlenecks and improvements
What breaks first
- Uploads through the API. every upload holds a request, bounded by the body limit and the proxy timeout. A large file fails in a way that looks like a network problem.
- Egress cost. at hundreds of gigabytes a day, bandwidth is the bill, and nothing in the application code is where it is decided.
- Orphans in both directions. objects with no row accumulate silently from failed requests; rows with no object break pages.
- Public buckets. a bucket made public to make serving simple is a bucket whose entire contents are enumerable by anybody who guesses the naming scheme.
- Unbounded uploads. no size limit and no type check turns an upload endpoint into a way to fill a disk and host arbitrary content.
What to do about it
- Pre-sign everything large. the browser uploads to the bucket, the API authorises and records. Removes the body limit, the timeout and the bandwidth from the application.
- CDN the public prefix. cache at the edge. It is a configuration change and it is the only thing that materially changes the egress bill.
- Sweep for orphans. a scheduled job comparing keys to rows in both directions, reporting rather than deleting until the counts are understood.
- Keep the bucket private, sign the reads. short-lived signed URLs for private files. Authorisation stays in the application and the bytes still come from the store.
- Limit and sniff. a size cap and an allowlist checked against the sniffed type. The extension is supplied by the uploader and means nothing.
