DEV Community

Cover image for Designing Dropbox: File Storage & Sync System Design
MANGESH MANDLIK
MANGESH MANDLIK

Posted on

Designing Dropbox: File Storage & Sync System Design

Uploading a file is straightforward. Keeping its bytes, metadata, permissions, history, and copies on several intermittently connected devices consistent is the real system-design problem.

complete architecture

Consider a document named report.docx. You edit it on your laptop, close the lid halfway through an upload, rename its folder on another device, and then open your phone after several days offline. Meanwhile, a collaborator may have edited the same document or lost access to it.

A file storage and synchronization system must answer more than “where do we store the bytes?” It must know which version is current, which changes each device has observed, what to do when histories diverge, and whether a user is still authorized to download a particular version.

We'll design a Dropbox- or Google Drive-like system from a simple upload API to a reliable sync architecture. This is a conceptual design, not a claim about either company's internal implementation.

1. Requirements and scale

The system supports uploading and downloading large files, resumable transfers, folders, rename and move, automatic synchronization across devices, offline edits, version history, deletion recovery, and viewer/editor sharing.

The priorities are durability, integrity, secure access, availability despite individual failures, reasonable synchronization freshness, and bandwidth efficiency. “Instant sync” and “never lose a file” are not defensible unconditional guarantees: network partitions, retention rules, disasters, and client availability impose real limits.

For a scale exercise, suppose we have 500 million registered users, 100 million daily active users, and 200 million uploads per day. With an illustrative average upload of 5 MB, that's approximately 1 PB of new uploaded data per day before compression, deduplication, replication, or retention effects. The mean arrival rate is about 2,315 uploads/second, but peak capacity must be sized above the mean. Assume individual files can reach 50 GB. These are invented design inputs, not published statistics for Dropbox or Google Drive.

Notice that upload requests, bytes ingested, metadata operations, download traffic, and device sync connections scale differently. We should not expect one service or database to solve them all.

2. Start with the simplest upload architecture

Client ── file bytes ──> Application server ──> Local disk
                             │
                             └───────────────> Metadata database
Enter fullscreen mode Exit fullscreen mode

start simple

This works for small workloads. But when a large upload is interrupted, a one-shot upload may have to restart. Application servers also spend bandwidth and connection capacity relaying bytes rather than handling application logic. Local disks complicate replication, capacity expansion, and recovery.

Proxying bytes is not inherently wrong: inspection, transformations, restricted networks, or specialized clients can justify it. For general large-file transfers at our assumed scale, however, separating the control plane from the byte-transfer path is a useful next step.

3. Separate file metadata from object storage

                       ┌──────────────────────┐
Client ── API calls ──> │ File / Metadata API  │ ──> Metadata DB
                       └──────────────────────┘      identity, folders,
                                  │                 ACLs, versions, state
                                  │ signed capability
                                  ▼
Client ── file bytes directly ──> Object storage
                                 immutable version objects
Enter fullscreen mode Exit fullscreen mode

A metadata record might contain file_id, owner_id, parent_folder_id, filename, current_version_id, lifecycle_state, and a concurrency token. A separate versions table stores version_id, file_id, parent_version_id, content checksum, size, and object key. The file's stable ID is independent of its path: renaming report.docx or moving it into another folder changes metadata, not the file's identity or necessarily its stored bytes.

Folders also have stable IDs and parent relationships. Listing a folder means querying its immediate children with pagination. A folder move can change the folder's parent pointer without rewriting every descendant, although subtree permissions, path caches, and asynchronous indexes may still require careful handling. If sibling names must be unique, enforce that invariant within the correct namespace, including the root-folder case.

Object storage holds large binary objects; the metadata store handles queryable relationships and conditional updates. This is a workload-driven decision, not a rule that databases can never store small blobs.

The metadata–object consistency gap

The database transaction that publishes a file version cannot ordinarily atomically include the object-storage upload. Therefore, use an explicit lifecycle:

CREATED → UPLOADING → VERIFYING → AVAILABLE
                ↘ FAILED / EXPIRED
AVAILABLE → DELETING → DELETED (subject to retention)
Enter fullscreen mode Exit fullscreen mode

A safe publication sequence is:

  1. Create a metadata record and an upload session in a non-available state.
  2. Upload bytes to a server-selected object key.
  3. Complete the object upload and verify the integrity evidence supported by the storage API.
  4. Atomically update the file's current-version pointer and record its sync change using a transaction or another durable, correctly ordered publication mechanism.
  5. Reconcile abandoned sessions and orphan objects asynchronously.

The fourth step matters. If a file becomes visible in metadata but its corresponding change event is lost, other devices may never discover it. A transactional outbox or a change log committed alongside metadata can bridge that boundary; publishing to an unrelated queue after the database commit without recovery is not enough.

If storage succeeds but the metadata commit fails, the object may be orphaned. A reconciler checks the session, object, checksum, and current metadata state before safely retrying publication or scheduling deletion after a grace period. If metadata says AVAILABLE but the object is missing, fail visibly and investigate or recover from a replica/backup; never serve an empty object as success.

For deeper treatment, see File Storage & Sync — Architecture Deep Dive.

4. Design the upload API and idempotency contracts

An illustrative API is:

Endpoint Purpose
POST /files Create file metadata; client supplies an idempotency key
POST /files/{id}/uploads Start or resume an upload session
POST /files/{id}/complete Complete and verify a session; safe to retry
GET /files/{id} Return metadata and current version
GET /files/{id}/download Authorize and issue a short-lived download capability
PATCH /files/{id} Rename/move with a version precondition
GET /sync/changes?cursor=... Fetch durable changes for a device

These are example contracts, not a mandated REST standard. A production implementation must also define deletion, restore, sharing, version retrieval, and session-status endpoints.

A retry of POST /files must not accidentally create two files. An idempotency key scoped to the caller and operation allows the server to return the same logical result. Completing an upload should similarly be retry-safe: a timeout may mean “completion succeeded but the response was lost.” Repeating that request must not publish a duplicate version.

Idempotency handles retries of the same operation. It does not solve two genuinely concurrent edits; those require version preconditions.

5. Direct uploads and presigned URL security

For the normal large-file path, the application authenticates the caller, checks permission and quota, creates the session, and returns a short-lived presigned upload URL. The client sends bytes directly to object storage.

A presigned URL is generally a bearer capability. Anyone who obtains it can potentially use it until expiration, subject to signature and storage-policy constraints. It is not automatically bound to the original user's identity. Therefore:

  • Authorize before issuing it; generate object keys on the server.
  • Scope it to a specific operation and object or multipart part.
  • Use HTTPS, short expirations, appropriate request constraints where supported, and avoid leaking URLs in logs.
  • Configure browser CORS separately; CORS is not a replacement for authorization.
  • Recheck authorization and the expected version before publishing a completed upload, because permissions or file state may have changed while bytes were in flight.

Revoking a user's ACL does not automatically revoke a previously issued, unexpired URL. If near-immediate revocation is mandatory, an edge or proxy path that validates current authorization on access may be necessary. Even then, bytes already downloaded cannot be recalled.

6. Resumable multipart uploads

A single 10 GB PUT still has an expensive failure mode: a dropped connection can force a complete restart. Native object-storage multipart upload lets the client send separately acknowledged parts.

Upload session U17
  part 1 ──────── ✓
  part 2 ──────── ✓
  part 3 ── X     retry part 3
  part 4 ──────── ✓
                ↓
  complete known parts → verify → publish new version
Enter fullscreen mode Exit fullscreen mode

A durable session tracks its owner, file ID, expected base version, storage upload ID, expiration, and status. The client can upload parts concurrently and, after reconnecting, reconcile confirmed parts against the storage provider's actual state where supported. Its own local progress counter may be stale if a part succeeded but the response was lost.

Part size is a trade-off: smaller parts reduce failed-retry cost but increase request and tracking overhead; larger parts reduce coordination but increase retry cost and buffering. Provider minimums, maximum part counts, file size, client memory, and network behavior determine the right choice. Don't treat a single part size as universal across providers.

An object-store multipart part is not necessarily the same abstraction as an application-level delta-sync chunk. Native multipart simplifies upload assembly; a content-addressed chunk store is a separate, substantially more complex design.

Integrity is not the same as upload completion

TLS protects bytes in transit, but it does not prove that the application selected the correct file. Where supported, validate per-part checksums and the assembled object's expected checksum. A mismatch in the final whole-object hash shows that something differs; it does not identify the corrupted part without finer-grained evidence.

Do not assume an object-storage ETag is always an MD5 digest of the entire file. ETag semantics vary by provider and upload/encryption method, especially for multipart objects. Use explicit checksum mechanisms and clearly specify what has been verified before publication.

Expired sessions and abandoned multipart uploads require cleanup. Unassembled parts may still consume storage and cost money; expiry policies and periodic reconciliation are part of the design, not optional housekeeping.

7. Downloads, range requests, and private CDN delivery

For downloads, the file API checks the current ACL and returns a short-lived signed capability for a specific immutable file version. The client downloads bytes from storage, or through an appropriately secured CDN where caching is useful.

HTTP Range requests can resume an interrupted download or retrieve part of a large file. A robust client should ensure the resumed ranges belong to the same immutable object version and verify the completed file's expected checksum. Simply appending bytes from a mutable URL can mix different versions.

CDN caching is workload-dependent. Frequently requested, immutable shared objects may benefit; private files downloaded once per user may not. A private CDN needs its own edge authorization and private-origin configuration. A presigned origin-storage URL does not magically make a CDN path private or authorized. Already-issued CDN capabilities have similar revocation windows to signed origin URLs.

8. Multi-device sync is a reconciliation protocol

A sync client needs persistent local state, not merely a filesystem watcher. For each file it tracks the stable remote ID, local path, last synchronized server version, local fingerprint, pending operations, and its durable checkpoint.

Laptop: edit report.docx
  local state: LOCAL_DIRTY
           ↓
  upload candidate version based on v5
           ↓
  server checks current version == v5
           ↓
  publish v6 + durable change record
           ↓
  push hint to other online devices
           ↓
Phone: GET /sync/changes?cursor=C
           ↓
  apply v6, then advance local cursor
Enter fullscreen mode Exit fullscreen mode

Local filesystem notifications can be missed, so a periodic reconciliation scan is a useful safety net. Server push via WebSocket, mobile push, or long polling is only a wake-up hint. The durable, cursor-based change log is the source of truth.

A cursor identifies a position in an ordered change stream within its defined scope. It is not simply the client's wall-clock timestamp. The client must not persist a cursor beyond changes it has durably applied; otherwise a crash can permanently skip an update. Applying duplicate events must be safe through operation IDs, version checks, and idempotent local updates.

One subtlety: when changes are paginated or generated concurrently, the server must define a consistent ordering/checkpoint contract so clients cannot skip a late-committed entry with an earlier cursor position. A snapshot-plus-change-stream handoff likewise needs a defined boundary; “take a snapshot, then start polling” without such a boundary can miss mutations.

If a device has been offline longer than change-log retention, its old cursor may be invalid. It must obtain a consistent namespace snapshot, compare it with local state, reconcile offline operations, and establish a fresh cursor without losing changes during the transition.

9. Whole-file hashing is not delta synchronization

A content hash answers: are these bytes equal? It does not answer: which blocks changed?

There are three practical strategies:

Strategy Works well when Main limitation
Whole-file replacement Files are small or binary deltas are ineffective Transfers the entire new file
Fixed-size block comparison Large files change in place without shifting offsets Insertions shift later block boundaries
Rolling-signature or content-defined chunking Large files have localized edits and reusable regions More CPU, signatures, chunk metadata, and GC complexity

An rsync-style rolling matching algorithm and content-defined chunking are related approaches, not identical algorithms. Both can locate reusable regions despite offset shifts, but their chunking and matching procedures differ.

For compressed archives, encrypted data, and some document formats, a tiny logical edit can change much of the serialized byte stream. In those cases, a delta algorithm may save little bandwidth. The client should choose based on measured transfer savings versus computation and coordination cost, not promise that every one-paragraph edit uploads only a few kilobytes.

10. Detect conflicts using version lineage

Suppose both laptop and phone begin with report.docx at version v5. The laptop edits and successfully publishes v6. The phone was offline and independently edited its copy based on v5.

             v5 (shared base)
             /            \
   laptop candidate     phone candidate
        accepted v6       base=v5
             │              │
       current=v6      compare-and-set fails
                            │
                      preserve as conflict
Enter fullscreen mode Exit fullscreen mode

The server accepts a new current version only when its expected base version matches the current version. A stale writer must not silently overwrite the newer version. This check must be atomic at the metadata commit boundary.

Possible policies include preserving a conflict copy, last-write-wins, or a three-way merge for supported text formats. Conflict copies reduce silent overwrites but require user intervention and do not guarantee zero data loss if the underlying storage or backup system fails. Last-write-wins is simple but can discard edits. Automatic merging is format-dependent; it is not a general solution for arbitrary binary files. Real-time collaborative editing with OT or CRDTs is a different problem from background file synchronization.

Metadata edits also race. Rename and move operate on stable IDs; delete should produce a tombstone that offline devices can observe. When a deleted file receives a stale offline edit, the product must explicitly choose to reject it, restore a conflict copy, or create a recovered file. A naive sync client that sees a missing file and uploads its old local copy risks resurrecting a deletion.

Content version checks and metadata-operation concurrency tokens may need separate treatment: a rename need not create a new content blob, yet it must still be protected against conflicting namespace mutations.

11. Sharing and access control

The metadata layer records ownership, direct shares, and possibly inherited folder permissions. Each operation checks whether the caller may read, edit, rename, delete, share, or restore the target. Folder inheritance and override rules must be defined rather than guessed.

For a file shared by Alice with Bob, Bob requests a download through the application, which checks Bob's access and issues a scoped signed URL. Direct object-store access after that check is intentional. The file bytes need not pass through the authorization service.

A recipient-oriented index for “shared with me” avoids scanning every owner's ACLs as metadata is sharded. That index must be maintained carefully, but an eventually updated discovery index must not become the sole authority for granting access. Authorization-sensitive paths should check authoritative permission state or an explicitly safe consistency mechanism. If current authorization cannot be determined, fail closed.

Malware scanning can also be a security gate. If the product requires scanning before sharing or downloading, a failed scanner must leave the file quarantined or pending, not silently mark it usable. Thumbnail generation and search indexing, by contrast, can often fail without making the file itself unavailable.

12. Scaling the metadata and sync planes

Object storage handles bulk bytes; the metadata database handles frequently changing file, folder, permission, and version records. Read replicas may help appropriate read workloads, but stale replicas are unsafe for decisions requiring the latest ACL or concurrency state.

Owner-based sharding keeps a user's namespace relatively local. Its weakness is a very large or active user/tenant, and cross-owner “shared with me” queries. Tenant-based partitioning improves isolation but can produce giant-tenant hotspots. File-ID sharding spreads load but requires secondary indexes for folder and owner queries. There is no universally best shard key.

A huge folder should be listed with pagination, not enumerated in a single response. Fan-out notifications and derived indexes should be processed asynchronously through durable jobs. Per-tenant fairness and backpressure protect other users when one tenant generates extreme traffic.

The sync change stream also has independent limits: write throughput, cursor retention, ordering scope, connected-device count, and notification fan-out. Partitioning by user can help localize ordered processing if the underlying stream actually provides that ordering, but it does not replace operation idempotency or atomic version preconditions.

13. Failure scenarios that reveal correctness bugs

Failure Safe behavior
Upload interrupted Reconcile storage-confirmed parts; retry missing parts only
Upload succeeds, metadata commit fails Idempotent completion plus reconciliation; never publish unverified bytes
Presigned URL expires Check session status and refresh URLs for remaining parts if still valid
Metadata primary unavailable Fail over safely or reject mutations; don't acknowledge writes buffered only in memory
Storage temporarily unavailable Bound retries with backoff and jitter; expose degraded byte-transfer state
Sync push unavailable Poll durable change log on reconnect or schedule
Duplicate/out-of-order sync delivery Use operation IDs, version preconditions, and safe cursor checkpoints
Device offline beyond cursor retention Full snapshot reconciliation with a safe change-stream handoff
ACL revoked after URL issuance Recognize capability expiry window; stricter validation if required
Abandoned multipart sessions Abort expired sessions and reclaim parts
Garbage collector thinks a live object is orphaned Grace period, durable deletion state, reference checks, and independent audit
Regional outage Follow defined RPO/RTO; prevent split-brain writes during failover

The garbage collector deserves special attention. An object may appear unreferenced because a metadata query failed, a replication lagged, or a deduplication reference has not yet been observed. Never hard-delete a blob solely because one transient lookup returns no references. Use retention windows, tombstones, durable GC state, rechecks, and recoverable backups.

Cross-user content deduplication is optional, not a free storage win. It complicates reference counting and deletion, and a naive “does this hash already exist?” endpoint can leak information about another user's private data. Start without global deduplication unless measured storage savings justify its security and operational complexity.

For a scenario-by-scenario breakdown, see File Storage & Sync — Failure & Scale Scenarios.

14. Architectural trade-offs worth defending

Direct transfer or application proxy? Direct signed upload/download reduces application bandwidth at scale. Proxying remains appropriate for mandatory inline inspection, transformation, or real-time authorization controls.

Native multipart or application-managed chunk store? Native multipart simplifies reliable large uploads. A content-addressed chunk store can support cross-version reuse and deduplication but introduces reference-counting, garbage-collection, and privacy complexity.

Push or polling? Push reduces notification delay for connected clients. Polling is a necessary fallback. Neither replaces the durable change log.

Strong consistency or eventual consistency? The acceptance of a new current version needs a correct atomic precondition. Other devices may observe that committed version later. These are separate consistency boundaries.

Single-region or multi-region? A single-region primary with tested backups and a defined disaster-recovery plan is a defensible baseline. Cross-region replication can improve recovery objectives but requires handling replication lag and preventing two primaries from accepting conflicting writes during a partition. Object durability is not the same as an RPO or RTO guarantee.

For alternatives and decision criteria, see File Storage & Sync — Trade-offs & Decisions.

15. System design interview follow-ups

Why not keep file bytes in PostgreSQL? Small blobs may be fine there. At the stated multi-GB-file and petabyte-scale ingest workload, separate object storage better fits byte capacity, transfer, and lifecycle requirements.

What happens when a 50 GB upload fails at 90%? Recover the durable multipart session, reconcile parts confirmed by storage, refresh expired upload capabilities if the session is valid, and retry only missing parts.

Can the client trust ETag as an MD5? No. Provider, multipart, and encryption behavior affect its meaning. Specify explicit checksum verification.

How do devices discover missed changes? Fetch from the durable change log using a cursor. Push is a hint; polling and reconnect reconciliation provide recovery.

Does SHA-256 identify changed blocks? A single whole-file SHA-256 does not. Delta algorithms need block-level or rolling/content-defined signatures.

What if two offline devices edit the same file? Both carry the base version they edited. Atomic compare-and-set accepts one current version; the divergent edit is handled by an explicit conflict policy.

Can you instantly revoke an issued presigned URL? Not by changing only the file ACL. The already-issued capability may remain usable until expiry unless the access path supports further revocation checks or invalidation.

What if a device reconnects after six months? If its cursor is older than retained history, take a consistent snapshot, reconcile queued local changes, and rejoin the change stream at a safe checkpoint.

How do you prevent sync from losing a metadata mutation? Persist the mutation and its change-log/outbox record atomically or use an equivalent recoverable publication protocol. A best-effort queue publish after commit can lose events.

How do you recover from accidental object deletion? Recovery depends on version retention, backups, and provider replication capabilities. A design that promises recovery must test its restore procedure and define acceptable loss and recovery time.

A larger question bank is available in File Storage & Sync — Interview Perspective.

Final architecture

                         ┌──────────────────────────┐
Laptop / Phone / Web ───> │ API / Auth / File Service│ ──> Metadata DB
        │                └──────────┬───────────────┘     file IDs, ACLs,
        │                           │                     versions, folders
        │                    signed upload/download
        │                           │
        └──────── file bytes ───────┴───────────────> Object Storage
        │                                           immutable version blobs
        │
        └── Sync client ──> Sync API ──> Durable change log / outbox
                               │                    │
                               └── cursor reads <───┘
                                      │
                           Push hints / polling fallback

              Background workers: scanning, previews, indexing,
                    session reconciliation, retention / GC
Enter fullscreen mode Exit fullscreen mode

The important insight is not the number of boxes. It is the boundaries between guarantees:

  • Object storage protects and serves bytes; metadata identifies which bytes belong to which logical file version.
  • A verified upload is not visible until its version is safely published.
  • An accepted edit is conditionally committed; its arrival on other devices is asynchronous.
  • Push improves freshness, but durable replay preserves synchronization correctness.
  • Authorization gates capability issuance; signed URLs have explicit lifetime and revocation limitations.
  • Garbage collection must not destroy data that another component still references.

A good file storage and sync design makes those boundaries explicit. That is what turns a basic upload feature into a system that can survive interrupted transfers, offline edits, permission changes, retries, and operational failures without silently corrupting user state.

References and further reading

Top comments (1)

Collapse
 
dev_in_the_fog profile image
Jason Y. (dev_in_the_fog) •

Insightful breakdown! The architectural considerations for deterministic outputs and cost control were spot on.