DEV Community

Galadd
Galadd

Posted on

Cutting Beacon State Storage by ~97%: Validator Sectors with Structural Sharing

A design prototype. I’d like your feedback before I build the reconstruction measurements.

TL;DR

Storing full beacon states at regular intervals costs roughly 3 TB. Diffing against a base reference (eth-state-diff) brings that down to about 400 GB. But ~85% of a beacon state is the validators list, which barely changes. By treating validators separately as sectors of structurally shared chunks, plus layered checkpoints of diffs, the estimate drops to ~70 GB, a reduction of roughly 97% from the naive approach.

All numbers are maximum approximations. This post is about the design prototype, not benchmarks.

Background

I’ve been looking into designs that reduce redundant on-disk data and speed up state reconstruction and historical state lookups. That work led to eth-state-diff, which computes the difference between two beacon states and can be reapplied to reconstruct one from the other.

It has a clear limitation: it always needs a base reference, and the farther the base is from the target, the larger the diff. A pure diff scheme is not a complete solution. It needs periodic full-state anchors (I’m thinking every 256 epochs).

Full anchors are still wasteful, because most of each one is almost identical to the last.

The observation

About 85% of the beacon state is the validators list, and that list is append-mostly:

  • New validators are appended, and the appends are clustered.
  • In-place mutations are extremely rare—as few as 0.015% of validators.

If the list were strictly append-only, a checkpoint would only need to store a length. It isn’t, but it’s close enough that storing a full copy every time is almost pure redundancy.

So the design question is: how do you exploit near-100% redundancy?

Validators need their own representation.



The design

Validators are stored in two cooperating layers: sectors (structural sharing) and checkpoints (diffs).

Layer 1: Sectors (structural sharing)

A sector spans 8 checkpoints of 256 epochs each → 2,048 epochs.

At every sector boundary the full validators list is materialised, but not as a monolithic blob. It is split into chunks of 64 validators (64 × 121 bytes ≈ 7.7 KB each). Each chunk is stored once and addressed by name (think content-addressed). A sector’s validator list is simply a list of chunk references:

  • The reference list is small (~250 KB), so a plain list is sufficient. A persistent tree would be over-engineering here.
  • When validators change, only the affected chunks are written. Unchanged chunks are shared with earlier sectors.

Cost at a maximum of 150 changes per sector:

150 changes × 121 B × 64 (whole chunk copied) ≈ 1.2 MB
Enter fullscreen mode Exit fullscreen mode

Appends are clustered, so they create few chunks (max ~25 KB). Across the full history the changed chunks total roughly 300 MB and the chunk-reference lists roughly 57 MB.

The tradeoff is intentional: a change to a single 8-byte field still copies its whole chunk, so this is not space-optimal in the small. But it never requires a deep copy of the validators list, and at this scale the overhead is tiny.

Layer 2: Checkpoints (layered diffs)

Inside a sector, checkpoints record diffs of the validators list:

  • Every 256 epochs a “major” checkpoint stores a diff against the sector base.
  • Every 32 epochs a “minor” checkpoint stores a diff against the nearest preceding major checkpoint.

Concrete layout (example at epoch 8768)

Sector 0 (epoch 0): full validators (chunks)
  ├─ minor 32:   diff(32, sector 0)
  ├─ minor 64:   diff(64, sector 0)
  ├─ … 
  └─ major 256:  diff(256, sector 0)
       ├─ minor 288: diff(288, major 256)
       ├─ … 
       └─ major 512: diff(512, sector 0)
            …

Sector 8192 (epoch 8192): full validators (new ChunksRef list)
  ├─ minor 8224: diff(8224, sector 8192)
  ├─ …
  ├─ major 8448: diff(8448, sector 8192)
  ├─ …
  ├─ major 8704: diff(8704, sector 8192)
  └─ minor 8768: diff(8768, major 8704)
Enter fullscreen mode Exit fullscreen mode

Reconstruction at epoch E

  1. Locate the sector base: floor(E / 2048) × 2048.
  2. Locate the nearest major checkpoint ≤ E: floor(E / 256) × 256.
  3. Locate the exact minor checkpoint (or the target epoch itself).
  4. Start from the sector’s full validators (resolved via its ChunksRef list).
  5. Apply the major-checkpoint diff.
  6. Apply the minor-checkpoint diff.

Example for epoch 8768:

validators@8768 = sector@8192 + diff@8704 + diff@8768
Enter fullscreen mode Exit fullscreen mode

Estimated storage

Component Size
Naive: full state every 32 epochs ~3 TB
Diffs every 32 epochs, base every 256 ~400 GB
This design
Sector bases (full state without validators) ~60 GB
Changed validator chunks ~300 MB
Chunk references ~57 MB
Validator patches / minor diffs ~550 MB
Diffs (remaining state + major checkpoints) ~10 GB
Total ~70 GB

≈ 97.7 % reduction versus the naive approach and ≈ 82 % reduction versus plain eth-state-diff.



What I still need to measure

Storage is only half the story. I will look into concrete reconstruction-cost numbers:

  • Latency vs distance from the sector base
  • Bytes read per reconstruction
  • Number of diffs applied and chunks fetched

Feedback wanted

  1. Is a 64-validator chunk the right granularity? Smaller chunks copy less on mutation but grow the reference list.
  2. Has anyone seen similar layered structural-sharing + diff approaches in other clients’ state stores?

All figures above are max approximates. The focus is the design prototype.

Top comments (0)