DEV Community

Cover image for RocksDB Inside Lioran S3: The Metadata Engine Behind Buckets, Objects and Uploads
Swaraj Puppalwar
Swaraj Puppalwar

Posted on

RocksDB Inside Lioran S3: The Metadata Engine Behind Buckets, Objects and Uploads

RocksDB Inside Lioran S3

Lioran S3 stores object bytes on the filesystem.

So why does it need RocksDB?

Because an object store contains much more state than raw payload bytes.

I’m Swaraj Puppalwar, Founder & CTO of Lioran Group and Lioran Developer Solutions. This article examines the current RocksDbMetadataStore.

The invariant

The source is explicit:

Object payload/image bytes are NEVER written to RocksDB.
Enter fullscreen mode Exit fullscreen mode

RocksDB is the metadata plane.

The filesystem is the object data plane.

The abstraction boundary

Higher layers depend on:

Arc<dyn MetadataStore>
Enter fullscreen mode Exit fullscreen mode

rather than directly depending on RocksDbMetadataStore.

The trait exists so persistence details stay inside the metadata crate.

That gives the engine a boundary like:

object engine
server
multipart
media workers
     ↓
MetadataStore trait
     ↓
RocksDbMetadataStore
     ↓
RocksDB
Enter fullscreen mode Exit fullscreen mode

Column families

The current engine defines dedicated column families for:

users
access_keys
buckets
objects
uploads
video_jobs
video_shares
video_manifests
system
Enter fullscreen mode Exit fullscreen mode

plus RocksDB's default column family.

This keeps logically different metadata classes separated inside the embedded database.

Object key encoding

For object metadata, the current key encoder is straightforward:

fn encode_object_key(bucket: &str, key: &str) -> Vec<u8> {
    format!("{bucket}/{key}").into_bytes()
}
Enter fullscreen mode Exit fullscreen mode

So the logical metadata key is based on:

bucket/key
Enter fullscreen mode Exit fullscreen mode

while the payload's physical filename remains an internal UUID path.

Those are two separate namespaces.

CPU-aware background work

When RocksDB opens, Bastion detects available parallelism:

std::thread::available_parallelism()
Enter fullscreen mode Exit fullscreen mode

The configured background job count is bounded between 2 and 8.

That is then used for RocksDB parallelism/background jobs.

The point is not to let metadata compaction spawn an absurd amount of background work simply because the host reports many CPUs.

Incremental syncing

Database-level options include:

bytes_per_sync = 1 MiB
wal_bytes_per_sync = 1 MiB
Enter fullscreen mode Exit fullscreen mode

The intent documented in the code is to reduce large OS page-writeback bursts.

WAL bound

The metadata database also bounds total WAL retention:

256 MiB
Enter fullscreen mode Exit fullscreen mode

This is a practical operational choice because unbounded logs become their own storage problem.

Shared block cache

The current configuration creates one:

64 MiB LRU block cache
Enter fullscreen mode Exit fullscreen mode

shared across column families.

For the high-volume objects family, table options include:

16 KiB block size
10-bit Bloom filter
cache index/filter blocks
pin L0 filter/index blocks
Enter fullscreen mode Exit fullscreen mode

The comments describe object metadata as relatively small JSON records, roughly hundreds of bytes.

Bloom filters

Object GET and HEAD are often point lookups.

Bloom filters help answer:

This key is definitely not in this table.

without always requiring the same disk work for negative lookups.

They do not tell you the value.

They reduce unnecessary lookup cost.

Objects column family

The objects family is configured more aggressively:

write buffer: 64 MiB
max write buffers: 4
minimum buffers to merge: 2
target file size: 64 MiB
level base: 256 MiB
dynamic level sizing: enabled
compression: LZ4
Enter fullscreen mode Exit fullscreen mode

This reflects the expectation that object metadata is a high-volume metadata class.

Smaller metadata families

Users, access keys, buckets, uploads, and video state use smaller write-buffer configurations.

For example, several use:

16 MiB write buffer
2 write buffers
LZ4
Enter fullscreen mode Exit fullscreen mode

They share the same general block-cache/table strategy.

Serialization

Records are serialized with Serde JSON before being written into RocksDB.

For a user, conceptually:

serde_json::to_vec(user)
Enter fullscreen mode Exit fullscreen mode

and on read:

serde_json::from_slice(&bytes)
Enter fullscreen mode Exit fullscreen mode

This is simple and inspectable, although binary formats could trade readability for other characteristics later.

Why not put payloads in RocksDB?

Large object payloads have different access patterns.

A 20 GiB video wants:

  • streaming
  • ranges
  • direct filesystem I/O
  • predictable memory use

Metadata wants:

  • point lookup
  • ordered iteration
  • compact records
  • indexes
  • transactional-ish state transitions

Using the same physical representation for both is unnecessary coupling.

The result

Lioran S3 uses RocksDB as the index and state engine around the object store, not as a giant blob container.

That distinction is the foundation of the rest of the architecture.

Top comments (0)