RocksDB Inside Lioran S3
Lioran S3 stores object bytes on the filesystem.
So why does it need RocksDB?
Because an object store contains much more state than raw payload bytes.
I’m Swaraj Puppalwar, Founder & CTO of Lioran Group and Lioran Developer Solutions. This article examines the current RocksDbMetadataStore.
The invariant
The source is explicit:
Object payload/image bytes are NEVER written to RocksDB.
RocksDB is the metadata plane.
The filesystem is the object data plane.
The abstraction boundary
Higher layers depend on:
Arc<dyn MetadataStore>
rather than directly depending on RocksDbMetadataStore.
The trait exists so persistence details stay inside the metadata crate.
That gives the engine a boundary like:
object engine
server
multipart
media workers
↓
MetadataStore trait
↓
RocksDbMetadataStore
↓
RocksDB
Column families
The current engine defines dedicated column families for:
users
access_keys
buckets
objects
uploads
video_jobs
video_shares
video_manifests
system
plus RocksDB's default column family.
This keeps logically different metadata classes separated inside the embedded database.
Object key encoding
For object metadata, the current key encoder is straightforward:
fn encode_object_key(bucket: &str, key: &str) -> Vec<u8> {
format!("{bucket}/{key}").into_bytes()
}
So the logical metadata key is based on:
bucket/key
while the payload's physical filename remains an internal UUID path.
Those are two separate namespaces.
CPU-aware background work
When RocksDB opens, Bastion detects available parallelism:
std::thread::available_parallelism()
The configured background job count is bounded between 2 and 8.
That is then used for RocksDB parallelism/background jobs.
The point is not to let metadata compaction spawn an absurd amount of background work simply because the host reports many CPUs.
Incremental syncing
Database-level options include:
bytes_per_sync = 1 MiB
wal_bytes_per_sync = 1 MiB
The intent documented in the code is to reduce large OS page-writeback bursts.
WAL bound
The metadata database also bounds total WAL retention:
256 MiB
This is a practical operational choice because unbounded logs become their own storage problem.
Shared block cache
The current configuration creates one:
64 MiB LRU block cache
shared across column families.
For the high-volume objects family, table options include:
16 KiB block size
10-bit Bloom filter
cache index/filter blocks
pin L0 filter/index blocks
The comments describe object metadata as relatively small JSON records, roughly hundreds of bytes.
Bloom filters
Object GET and HEAD are often point lookups.
Bloom filters help answer:
This key is definitely not in this table.
without always requiring the same disk work for negative lookups.
They do not tell you the value.
They reduce unnecessary lookup cost.
Objects column family
The objects family is configured more aggressively:
write buffer: 64 MiB
max write buffers: 4
minimum buffers to merge: 2
target file size: 64 MiB
level base: 256 MiB
dynamic level sizing: enabled
compression: LZ4
This reflects the expectation that object metadata is a high-volume metadata class.
Smaller metadata families
Users, access keys, buckets, uploads, and video state use smaller write-buffer configurations.
For example, several use:
16 MiB write buffer
2 write buffers
LZ4
They share the same general block-cache/table strategy.
Serialization
Records are serialized with Serde JSON before being written into RocksDB.
For a user, conceptually:
serde_json::to_vec(user)
and on read:
serde_json::from_slice(&bytes)
This is simple and inspectable, although binary formats could trade readability for other characteristics later.
Why not put payloads in RocksDB?
Large object payloads have different access patterns.
A 20 GiB video wants:
- streaming
- ranges
- direct filesystem I/O
- predictable memory use
Metadata wants:
- point lookup
- ordered iteration
- compact records
- indexes
- transactional-ish state transitions
Using the same physical representation for both is unnecessary coupling.
The result
Lioran S3 uses RocksDB as the index and state engine around the object store, not as a giant blob container.
That distinction is the foundation of the rest of the architecture.
Top comments (0)