Measuring the Lioran S3 Rust Storage Engine
A benchmark that says:
WRITE: 400 MB/s
creates more questions than it answers.
Where did time go?
Network?
SHA-256?
Filesystem writes?
fsync?
Directory creation?
Rename?
RocksDB metadata?
Lioran S3's normal PUT path contains explicit timing instrumentation for these stages.
I’m Swaraj Puppalwar, Founder & CTO of Lioran Group / Lioran Developer Solutions.
Enabling detailed PUT traces
The engine checks:
BASTION_TRACE_PUT_TIMINGS
Accepted truthy forms include values such as:
1
true
yes
The setting is cached in an atomic so the environment is not reparsed on every request.
Streaming timings
stream_to_staging measures:
recv_duration
write_duration
sha256_duration
flush_duration
fsync_duration
total_stream_duration
Then the outer PUT path adds:
close_duration
mkdir_duration
rename_duration
metadata_duration
total_duration
A conceptual trace
A PUT can therefore be decomposed like:
TOTAL PUT
├── receive from AsyncRead
├── SHA-256
├── staging writes
├── flush
├── fsync
├── close
├── mkdir
├── rename
└── RocksDB metadata commit
This is exactly the decomposition you want when tuning a storage engine.
Case 1: receive dominates
Suppose:
recv_ms 900
write_ms 100
sha256_ms 30
fsync_ms 20
metadata_ms 2
The server is mostly waiting for incoming bytes.
Optimizing RocksDB will not magically fix that.
Investigate:
- client upload speed
- network
- TLS
- reverse proxy buffering
- request-body behavior
Case 2: fsync dominates
Suppose:
recv_ms 100
write_ms 80
fsync_ms 700
Now durability synchronization is expensive.
Questions include:
- storage device latency
- filesystem
- virtualization
- durability mode
- write cache behavior
Case 3: metadata dominates
If:
metadata_ms
grows significantly under load, inspect the metadata path:
- RocksDB compaction
- WAL behavior
- block cache
- write stalls
- disk contention
- concurrent metadata operations
Case 4: hashing matters
SHA-256 is calculated incrementally on every incoming chunk.
At very high throughput, checksum CPU cost can become measurable.
That does not mean “remove checksums”.
It means measure the actual cost before inventing a bottleneck.
Aggregated metrics
LocalObjectStore also owns:
Arc<PutMetrics>
and records successful and failed PUT performance.
This lets the engine expose more than one request's anecdote.
Why high-resolution timing matters
Without stage timings, developers often optimize whatever subsystem they personally distrust.
Database engineer:
must be RocksDB
Networking engineer:
must be TLS
Systems engineer:
obviously fsync
The stopwatch is less political.
Benchmark contract
When publishing Lioran S3 numbers, record:
CPU
RAM
disk
filesystem
operating system
container/native
network topology
reverse proxy
TLS
object size distribution
concurrency
durability mode
software commit/version
Otherwise two benchmark runs may not be measuring the same system.
Optimization order
A sane loop is:
measure
↓
identify dominant stage
↓
change one thing
↓
measure again
not:
change 14 knobs
↓
number improved
↓
no idea why
Lioran S3's timing instrumentation exists so performance work can stay evidence-driven.
Top comments (0)