DEV Community

Cover image for Strict vs Balanced Durability in Lioran S3: What fsync Actually Changes
Swaraj Puppalwar
Swaraj Puppalwar

Posted on

Strict vs Balanced Durability in Lioran S3: What fsync Actually Changes

Strict vs Balanced Durability in Lioran S3

A storage benchmark without a durability contract is incomplete.

Two systems can both claim:

500 MB/s writes
Enter fullscreen mode Exit fullscreen mode

while promising very different things when the machine loses power.

Lioran S3 exposes that tradeoff explicitly through DurabilityMode.

I’m Swaraj Puppalwar, Founder & CTO of Lioran Group / Lioran Developer Solutions. Here is what the current Rust path actually does.

Streaming always writes and flushes

At the end of stream_to_staging, Bastion calls:

staging_file.flush().await
Enter fullscreen mode Exit fullscreen mode

This pushes buffered application writes toward the operating system.

But flush is not automatically equivalent to stable physical persistence.

Strict durability

The streaming path asks:

durability.requires_fsync()
Enter fullscreen mode Exit fullscreen mode

If true:

staging_file.sync_all().await
Enter fullscreen mode Exit fullscreen mode

is issued before the streaming helper returns.

Conceptually:

application
   ↓ write
userspace / runtime buffering
   ↓ flush
OS / filesystem cache
   ↓ fsync / sync_all
durable-storage boundary
Enter fullscreen mode Exit fullscreen mode

Strict mode pays the latency cost of that synchronization.

Balanced durability

When the mode does not require fsync, the explicit sync_all() stage is skipped.

The write can therefore depend more heavily on operating-system/filesystem writeback behavior.

That can improve throughput.

It also changes the failure contract.

Why fsync can dominate benchmarks

Imagine a client can send data at:

1 GB/s
Enter fullscreen mode Exit fullscreen mode

and the filesystem can absorb writes into cache quickly.

If each completed object requires a durable synchronization, the final latency can still be dominated by the storage device and filesystem's persistence behavior.

That is not necessarily inefficiency.

It is work required by the selected contract.

Bastion measures it separately

The streaming timing structure includes:

flush_duration
fsync_duration
Enter fullscreen mode Exit fullscreen mode

The full PUT path can emit:

recv_ms
write_ms
sha256_ms
flush_ms
fsync_ms
close_ms
mkdir_ms
rename_ms
metadata_ms
total_ms
Enter fullscreen mode Exit fullscreen mode

So if strict mode suddenly cuts throughput, we can inspect whether fsync_ms is actually responsible.

Staging matters

The fsync happens while the payload is still represented as the staging file.

Only after streaming succeeds does the normal PUT path move toward final placement and metadata commit.

This helps prevent a partially received network stream from being treated as a normal committed object.

Current commit ordering

One subtle implementation detail matters.

The current normal PUT path:

  1. streams and optionally fsyncs staging
  2. closes the staging handle
  3. verifies capacity/quota
  4. renames staging into final object placement
  5. writes metadata to RocksDB
  6. removes the physical file if metadata commit fails

That is the implementation today.

It is worth documenting because durability discussions become meaningless when commit ordering is described only at a diagram level.

Power loss is different from process failure

A graceful shutdown is one thing.

A process crash is another.

A host power loss is another.

An SSD lying about persistence is another.

fsync strengthens the contract, but “durable storage” still ultimately depends on the operating system, filesystem, device, controller, and hardware behavior underneath.

How to benchmark honestly

Always record:

durability mode
filesystem
disk type
object size
concurrency
reverse proxy
network path
CPU
RAM
software version
Enter fullscreen mode Exit fullscreen mode

A throughput number detached from these details is mostly decoration.

Why keep both modes?

Different workloads value different things.

A disposable cache-like workload may prefer throughput.

A source-of-truth object workload may prefer stronger synchronization.

The engine should make that decision explicit rather than hiding it behind one unexplained benchmark preset.

That is why Lioran S3 treats durability as part of configuration, not as a footnote.

Top comments (0)