DEV Community

Cover image for Durability, Crash Safety and Disk Guardrails in Lioran S3
Swaraj Puppalwar
Swaraj Puppalwar

Posted on

Durability, Crash Safety and Disk Guardrails in Lioran S3

Durability, Crash Safety and Disk Guardrails in Lioran S3

Object storage is easy if the process never crashes, disks never fill, power never disappears, and nobody uploads large files.

Unfortunately, reality has an excellent QA team.

This article documents the durability model currently exposed by Lioran S3 V1 Pre-Alpha, built by Lioran Developer Solutions under Lioran Group, led by Swaraj Puppalwar.

Storage success has a meaning

When an API returns success for a write, the system must define what has actually become durable.

Without that definition, “200 OK” is vibes.

Lioran S3 provides explicit durability modes.

Strict mode

BASTION_DURABILITY=strict
Enter fullscreen mode Exit fullscreen mode

Strict mode uses explicit flush/sync boundaries before storage metadata is committed as durable state.

Conceptually:

receive bytes
   ↓
write staged file
   ↓
flush
   ↓
fsync
   ↓
commit metadata
   ↓
promote
   ↓
acknowledge success
Enter fullscreen mode Exit fullscreen mode

This costs throughput because synchronous persistence has a physical price.

That cost is the point.

Balanced mode

BASTION_DURABILITY=balanced
Enter fullscreen mode Exit fullscreen mode

Balanced mode avoids the same per-write synchronous persistence boundary and relies more on operating-system/filesystem writeback behavior.

It can improve throughput.

It also changes the power-loss durability assumptions.

Temporary objects

Incomplete writes remain temporary.

They should not become visible through normal object APIs.

That means an interrupted request should not leave an object that GET reports as if it were fully committed.

Multipart state

Multipart uploads are another place where explicit state matters.

A multipart session can contain uploaded parts without exposing a final assembled object.

Only completion transitions the upload into committed object state.

Existing objects

An interrupted replacement should not corrupt the previously committed object.

Storage mutation paths need staging and promotion semantics specifically to avoid mixing “new object is incomplete” with “old object no longer exists”.

Disk low-watermark protection

Object stores are extremely capable of filling disks.

That is their job.

The server therefore exposes minimum free-space guardrails.

For example:

BASTION_MIN_FREE_SPACE_BYTES=536870912
Enter fullscreen mode Exit fullscreen mode

That value represents 512 MiB.

Production deployments should choose a threshold appropriate to their disk size and neighboring workloads.

An optional percentage-based threshold can also be used.

Bucket quotas

Host headroom and bucket quotas solve different problems.

Bucket quota:

How much may this logical tenant/bucket consume?
Enter fullscreen mode Exit fullscreen mode

Disk low-watermark:

How close may the entire host get to physical exhaustion?
Enter fullscreen mode Exit fullscreen mode

Use both.

Checksums

The object pipeline tracks integrity metadata such as SHA-256.

Checksums are useful for:

  • verifying uploads
  • detecting accidental corruption
  • validating downloaded content
  • establishing ETags/integrity identifiers

They are not a substitute for backups or replication.

Graceful shutdown

The server handles process shutdown so active requests can drain before metadata stores close.

That improves controlled restart behavior.

It does not eliminate the need to design for uncontrolled termination.

kill -9 and power loss do not care how polite your shutdown handler is.

Crash testing

A serious storage test matrix should include:

PUT during process termination
multipart during termination
GET during termination
disk full
quota exceeded
invalid range
metadata restart
reverse proxy restart
host reboot
filesystem remount
credential rotation
Enter fullscreen mode Exit fullscreen mode

The point is not merely whether the process starts again.

The point is whether previously acknowledged committed data remains correct and incomplete data remains distinguishable.

Operational recommendation

During pre-alpha evaluation:

  • use strict when validating durability behavior
  • run repeated restart tests
  • compare checksums
  • fill disks intentionally in disposable environments
  • test quota boundaries
  • test interrupted multipart uploads
  • inspect orphan cleanup
  • measure memory during large streams
  • pin versions before testing so behavior is reproducible

Throughput versus durability

There is no honest storage benchmark without describing durability mode.

A write benchmark with aggressive OS buffering and a benchmark with synchronous durability are measuring different contracts.

When publishing numbers, include the contract.

The next Lioran S3 version is planned for November 17, 2026, and durability behavior will continue to be one of the core areas under evaluation.

Docs: https://docs.liorans3.sbs

Top comments (0)