Durability, Crash Safety and Disk Guardrails in Lioran S3
Object storage is easy if the process never crashes, disks never fill, power never disappears, and nobody uploads large files.
Unfortunately, reality has an excellent QA team.
This article documents the durability model currently exposed by Lioran S3 V1 Pre-Alpha, built by Lioran Developer Solutions under Lioran Group, led by Swaraj Puppalwar.
Storage success has a meaning
When an API returns success for a write, the system must define what has actually become durable.
Without that definition, “200 OK” is vibes.
Lioran S3 provides explicit durability modes.
Strict mode
BASTION_DURABILITY=strict
Strict mode uses explicit flush/sync boundaries before storage metadata is committed as durable state.
Conceptually:
receive bytes
↓
write staged file
↓
flush
↓
fsync
↓
commit metadata
↓
promote
↓
acknowledge success
This costs throughput because synchronous persistence has a physical price.
That cost is the point.
Balanced mode
BASTION_DURABILITY=balanced
Balanced mode avoids the same per-write synchronous persistence boundary and relies more on operating-system/filesystem writeback behavior.
It can improve throughput.
It also changes the power-loss durability assumptions.
Temporary objects
Incomplete writes remain temporary.
They should not become visible through normal object APIs.
That means an interrupted request should not leave an object that GET reports as if it were fully committed.
Multipart state
Multipart uploads are another place where explicit state matters.
A multipart session can contain uploaded parts without exposing a final assembled object.
Only completion transitions the upload into committed object state.
Existing objects
An interrupted replacement should not corrupt the previously committed object.
Storage mutation paths need staging and promotion semantics specifically to avoid mixing “new object is incomplete” with “old object no longer exists”.
Disk low-watermark protection
Object stores are extremely capable of filling disks.
That is their job.
The server therefore exposes minimum free-space guardrails.
For example:
BASTION_MIN_FREE_SPACE_BYTES=536870912
That value represents 512 MiB.
Production deployments should choose a threshold appropriate to their disk size and neighboring workloads.
An optional percentage-based threshold can also be used.
Bucket quotas
Host headroom and bucket quotas solve different problems.
Bucket quota:
How much may this logical tenant/bucket consume?
Disk low-watermark:
How close may the entire host get to physical exhaustion?
Use both.
Checksums
The object pipeline tracks integrity metadata such as SHA-256.
Checksums are useful for:
- verifying uploads
- detecting accidental corruption
- validating downloaded content
- establishing ETags/integrity identifiers
They are not a substitute for backups or replication.
Graceful shutdown
The server handles process shutdown so active requests can drain before metadata stores close.
That improves controlled restart behavior.
It does not eliminate the need to design for uncontrolled termination.
kill -9 and power loss do not care how polite your shutdown handler is.
Crash testing
A serious storage test matrix should include:
PUT during process termination
multipart during termination
GET during termination
disk full
quota exceeded
invalid range
metadata restart
reverse proxy restart
host reboot
filesystem remount
credential rotation
The point is not merely whether the process starts again.
The point is whether previously acknowledged committed data remains correct and incomplete data remains distinguishable.
Operational recommendation
During pre-alpha evaluation:
- use
strictwhen validating durability behavior - run repeated restart tests
- compare checksums
- fill disks intentionally in disposable environments
- test quota boundaries
- test interrupted multipart uploads
- inspect orphan cleanup
- measure memory during large streams
- pin versions before testing so behavior is reproducible
Throughput versus durability
There is no honest storage benchmark without describing durability mode.
A write benchmark with aggressive OS buffering and a benchmark with synchronous durability are measuring different contracts.
When publishing numbers, include the contract.
The next Lioran S3 version is planned for November 17, 2026, and durability behavior will continue to be one of the core areas under evaluation.
Top comments (0)