Failure Windows in Lioran S3 V1
Storage-engine design becomes interesting at the word:
crash.
Happy-path code is only half the system.
The current Lioran S3 V1 Pre-Alpha write path has a clear sequence of staging, promotion, metadata commit, and rollback. That sequence also defines its failure windows.
I’m Swaraj Puppalwar, Founder & CTO of Lioran Group and Lioran Developer Solutions.
This article describes the current implementation, including where future hardening can go deeper.
Current normal PUT sequence
The implementation today is:
1. validate bucket/key
2. check capacity/quota
3. create staging file
4. stream payload + SHA-256
5. flush
6. fsync in strict mode
7. close staging file
8. re-check capacity/quota
9. create final parent directory
10. rename staging → final object path
11. build ObjectMetadata
12. write metadata to RocksDB
13. if metadata write fails, remove final payload
14. return success
The most important ordering detail is:
physical promotion
BEFORE
metadata commit
Why document the real ordering?
An architecture document may describe an intended invariant at a higher level.
The source code is what defines the current executable behavior.
For pre-alpha infrastructure, publishing the actual order is more useful than pretending the engine has already reached its final crash-consistency model.
Failure during streaming
If reading from the incoming stream fails:
staging file is removed
metadata was never committed
object is not visible as committed
That is a clean failure case.
Failure after streaming but before promotion
Capacity/quota checks can still reject the upload after its exact size is known.
The staging file is removed.
Again, no committed metadata has been created.
Rename failure
If:
fs::rename(staging, final)
fails, the operation returns a storage error and attempts staging cleanup.
Metadata has not yet been committed.
Metadata write failure
This case is more interesting.
At that moment:
physical payload has already been promoted
metadata has not successfully committed
The implementation responds by removing the physical payload:
fs::remove_file(&abs_path)
This is application-level rollback.
Process crash between rename and metadata
Now consider an uncontrolled process termination exactly after physical rename but before metadata commit.
The rollback code does not execute because the process is gone.
That can theoretically leave physical bytes without a corresponding metadata record.
This is a classic cross-resource atomicity problem:
filesystem namespace
+
embedded metadata database
are not one shared transaction.
Why this matters
Normal reads begin from metadata.
So an unreferenced physical file is not automatically exposed as a committed object.
But it can become orphaned disk usage.
That is different from exposing partial data, but it is still a recovery concern.
The inverse ordering has its own problem
Suppose metadata were committed first and the process died before physical promotion.
Then metadata could point to missing physical bytes.
That is usually worse for read correctness.
This is why robust storage systems often introduce additional durable intent/state machinery instead of pretending two independent persistence domains can be made atomic with ordering alone.
V1's current strategy
The current code uses:
- staging
- strict/balanced durability
- atomic filesystem rename semantics where supported
- metadata status checks
- rollback on synchronous metadata failure
- logical/physical separation
It does not magically turn filesystem + RocksDB into one atomic transaction.
That would be an inaccurate claim.
Future hardening direction
A stronger recovery protocol could involve concepts such as:
durable publish intent
recovery journal
startup reconciliation
orphan scanning
transaction state records
idempotent repair
Those are design directions, not claims about the current V1 implementation.
Why pre-alpha is the right time to discuss this
Pre-alpha should expose uncomfortable edge cases.
That is the phase for:
kill -9
power-loss simulation
disk-full testing
metadata failure injection
rename failure injection
orphan detection
checksum verification
restart loops
The purpose is not to prove the code never fails.
The purpose is to make every failure state defined and recoverable.
Correctness before architecture theater
Distributed replication would not remove this local commit problem.
It could multiply it.
That is why Lioran S3 intentionally focuses on single-node correctness before adding clustering and consensus layers.
A storage engine becomes trustworthy when its failure states are boring, documented, testable, and repairable.
That is the target.
Top comments (0)