GET and Range Reads Inside Lioran S3
Writes get most of the dramatic engineering diagrams.
Reads are where object storage spends much of its life.
The current Lioran S3 local engine keeps the read path deliberately direct:
logical key
↓
metadata lookup
↓
physical path
↓
Tokio file
↓
async stream
I’m Swaraj Puppalwar, Founder & CTO of Lioran Group / Lioran Developer Solutions.
Metadata first
A GET begins with:
self.metadata.get_object(bucket, key)?
If no metadata record exists:
ObjectNotFound
is returned.
The engine also verifies that the record is:
not deleted
status == Committed
So an incomplete/non-committed state does not become a normal GET response.
Resolve placement
The metadata record contains a relative physical path.
The engine resolves it underneath the configured storage root:
self.layout.resolve_absolute_path(&meta.physical_path)
Then Tokio opens the file:
File::open(&abs_path).await
Missing physical payload
An interesting corruption/error case is:
metadata says committed
but physical file is missing
The implementation does not silently convert that into a normal object-not-found condition.
It reports a storage error describing the missing physical payload for the committed key.
That distinction is valuable operationally.
There is a difference between:
user requested a key that never existed
and:
metadata references bytes that disappeared
Return a reader, not a Vec
The engine returns:
Box<dyn AsyncRead + Send + Unpin>
with the metadata.
Again, the object layer does not need to materialize the complete payload into memory.
The HTTP layer can stream from the reader.
Range reads
get_object_range follows the same metadata validation and physical-path resolution.
Then:
file.seek(SeekFrom::Start(start)).await
moves the file cursor.
The returned reader is limited:
file.take(length)
Conceptually:
large object
|----------------------------------------------|
^ start
|---------- requested ----------|
Only the requested range needs to flow through the response.
Why ranges matter
Byte ranges enable:
- video seeking
- media players
- resumable downloads
- archive inspection
- partial dataset reads
- clients that need headers or tails of large objects
For large media, range support is not a cute extra.
It is part of practical delivery.
HEAD
head_object stops after metadata validation.
It does not need to open and stream the physical payload.
That makes HEAD the correct operation when the caller only needs:
- size
- content type
- checksum/ETag-style metadata
- object existence
Downloading an object to learn its size would be infrastructure slapstick.
Read-path complexity
The normal read path avoids:
- copying the full payload into RocksDB
- copying the full payload into a Rust
Vec - reconstructing a filesystem path from the raw user key
- doing media-specific processing for normal objects
It resolves metadata, opens a file, and streams.
That simplicity matters.
Where performance can move
For a GET, latency can come from:
metadata lookup
filesystem open
seek
disk/cache read
network send
TLS/reverse proxy
client receive
If the payload is already in the OS page cache, the physical storage device may barely participate.
If it is a cold read, disk behavior matters much more.
That is why “read speed” is not one universal number.
Read correctness first
A fast read from the wrong physical object is not a successful optimization.
The metadata-to-placement mapping is the core correctness step.
Everything after that is moving bytes.
Top comments (0)