Bounded-Memory Streaming in Lioran S3
A storage engine that can accept a 100 GiB object should not need a 100 GiB heap.
That sounds obvious. It is also one of the easiest properties to accidentally destroy when abstractions begin buffering data behind your back.
In Lioran S3 / Lioran Bastion, the core streaming helper is deliberately small and explicit.
I’m Swaraj Puppalwar, Founder & CTO of Lioran Group and Lioran Developer Solutions. This article examines that Rust data path.
The invariant
The source states the rule directly:
Object uploads and downloads strictly utilize fixed-size buffer chunks.
Under no circumstances are entire payloads buffered into heap memory.
For normal object PUTs, the default buffer is:
pub const DEFAULT_STREAM_CHUNK_SIZE: usize = 256 * 1024;
That is 256 KiB.
The function
The core helper accepts:
pub async fn stream_to_staging(
reader: &mut (dyn AsyncRead + Send + Unpin),
staging_file: &mut File,
durability: DurabilityMode,
chunk_size: usize,
) -> Result<StreamWriteResult>
Notice what it does not accept.
There is no:
Vec<u8>
representing the entire object.
The source is an asynchronous reader.
Allocate once
The effective chunk size is selected and one buffer is allocated:
let mut buffer = vec![0u8; effective_chunk_size];
The loop then reuses it.
Conceptually:
network
↓
[ 256 KiB buffer ]
↓
filesystem
A 10 MiB object and a 100 GiB object can therefore use the same primary streaming-buffer size.
Concurrency still multiplies memory usage, of course.
If 100 uploads each own a 256 KiB buffer, buffers alone are roughly:
100 × 256 KiB ≈ 25 MiB
That is why bounded memory and bounded concurrency belong in the same conversation.
Read
Each iteration calls:
reader.read(&mut buffer).await
EOF is represented by:
n == 0
The engine tracks cumulative receive duration separately.
Hash
Only the valid slice is hashed:
let chunk = &buffer[..n];
hasher.update(chunk);
SHA-256 is incremental.
The entire object is never reconstructed in memory merely to calculate its digest.
Write
The same chunk is written:
staging_file.write_all(chunk).await
Then:
total_bytes += n as u64;
The loop repeats.
Progress visibility
Every five seconds the streaming path can log:
- bytes received
- elapsed time
- MiB received
- effective MiB/s
This is useful when debugging the classic storage question:
Is the server slow, or is the server waiting for the client/network?
The implementation separately accumulates recv_duration and write_duration, which makes that distinction measurable.
Flush is not fsync
After EOF:
staging_file.flush().await
Flush and physical durability are not treated as synonyms.
If the selected durability mode requires it:
staging_file.sync_all().await
is also issued.
That means the data path explicitly separates:
application → OS
from:
OS → durable storage boundary
Timing structure
The function returns a StreamWriteResult containing:
pub struct StreamWriteResult {
pub bytes_written: u64,
pub sha256_hex: String,
pub timings: StreamTimings,
}
StreamTimings tracks:
recv_duration
write_duration
sha256_duration
flush_duration
fsync_duration
total_stream_duration
That instrumentation is valuable because throughput can be limited by very different things.
Network-bound
recv_duration dominates
Disk-bound
write_duration / fsync_duration dominates
CPU hashing pressure
sha256_duration becomes significant
Why AsyncRead
The object engine does not need to know whether bytes originated from:
- an HTTP request body
- a file
- a test stream
- another asynchronous producer
It knows only that it can asynchronously read bytes.
That keeps the object engine below the HTTP layer.
Multipart uses the same philosophy
The multipart engine uses its own bounded copy buffer:
const STREAM_BUFFER_SIZE: usize = 128 * 1024;
and adds global concurrency control.
So the same design principle survives when one object becomes many independently uploaded parts.
Memory complexity
For one normal upload, the dominant explicit streaming buffer is effectively:
O(chunk_size)
rather than:
O(object_size)
That difference is the difference between a storage engine and a very expensive read_to_end().
The bigger rule
When building infrastructure, performance is not just “make the loop fast”.
It is often:
make resource consumption predictable as input size grows.
Lioran S3's bounded streaming path is intentionally boring code.
For storage software, boring data paths are beautiful.
Top comments (0)