Multipart Upload Internals in Lioran S3
Large uploads are not just normal PUTs with bigger numbers.
They need resumability, part integrity, state persistence, concurrency limits, expiry, quota accounting, and final assembly.
Lioran S3 implements this in MultipartUploadManager.
I’m Swaraj Puppalwar, Founder & CTO of Lioran Group and Lioran Developer Solutions.
Configuration
The current defaults include:
recommended part size 32 MiB
minimum part size 5 MiB
maximum part size 512 MiB
recommended parallelism 8
session TTL 24 hours
global part concurrency 64
stream copy buffer 128 KiB
These values are not all hard-coded client behavior. Some are recommendations or configurable guardrails.
Manager state
The manager owns:
pub struct MultipartUploadManager {
layout: StorageLayout,
metadata: Arc<dyn MetadataStore>,
config: UploadConfig,
durability: DurabilityMode,
global_limiter: Arc<Semaphore>,
session_update_lock: Arc<Mutex<()>>,
}
Two synchronization primitives stand out:
Semaphore
Mutex
They solve different problems.
Global semaphore
Every part upload acquires a permit:
self.global_limiter.acquire().await
This provides global backpressure.
Without it, 1,000 clients each deciding to upload 64 parts concurrently could create a hilarious and short-lived benchmark.
The semaphore limits active part-streaming work across sessions.
Session update lock
Multipart parts can finish concurrently.
But shared session metadata such as:
bytes_received
parts
updated_at
state
must be updated coherently.
That is why the manager also owns a session update mutex.
Concurrency belongs on the data path.
Serialization belongs around shared mutable state that actually requires it.
Initializing a session
init_upload validates:
- object key
- bucket existence
- host free space
- bucket quota
- active upload-session count
- expected payload size
Then it selects a part size.
The current adaptive recommendation is conceptually:
<= 1 GiB → 16 MiB
> 1 GiB → 32 MiB
> 10 GiB → 64 MiB
subject to configured bounds.
Persistent identity
A session receives:
upload_id
video_id
and expiration metadata.
The session directory is created under:
staging/uploads/<upload-id>/
Then the session record is persisted through MetadataStore.
That is what makes resumability a state problem, not merely an HTTP retry trick.
Part numbers
The current implementation accepts:
1..=10000
for part numbers.
Each part has a final temporary part path and an .uploading path while bytes are arriving.
Conceptually:
part_00001.uploading
↓ successful stream
part_00001.tmp
This helps distinguish incomplete part transfer from a completed part.
Bounded part streaming
Multipart uses a 128 KiB internal copy buffer.
It incrementally calculates SHA-256 while writing.
The entire part does not need to become one giant heap allocation.
Quota accounting
Quota checks include:
current committed bucket usage
- existing object being replaced
+ bytes already received by session
+ incoming part length
This matters because multipart data can consume real disk space before final object completion.
Ignoring in-progress parts would make quotas fictional.
Session TTL
The default session TTL is 24 hours.
A resumable upload cannot be immortal by accident.
Expired sessions need cleanup semantics because abandoned part files are still physical storage.
The concurrency lesson
Multipart performance is not:
more concurrency = more speed
It is closer to:
use enough concurrency to keep network + disk busy
without turning memory, metadata locks, descriptors,
or storage queues into the bottleneck
That is why the engine has both client-facing parallelism guidance and a server-side global limit.
Why this belongs inside the engine
A robust multipart implementation must coordinate:
- physical part files
- persistent session state
- integrity hashes
- quota
- capacity
- concurrency
- final object visibility
Leaving all of that to clients would make every SDK reinvent storage semantics.
The server needs to own the truth.
Top comments (0)