TL;DR: send a file in 5 MB chunks, let the server's offset decide where to continue, and do the slow final copy in a queue job on its own queue connection, with a retry_after longer than the job. That is the whole design, and it needs no package.
Uploading a 2 GB file through a normal PHP form is a bad idea. A proxy times out, PHP's limits get in the way, and one dropped connection at 95% means starting again. When I built a self-hosted file server (Laravel, Livewire, Flux UI) I wanted uploads that survive all of that, including a final copy to a remote disk (SFTP, S3) that can take minutes.
If you need interoperability with existing clients, the tus protocol is the standard and there are Laravel packages for it. What follows is a simpler, application-specific protocol that I could test end to end.
The protocol: four endpoints
All uploads, big or small, use the same flow, so there is no separate "simple" path to maintain.
-
POST /uploadswith the file name, size and a fingerprint. The server creates anuploadsrow and answers with anid, the currentoffset(0) and thechunk_size. Starting the same file again resumes the unfinished upload instead of creating a new one. -
PATCH /uploads/{id}with the raw chunk as the request body and anUpload-Offsetheader. -
GET /uploads/{id}returns the state:receiving,processing,doneorfailed. -
DELETE /uploads/{id}cancels.
Chunks are appended to a temporary file named after the upload id. The default chunk size is 5 MB, so each request only has to fit your proxy's body limit, not the whole file. upload_max_filesize stops mattering because the body is not a form upload, but keep post_max_size and your proxy's body limit above the chunk size.
The client: ask the server where you are
The browser slices the file and sends one chunk at a time. The key idea is that the client never trusts its own idea of progress. The server's offset is the truth.
const end = Math.min(state.offset + state.chunk_size, file.size);
const response = await fetch(`/uploads/${state.id}`, {
method: 'PATCH',
body: file.slice(state.offset, end),
headers: { 'Content-Type': 'application/octet-stream', 'Upload-Offset': String(state.offset) },
});
if (response.status === 409) {
// Out of step: the server says where it really is, and we carry on from there.
state = { ...state, ...(await response.json()) };
continue;
}
If a request fails, the client waits (1, 2, 4 seconds and so on, up to five retries), asks GET /uploads/{id} for the real offset and continues from it. That one rule gives you resume after a network drop, after a tab reload (the fingerprint, here the file's last-modified time, finds the unfinished upload) and after a duplicated request.
The server: one writer at a time
Two requests for the same upload must never write at the same time, and a chunk must only be accepted at the exact offset the server expects:
$lock = Cache::lock("upload:{$upload->id}", 120);
if (! $lock->get()) {
return response()->json($this->state($upload), 409);
}
try {
$upload->refresh();
$header = $request->header('Upload-Offset');
$offset = is_numeric($header) ? (int) $header : -1;
// Wrong offset, or already handed over: tell the client where we are.
if (! $upload->isReceiving() || $offset !== $upload->offset) {
return response()->json($this->state($upload), 409);
}
// append the chunk to the temporary file at that offset ...
} finally {
$lock->release();
}
A 409 carrying the real state is not an error to the client. It is how it finds its place. I also refuse a chunk that would go past the declared size (413), and every chunk request checks that the upload belongs to the current user.
Finishing: do the slow part in a queue job
When the last chunk arrives, the request does one thing: it marks the upload processing, queues a job and answers 202 at once. The job hashes the file (SHA-256), sniffs the MIME type from the content (the client's claim is ignored), copies it to its disk under a random key and creates the database record.
Why not in the request? On a remote disk the final copy can take minutes, and a proxy will cut the request long before that. Meanwhile the browser shows "Storing the file…", starts sending the next file, and polls GET /uploads/{id} (every second at first, slowing to five) until the state says done or failed.
The job has two unusual settings:
class FinalizeUpload implements ShouldQueue
{
public int $tries = 1; // a half-finished copy is cleaned up; the user retries
public int $timeout = 21600; // six hours
public bool $failOnTimeout = true;
public function __construct(public string $uploadId)
{
$this->onConnection('uploads');
}
}
The gotcha that cost me the most time: retry_after
A queue worker puts a job back on the queue if it has not finished within the connection's retry_after seconds. The default is 90. A copy that legitimately takes an hour would be picked up and started a second time while the first is still running.
The fix is a separate queue connection just for these jobs, with a retry_after larger than the job's timeout, and a worker started on that connection:
// config/queue.php
'uploads' => [
'driver' => 'database',
'table' => 'jobs',
'queue' => 'uploads',
'retry_after' => 21700, // above the job's 21600 s timeout
],
php artisan queue:work uploads --queue=uploads --tries=1 --timeout=0
The first uploads is the connection name, not the queue name. A worker started without it would use the default connection and its 90-second window.
Other things that bit me
-
Clean up on failure. Note the destination key on the upload row before you write, so a crash halfway can still remove the partial file. A failed job deletes the temporary file and any file that no database record uses, and marks the upload
failedwith a reason the browser can show next to a Retry button. - Quota. I check it when the upload starts and again, inside a transaction that locks the user's row, when the file is stored. Parallel uploads can each pass the first check, and the second check keeps the books right.
- Temporary space. The temp folder needs room for the largest file times the uploads in flight, and the target disk needs the file again.
- Deploys. Workers keep the old code in memory, so restart them after every deploy. A scheduled prune command removes unfinished uploads older than a day and gives up on jobs that died.
- Proxy limits. The proxy's body limit and timeout only have to cover one chunk. Long timeouts are for downloads, not uploads.
What I would still do differently
- If you need to resume across different clients and tools, use tus instead of a custom protocol.
- Hash on the client if you want to skip uploads of files you already have. I did not.
- I tested the remote path against one SFTP server (a Hetzner Storage Box) and a local OpenSSH server, not against a real S3 service, so treat that part as less proven.
The full implementation (controller, job, the browser store) is open source in the project where I built this: Ferrite. The upload design is written up in docs/uploads.md.
*Disclosure: I wrote the code. An AI assistant helped me draft and edit this article, and I checked every claim against the code.
Top comments (0)