DEV Community

Mahir Amaan
Mahir Amaan

Posted on

Document Management System: Fixing Large File Uploads

A legal Document Management System can work perfectly in development and still struggle when real users upload contracts, pleadings, exhibits, and scanned case files.

The failure usually starts when every large file passes through the application server. The browser sends the document to the API, the API processes the request, and the server becomes part of every file transfer.

As file sizes and concurrent uploads increase, request duration, bandwidth consumption, and server resource usage increase with them.

We faced this architecture decision while building an AI-powered legal Document Management System with secure document handling, version control, client portals, workflow automation, and document processing.

The solution was to separate document metadata from binary storage. The API handles authorization and metadata, while object storage handles the actual file transfer.

1. Separate document metadata from document bytes

The first problem was treating the document as a single database object.

A better Document Management System keeps business metadata in PostgreSQL while storing the actual files in object storage.

A simplified schema looks like this:

-- The Document Management System stores business state separately from file bytes.
CREATE TABLE documents (
    id UUID PRIMARY KEY,
    organization_id UUID NOT NULL,
    uploaded_by UUID NOT NULL,
    original_name TEXT NOT NULL,
    storage_key TEXT NOT NULL UNIQUE,
    mime_type TEXT NOT NULL,
    size_bytes BIGINT NOT NULL,
    status TEXT NOT NULL DEFAULT 'pending',
    created_at TIMESTAMPTZ NOT NULL DEFAULT NOW()
);
Enter fullscreen mode Exit fullscreen mode

The storage_key is the important part.

It gives the application a stable reference to the actual document without putting PDFs, DOCX files, or scanned images inside PostgreSQL.

The database can then manage states such as pending, uploaded, processing, ready, and rejected.

This separation also makes later document processing easier. A worker can scan or classify a file without changing the application's core document record.

2. Keep large uploads out of the API server

Once metadata and binary storage are separated, the upload architecture changes.

The naive approach looks like this:

Browser
   |
   | 20 MB PDF
   v
Node.js API
   |
   v
Object Storage
Enter fullscreen mode Exit fullscreen mode

The API now receives a file it does not actually need to process synchronously.

For a Document Management System handling large legal files, this creates unnecessary work for the application layer.

Instead, the API can generate a short-lived presigned URL:

Browser
   |                    \
   | metadata            \ 20 MB PDF
   v                      \
Node.js API               Object Storage
   |                          |
   +---- storage key ---------+
Enter fullscreen mode Exit fullscreen mode

The API authorizes the operation and generates the URL. The browser then sends the document directly to object storage.

For example:

// The Document Management System creates a short-lived URL for one storage object.
import { S3Client, PutObjectCommand } from "@aws-sdk/client-s3";
import { getSignedUrl } from "@aws-sdk/s3-request-presigner";

const s3 = new S3Client({
  region: process.env.AWS_REGION,
});

export async function createUploadUrl(
  storageKey: string,
  contentType: string
) {
  const command = new PutObjectCommand({
    Bucket: process.env.AWS_BUCKET!,
    Key: storageKey,
    ContentType: contentType,
  });

  return getSignedUrl(s3, command, {
    expiresIn: 300,
  });
}
Enter fullscreen mode Exit fullscreen mode

The client can then upload directly:

// The file payload bypasses the application server and goes directly to storage.
await fetch(uploadUrl, {
  method: "PUT",
  headers: {
    "Content-Type": file.type,
  },
  body: file,
});
Enter fullscreen mode Exit fullscreen mode

The important part is the responsibility split.

The API authorizes the upload. Object storage handles the file transfer. PostgreSQL records the document's business state.

That pattern keeps the Document Management System from turning every large upload into a long-running application request.

3. Validate files before making them trusted

Direct uploads solve one architectural problem, but they do not solve document security.

A Document Management System should not mark a file as trusted immediately after the browser reports a successful upload.

Instead, the document should move through explicit processing states:

pending
   |
   v
uploaded
   |
   v
scanning
   |
   +---- rejected
   |
   v
ready
Enter fullscreen mode Exit fullscreen mode

The API should validate the user's authorization, permitted file types, maximum file size, and generated storage key before issuing an upload URL.

A basic validation layer might look like this:

// Application validation runs before the Document Management System accepts the upload.
const ALLOWED_TYPES = new Set([
  "application/pdf",
  "application/vnd.openxmlformats-officedocument.wordprocessingml.document",
]);

const MAX_FILE_SIZE = 25 * 1024 * 1024;

if (!ALLOWED_TYPES.has(contentType)) {
  throw new Error("Unsupported document type");
}

if (sizeBytes > MAX_FILE_SIZE) {
  throw new Error("Document exceeds the maximum allowed size");
}
Enter fullscreen mode Exit fullscreen mode

This is only an initial validation layer.

A production Document Management System may also need malware scanning, content inspection, OCR, text extraction, and asynchronous classification.

The important distinction is between upload completion and document trust.

A file can exist in storage without being available to every user.

4. Keep business permissions in the database

After moving uploads directly to object storage, another question appears: who decides whether a user can access a document?

Storage permissions alone usually do not represent the complete business relationship.

In a legal Document Management System, access might depend on organization membership, matter assignment, client relationships, user roles, or document ownership.

Those rules belong in the application database.

For example:

-- Access is resolved from application membership, not from the storage URL.
SELECT d.id, d.storage_key
FROM documents d
JOIN matter_members mm
  ON mm.matter_id = d.matter_id
WHERE d.id = $1
  AND mm.user_id = $2
  AND d.status = 'ready';
Enter fullscreen mode Exit fullscreen mode

Only after this authorization check should the API generate a download URL.

This gives the Document Management System one authorization layer across web users, internal services, and API clients.

The same architecture can also fit broader ERP Solutions, where document workflows may eventually connect with CRM, HR, finance, procurement, or other business modules.

5. Model document versions explicitly

The upload problem becomes more interesting when users edit existing contracts.

Overwriting the same storage object makes audit history harder to manage. A better Document Management System creates a new version for each revision.

A storage structure could look like this:

organizations/
  {organizationId}/
    documents/
      {documentId}/
        versions/
          1/original
          2/original
          3/original
Enter fullscreen mode Exit fullscreen mode

The application can maintain the relationship through a separate table:

-- Each revision gets its own immutable storage reference.
CREATE TABLE document_versions (
    id UUID PRIMARY KEY,
    document_id UUID NOT NULL REFERENCES documents(id),
    version_number INTEGER NOT NULL,
    storage_key TEXT NOT NULL UNIQUE,
    size_bytes BIGINT NOT NULL,
    created_by UUID NOT NULL,
    created_at TIMESTAMPTZ NOT NULL DEFAULT NOW(),
    UNIQUE(document_id, version_number)
);
Enter fullscreen mode Exit fullscreen mode

This gives the Document Management System an explicit audit trail.

The application knows who created each version, when it was created, and which storage object represents it.

Storage-level versioning can provide another recovery mechanism, but it should not replace the business-level version model when authorship, approval status, or workflow history matters.

Real-World Application

The direct-upload decision came directly from the trade-off above: keeping the API responsible for authorization without making it responsible for transferring every document.

We implemented this architecture for an AI-powered legal Document Management System where users needed secure document storage, collaboration, document classification, contract lifecycle tracking, and workflow processing.

The application database stored document metadata, permissions, and processing state. Object storage handled the actual files, while background workers handled scanning and downstream processing.

Consider a 20 MB document.

With an application-mediated upload, the application server receives the 20 MB payload before storing or forwarding it. With direct upload, the browser sends that 20 MB payload directly to object storage.

The network transfer still happens, but the application server is removed from the file-transfer path.

That distinction becomes important as concurrent uploads increase.

It also gives the processing pipeline a cleaner handoff. Once the upload finishes, a worker can scan the file, extract text, classify the document, generate embeddings, or trigger an AI workflow without keeping the original HTTP request open.

For an AI-powered Document Management System, this separation is especially useful when document processing takes considerably longer than the initial upload.

Conclusion: What We Learned Building a Document Management System

  • A Document Management System should separate binary files from transactional metadata.
  • Large uploads should bypass the application server whenever the architecture permits it.
  • Presigned URLs let the API authorize uploads without carrying the file payload.
  • PostgreSQL should own permissions, document state, relationships, and business-level versions.
  • Object storage should handle document bytes while background workers handle scanning and processing.
  • Upload completion should not automatically make a document trusted.
  • A scalable Document Management System should keep authorization, storage, and processing responsibilities clearly separated.

If you are building a Document Management System, what approach are you using for large uploads, versioning, and asynchronous document processing? You can explore our Document Management System solutions or contact our team to discuss similar architecture challenges.

FAQ

Why use object storage in a Document Management System?

Object storage is designed for binary files, while the application database can focus on metadata, permissions, workflows, and relationships.

Should documents be stored directly in PostgreSQL?

It is possible, but separating document bytes from application metadata can simplify large-file transfers, storage management, and asynchronous processing.

Are presigned URLs enough for security?

No. A Document Management System should still authorize the user, restrict upload scope, validate file metadata, limit file size, and apply appropriate content or malware scanning.

How should a Document Management System handle document versions?

Treat every revision as a separate version with its own storage reference and application-level metadata. This preserves authorship, timestamps, workflow state, and audit history.

Top comments (0)