DEV Community

Shahzaib
Shahzaib

Posted on

Split by Size: The PDF Feature Nobody Else Has (And How I Built It)

Series: Building PdfWord — a free, no-backend PDF tools site (Part 13)

Every PDF splitter in the world works the same way: pick page ranges, or split into fixed page counts. But the actual problem people have is different. It's not "give me pages 1–10." It's "my PDF is 68MB and Gmail won't send it." Email caps at 25MB. WhatsApp gets flaky with big files. The limit that matters is megabytes, not pages.

So I built Split by Size: tell it a target size (under 10MB, 25MB, or 50MB), and it packs pages into the fewest files that each fit under that limit. No other free tool I know does this.

Try it: Split by Size


The problem with page-based splitting

A 50-page PDF of scanned images might be 80MB. The same 50 pages of plain text might be 800KB. "Split every 10 pages" gives you eight 10MB chunks in one case and eight 100KB chunks in the other — wildly inconsistent, and the second case didn't need splitting at all. Pages are a meaningless unit for the size problem.

What you actually want is a bin-packing problem: pack pages into bins (files) such that no bin exceeds the target size, and you use as few bins as possible. That's the whole feature.

Estimating page sizes without rendering everything

Here's the catch: to know a page's size, you have to serialize it. pdf-lib's PDFDocument.copyPages() + save() per page would be accurate but slow — save a 200-page document page-by-page and the user watches a progress bar for a minute.

My approach: estimate each page's size from its content streams and embedded images (the PDFPage node's byte length in the original file is a decent proxy), then do a greedy pack: walk pages in order, accumulate estimated size, and start a new chunk when adding the next page would exceed the target. Then verify each chunk by actually serializing it — if a chunk overshoots (estimates aren't perfect), split it further.

// Greedy bin packing: pages stay in order, chunks never exceed target
const chunks = [[]];
let running = 0;
for (const p of pagesWithSize) {
  if (running + p.size > TARGET && chunks[chunks.length - 1].length > 0) {
    chunks.push([]);  // start a new chunk
    running = 0;
  }
  chunks[chunks.length - 1].push(p);
  running += p.size;
}
Enter fullscreen mode Exit fullscreen mode

The greedy choice keeps pages in reading order — a fancier optimal-packing algorithm could use fewer chunks, but it would shuffle your document. Nobody wants page 47 of their contract in chunk 1.

The verification pass (because estimates lie)

Size estimation from content streams is approximate. An image shared between pages (like a letterhead logo) gets counted twice in estimates but stored once in the real output. Font subsets can grow or shrink depending on which pages land together. So after packing, I serialize each chunk with pdf-lib and check its actual byte size. Overshoots get split further; tiny last chunks get merged into the previous one if they'd be silly-small.

This two-pass design — fast estimate, then verify — is the whole trick. One pass is fast but wrong; pure per-page serialization is right but slow. Estimate-then-verify is fast and right.

The honest limitations

  • It's still page-granular. If a single page is bigger than your target (one giant scanned blueprint), no packing algorithm can help — the tool tells you so instead of producing a "chunk" that violates the limit.
  • Estimates can undershoot on weird files. Files with lots of shared resources pack tighter than estimated, which is harmless (smaller chunks). Overshoots are caught by the verification pass.
  • Compression first is usually smarter. An 80MB scanned PDF splits into eight 10MB chunks — or compresses into one 12MB file. The tool links to the compressor for exactly this reason.

Why nobody else has this

Honestly? Because it's a weird feature. It doesn't demo well in a feature grid ("Split: ✓"). It only matters in the exact moment you're staring at a "file too large" error. But that moment is miserable, and every existing tool answers it with page ranges — which is the wrong unit entirely.

The best features are like this: unsexy, un-demoable, and exactly what you need at 11pm when a deadline is attached.

Try it: Split by Size — set your target, get chunks that actually fit.

What's the most annoying file-size limit you've ever hit? Email's 25MB? Some government portal's 5MB? I'm collecting war stories.

Top comments (0)