Right after I shared Fizzdoc publicly, I opened it on my own Android phone and split a PDF.
It took six minutes.
Six. On my laptop the same kind of job took about a second. The link was already out in public, so somewhere out there, people were staring at a progress bar that didn't exist and thinking "nice project, Rishab."
I'll tell you how I fixed it. But first, what this thing even is, and why I think it fits this year's Hacktoberfest theme, "AI belongs to everyone", better than I planned.
What is Fizzdoc?
Fizzdoc is 64 tools for PDF, Word, Excel, PowerPoint, images and audio. Merge, compress, convert, edit, redact, OCR, cut an MP3, turn a voice note into text. The usual suspects, in 16 languages.
None of it uploads your file, though. There's no backend. Not a small one, not a "we only keep it for an hour" one. Nothing. It's a static site, and every job runs inside your browser tab, on your own device.
Most "free PDF" sites work like this: you send them your passport scan (or your Aadhaar card, if you're in India), they send you back a slightly smaller one, and you hope for the best. That always felt like a strange deal for jobs a phone can do on its own. So I built the other version.
Oh, and it runs an open-weight AI model (Whisper) right in the tab to turn audio into text. No API key, no GPU, no server. If AI belongs to everyone, it should also run on everyone's phone.
It's open source (Apache-2.0): github.com/kingrishabdugar/fizzdoc. Everything in this post is in that repo, and every number comes from the code or something I actually measured.
Where Fizzdoc fits
The open-source world already has great PDF tools, and I've used several. Some are big platforms you self-host with Docker and run for a whole team, with pipelines, APIs and logins. Some are lean, browser-only PDF kits. Both are good at what they do.
Fizzdoc makes a different bet:
- Nothing to install. Not even Docker. Open a link on any phone or laptop and it works. That matters when the person who needs to unlock a PDF is your dad, not your homelab.
- Not just PDF. 30 PDF tools, 14 for images, 14 for Word, Excel and PowerPoint, and 6 for audio, including speech to text and subtitles. Plus developer odds and ends like Excel to JSON and Mermaid to PNG.
- Your language, not just English. 16 languages, from Spanish, French, German and Portuguese to Hindi, Bengali, Tamil and Telugu, each with its own URLs, so people can find it by searching the way they actually talk.
- Built for cheap phones first. More on that below, because I learned it the hard way.
- Privacy you can verify. A strict CSP and a test that fails if any request leaves the site.
If you need automation pipelines or an admin panel for your company, a self-hosted platform is the right call. If you need to shrink a photo to 50 KB for a government form on your phone in the next 30 seconds, that's what this is for.
The whole thing on one slide
Three layers, and only the last one ever touches your file.
Build time. One TypeScript file, src/site.ts, describes all 64 tools: slug, formats, options, page text. A renderer turns that into about 1,050 plain HTML pages (64 tools × 16 languages, plus a few extras), with the sitemap and llms.txt thrown in. Adding a tool is one registry entry, one engine function and tests. The pages write themselves, which is the only kind of page I enjoy writing.
Hosting. Cloudflare Pages. It serves files and answers GET requests. That's its entire personality.
Your browser. A small script handles the boring parts: picking files, options, progress, download. It has no idea what a PDF is. When you press the button, it borrows a specialist:
| Engine | Job | Download (gzipped) |
|---|---|---|
| qpdf (C++, compiled to WebAssembly) | merge, split, rotate, passwords | ~440 KB |
| pdf.js | drawing pages, reading text | ~390 KB |
| pdf-lib | page numbers, watermarks, building PDFs | small |
| Tesseract.js | OCR | ~5.8 MB |
| Whisper tiny | speech to text | ~58 MB, once |
| fflate | Word, Excel, PowerPoint (secretly ZIP files) | small |
| Mermaid | diagrams to PNG/SVG | per diagram type |
"Just use a server, bro"
I know. A server would have been easier. I chose pain for three reasons.
You can check it. "We delete your files" is a promise. This is a rule your browser enforces:
default-src 'self';
script-src 'self' 'wasm-unsafe-eval';
connect-src 'self';
img-src 'self' data: blob:;
object-src 'none'; base-uri 'none'; form-action 'none'
connect-src 'self' means the page can only talk to fizzdoc.com, and fizzdoc.com only serves the website. Open DevTools, go to Network, run a job, and watch nothing interesting happen. The test suite does the same thing on every pull request and fails if a single request goes anywhere else.
It's free to run. A file-processing server gets more expensive with every user. A static site doesn't care how popular it gets. That's the only reason I can say "free" without a pricing page hiding behind it.
No waiting in line. No upload bar, no "processing on our servers", no download. Your file goes from your disk to your RAM and back.
The catch is obvious: your device does the work, and not everyone has a laptop. Which brings us back to my six-minute disaster.
The Android bug
qpdf runs on Emscripten, which gives it a small pretend file system. I was handing it your file through WORKERFS, which reads the file lazily, only when qpdf asks for a piece. Sounds efficient, very "don't load what you don't need", and I was proud of it.
What I didn't know: qpdf doesn't ask for a few big pieces. It asks for lots and lots of tiny ones. On a 10 MB, 28-page test PDF, I counted 1,448 separate reads: 434 just to count the pages, and 1,014 more for the split.
On a laptop, 1,448 small reads is nothing. On Android, reading a picked file in small pieces is slow, and 1,448 slow things in a row add up to minutes. My "efficient" design was death by a thousand paper cuts (literally, since it was a PDF).
The fix was almost embarrassing:
// qpdf reads its input in thousands of small pieces, and on Android every small read
// of a picked file is slow. Only inputs too big to copy safely are still read lazily.
const COPY_UP_TO = 128 * 1024 * 1024;
if (totalSize <= COPY_UP_TO) await copyIn(q, files, run); // read once, in order, into memory
else q.FS.mount(q.WORKERFS, { blobs }, '/in'); // huge files: no second copy
If your files add up to 128 MB or less, the worker reads them once, start to finish, into memory. All those tiny reads now hit RAM, which doesn't care. Same test file: 1,448 reads became zero, and even on my laptop the split went from 1.4 s to 0.25 s. So the bug was slowing everyone down; phones were just louder about it.
Above 128 MB, making a second full copy could run a phone out of memory, so big files still get the lazy path. Slower, but it finishes. I tested that with a 140 MB PDF, and yes, I waited.
Being nice to a ₹8,000 phone
That bug taught me to design for the cheapest phone first, not my laptop. This is what changed because of it.
Load almost nothing
The first page load is about 15 KB of gzipped JavaScript and 7 KB of CSS. No framework. Every page is prerendered HTML, so you can read it before any script runs (Google can too).
The Inter font is split into 7 unicode-range pieces with font-display: swap, so your browser only downloads the characters a page actually uses, and shows text in a system font while it waits. A tiny theme.js sets dark mode before the first paint, so no white flash at 2 a.m. Hashed files are cached for a year, so the second visit is mostly just HTML.
Load engines only when asked
Each tool's engine hides behind a dynamic import():
'compress-pdf': async (f, o) => (await import('./engine/compress')).compressPdf(f[0], { ... }),
'ocr-pdf': async (f) => (await import('./engine/ocr')).ocrPdf(f[0], progressLabel),
If all you do is rotate a PDF, you never download the OCR engine or a 58 MB speech model. Mermaid is even pickier: every diagram type is its own chunk, so drawing a flowchart doesn't drag in the code for sequence diagrams.
Check the budget before you spend it
Before a job starts, the page does a quick sanity check:
// 128 MiB per GiB of reported device memory, capped at 1 GiB.
function inputBudget() {
const gib = navigator.deviceMemory;
return gib ? Math.min(gib * 128, 1024) * MiB : 512 * MiB;
}
A 4 GB phone gets 512 MiB, an 8 GB laptop gets 1 GiB. Firefox and Safari don't report device memory, so they get a sensible 512 MiB. The cap exists because qpdf's WebAssembly memory tops out at 2 GiB and a job needs room for more than its input.
Is it a perfect measurement? No, and the code comment admits that. But "this file is too big for this device" is a much better message than a tab that quietly dies after five minutes.
Show real progress
A spinner on a slow phone is basically a "maybe" button. So the progress bar is real.
The first 30% is the file being read into memory, byte by byte. The other 70% comes from qpdf itself. qpdf has a --progress flag that prints lines like write progress: 42%. This WebAssembly build ignores the option meant to catch that output, so it just goes to console.log. Okay then, I'll grab it from there:
console.log = (...args) => {
const pct = /: write progress: (\d+)%$/.exec(String(args[0]))?.[1];
if (pct === undefined) return log(...args);
post({ type: 'progress', fraction: READ_SHARE + ((1 - READ_SHARE) * Number(pct)) / 100 });
};
Is hijacking console.log elegant? Absolutely not. Does the bar move now? It does. I sleep fine.
One job, one worker, then gone
Every qpdf job gets its own fresh Web Worker, so the page stays smooth while it works. When the job ends, or you hit Cancel, the page calls worker.terminate(). That one line throws away the file, any password you typed, and the engine's whole memory, so the cancel button doubles as cleanup and leak prevention. No state survives between jobs, so memory doesn't slowly creep up while you merge your way through tax season.
Before running, the worker also asks qpdf what's in the file: bookmarks, forms, signatures, encryption. If you merge a signed PDF, Fizzdoc tells you the signature won't be valid anymore, instead of letting you find out from your CA.
Only draw what's on screen
The editor and the redaction tool show every page of your PDF. On a 200-page file that could mean 200 canvases fighting for memory. Instead, a page is drawn only when it's about to scroll into view:
const io = new IntersectionObserver(
(entries) => { for (const e of entries) if (e.isIntersecting) renderPage(pages[e.target.dataset.index]); },
{ root: scroller, rootMargin: '600px 0px' },
);
Tools that must touch every page (PDF to JPG, scanned PDF, OCR) go one page at a time and clean each page up before moving on.
Do less work
- MP3 and M4A are cut without re-encoding. An MP3 is a row of little frames, each saying how long it is. Cutting is just picking the frames between two timestamps. No decoding, no quality loss, and fast even on a phone.
- Compression never makes things bigger. Each image inside a PDF is kept only if the new version is smaller. If the whole file doesn't shrink, you get your original back, not a "compressed" file that's 3% larger. (We've all seen that one.)
- "Compress image to 50 KB" is smart about it. It keeps the photo large and lowers quality first, never below 50% while it's 1000 px or wider, using a 6-step binary search. Only if that can't fit does it shrink the size. Roughly how messaging apps do it.
And when it's slow anyway, say so
While a job runs, the page tells you speed depends on your device. If something takes over 20 seconds or fails, a "Slow or not working? Report it on GitHub" link lights up. It opens a ready-made issue with the tool, browser, time taken, file size, and the device's RAM and CPU cores. Never the file, never its name. That link is how I find the next bug instead of guessing.
What's still not great: PDF to JPG keeps every page image in memory until the ZIP is built (a 1,000-page PDF on a phone would be a bad day), some jobs still run on the main thread, and Whisper runs on a single thread. Speaking of which.
Speech to text, without a server
This was the feature I wasn't sure a static site could pull off.
Fizzdoc runs Whisper tiny (8-bit quantized) with transformers.js and ONNX Runtime, on the CPU, inside a worker. It took three fights to get there.
Cloudflare Pages allows 25 MiB per file. The model is bigger. So a build script chops the model into parts. The first time you use the tool, the browser downloads the parts, glues them back together and keeps them in Cache Storage. After that it starts in seconds, even offline.
My own security rules blocked it. ONNX Runtime likes to load its WebAssembly through a blob: URL, which my CSP refuses. So the worker reads the file from the cache itself and hands it over directly. Remote model loading is switched off too, so it can only read from fizzdoc.com.
Nobody trusts a spinner for 58 MB. The first run says "Preparing… one-time setup: 23 of 58 MB" with a real byte counter, and tells you it only happens once.
Audio is turned into 16 kHz mono (what Whisper expects) and fed in 30-second windows with 5 seconds of overlap, so words at the edges don't get chopped in half.
Whisper tiny is, well, tiny. Clear English comes out well. Hinglish, sadly, does not. Give it Hindi and English in the same sentence and it answers with great confidence and very little accuracy. The tool now says it works best on clear English. Multi-threading would help speed, but it needs cross-origin isolation, and my first attempt made the worker hang. So it's single-threaded for now. If you've solved this, please come find me.
Smaller decisions I'd defend in a code review
Redaction that actually redacts. Drawing a black box over text doesn't remove it. People still do that, and people still copy the text out from under the box. Fizzdoc turns each page into an image with the boxes burned in, then builds a new PDF from those images. The text is really gone. The trade-off is the result isn't selectable text, and the tool says so. There's also an optional "find personal info" button that marks emails, UPI IDs, phone numbers, PAN, IFSC, IBAN and long ID numbers. It's plain regex, not AI, and you check every mark before saving.
Editing text in place, in the PDF's own font. Click a line in a PDF and retype it. On save, Fizzdoc parses the page's content stream, tracks the text state (matrices, font, spacing) to find the exact operators that draw that line, and rewrites them using the font that's already embedded in the file. Same font, same size, same colour, no white box. The catch: embedded fonts are usually subsets, so you can only use letters the file already contains. Change "Rs 4,250" to "Rs 2,450"? Easy. Type a "Z" into a PDF that never had one? The glyph simply isn't there, so that line falls back to covering the old text and drawing new text in the closest standard font. Hindi and Arabic always take the fallback too, because their glyphs are drawn in a different order than they're typed, and I'd rather be boring than garbled.
Word to PDF via the print dialog. Drawing Hindi, Tamil or Arabic correctly inside a JavaScript PDF writer is a whole career. Every browser already ships a great PDF writer: Print → Save as PDF. So Fizzdoc turns your .docx into clean HTML and opens the print dialog. One extra click, and every script comes out right. Lazy, yes, but it works.
qpdf instead of a JavaScript PDF library. I started with pdf-lib for merging. It works, but copying pages that way drops bookmarks, form fields and internal links. qpdf rewrites the file's structure instead, and has been doing it reliably for years.
Bonus: your coding agent can use it too
Hacktoberfest is big on agents this year, so here's a small one. The repo has an MCP server (in /mcp) that runs the same qpdf and audio engines locally. Hook it up to Claude Code, Codex or Cursor and you can say "merge these PDFs" or "cut the first 30 seconds of this MP3", and the agent does it on your machine. No upload, no throwaway Python script, and it never overwrites your files.
Lots of languages, one source file
Each language has real URLs (/hi/merge-pdf/, /ta/compress-pdf/), translated buttons, FAQs and error messages, and hreflang links between them: Hindi, Bengali, Marathi, Tamil, Telugu, Spanish, Portuguese, French, German, Italian, Dutch, Polish, Turkish, Indonesian, Vietnamese and English.
Because every page is generated, the tests can check all ~1,050 of them for their CSP, title, description, canonical link, a single h1 and valid structured data. I'm not checking 1,050 pages by hand.
Tests, because I like sleeping
209 unit tests and 1,087 browser tests. Engines are plain functions (bytes in, bytes out), so most run in Node against small weird files: encrypted PDFs, forms, bookmarks, MP3s. The browser tests click through the real UI, including the "no request leaves the site" check. CI runs both on every pull request, and nothing merges red.
Come build it with me
Right now Fizzdoc is one person and a lot of tabs. There's plenty left, and some of it is really fun engineering. Hacktoberfest doesn't count pull requests anymore, which honestly suits me. I'd much rather have one good PR that someone understands than ten that just exist.
Good first issues
- Crop PDF, Flatten PDF, Split every N pages
- Pick the DPI for PDF → JPG, watermark position and colour
- Word → Markdown / HTML
- Translations: Japanese, Korean, Gujarati, Kannada, Malayalam, Punjabi
Bigger ones
- Sign PDF
- OCR in more languages, downloaded on demand like Whisper
- Install it as an offline app (PWA)
- Page thumbnails with drag and drop
- WAV, FLAC, OGG and OPUS in the audio tools
- Remove image background, on the device
- Right-to-left languages
Problems I'd love your brain on
- Faster Whisper: threads (COOP/COEP without the hang) or WebGPU. If you've shipped either, tell me everything.
- A bigger or multilingual open-weight speech model that still fits a phone, so Hinglish stops being a comedy show.
- More MCP tools: the image, OCR, redaction and Office engines in Node.
- Streaming the ZIP for PDF to JPG, so huge PDFs don't sit in memory.
- Real text deletion in Edit PDF, rewriting the content stream instead of covering it.
Setup takes about as long as reading this sentence twice:
git clone https://github.com/kingrishabdugar/fizzdoc
cd fizzdoc
npm install
npm run dev
CONTRIBUTING.md shows how a tool goes from the registry to its engine. Small pull requests get reviewed fastest. Not sure where to start? Comment on any issue and I'll help you pick one.
kingrishabdugar
/
fizzdoc
Free, private PDF, image, audio, Word, Excel & PowerPoint tools that run 100% in your browser. Compress, convert, edit, redact, OCR, cut MP3, Excel↔JSON, Mermaid to PNG/SVG. No uploads, no sign-up, no watermark. Free forever · 64 tools · 16 languages.
Free, private tools for PDF, images, audio, Word, Excel and PowerPoint that never see your files.
64 tools — compress, convert, edit, redact, merge, split, cut, OCR, transcribe — running 100% in your browser.
No uploads. No sign-up. No watermark. No limits. Free forever. 16 languages.
Open source. Edits PDF text in the file’s own font, transcribes audio with Whisper on your device, and never hands your files to anyone.
Open Fizzdoc → · All tools · How it works · Contribute
✨ Why Fizzdoc
Most free online file tools ask you to upload your contract, payslip, bank statement, passport photo or voice recording to someone else's server. Fizzdoc doesn't have a server to upload to. Everything runs inside your browser tab with WebAssembly and JavaScript, and the page's Content Security Policy blocks it from connecting to any other site — you can check it yourself in the Network…
That's it
Try it at fizzdoc.com. Throw your weirdest PDF at it. If it's slow on your phone, hit the report link and tell me which phone, so I can have another humbling morning.
And if you think it's useful, a star on kingrishabdugar/fizzdoc helps more people find it. It also makes me unreasonably happy.
Thanks for reading.









Top comments (0)