<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: AnyFileKit</title>
    <description>The latest articles on DEV Community by AnyFileKit (@anyfilekit).</description>
    <link>https://dev.to/anyfilekit</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4147356%2F1d43328f-3324-4163-aeea-77691929d84c.png</url>
      <title>DEV Community: AnyFileKit</title>
      <link>https://dev.to/anyfilekit</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/anyfilekit"/>
    <language>en</language>
    <item>
      <title>What broke when I moved my file tools into the browser</title>
      <dc:creator>AnyFileKit</dc:creator>
      <pubDate>Mon, 28 Sep 2026 14:38:55 +0000</pubDate>
      <link>https://dev.to/anyfilekit/what-broke-when-i-moved-my-file-tools-into-the-browser-2ckh</link>
      <guid>https://dev.to/anyfilekit/what-broke-when-i-moved-my-file-tools-into-the-browser-2ckh</guid>
      <description>&lt;p&gt;I built this because I had to compress a signed contract, and every "free online PDF compressor" I found wanted me to upload it first.&lt;/p&gt;

&lt;p&gt;So I set myself one rule for &lt;a href="https://anyfilekit.com/" rel="noopener noreferrer"&gt;AnyFileKit&lt;/a&gt;: the file never leaves the user's computer. No upload endpoint, no temporary storage, no "we delete it after an hour." Everything happens in the tab.&lt;/p&gt;

&lt;p&gt;The stack is boring on purpose. Static Astro pages on Cloudflare Workers, pdf-lib and pdf.js for PDFs, ffmpeg.wasm for video, tesseract.js for OCR, SheetJS for Excel. It now runs 23 tools.&lt;/p&gt;

&lt;p&gt;The boring stack was the easy part. Here's what actually broke.&lt;/p&gt;

&lt;h2&gt;
  
  
  The ffmpeg core was too big to deploy
&lt;/h2&gt;

&lt;p&gt;The single-threaded &lt;code&gt;ffmpeg-core.wasm&lt;/code&gt; is 30.7 MB. Cloudflare Workers static assets have a 25 MiB limit per file, so the first deploy was simply rejected.&lt;/p&gt;

&lt;p&gt;The obvious fix is loading ffmpeg from a public CDN. I didn't want that. "Your file stays on your machine" is much easier to believe when the page makes no third-party requests at all, and people do open the Network tab to check.&lt;/p&gt;

&lt;p&gt;So the wasm ships gzipped (about 9.8 MB) and the page decompresses it itself:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;gz&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;fetch&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;/ffmpeg/ffmpeg-core.wasm.gz&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;wasm&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Response&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="nx"&gt;gz&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;body&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;pipeThrough&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;DecompressionStream&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;gzip&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)),&lt;/span&gt;
&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;blob&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;wasmURL&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;URL&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;createObjectURL&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;wasm&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;ffmpeg&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;load&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;classWorkerURL&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;/ffmpeg/worker.js&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;coreURL&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;/ffmpeg/ffmpeg-core.js&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;wasmURL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="nx"&gt;URL&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;revokeObjectURL&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;wasmURL&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;DecompressionStream&lt;/code&gt; is in every current browser, so this needs no library. The download only starts when someone clicks "start" on a video tool. Nobody pays 10 MB just to look at the page.&lt;/p&gt;

&lt;p&gt;It's still the worst part of the experience. The first video job waits for that download.&lt;/p&gt;

&lt;h2&gt;
  
  
  I picked the single-threaded ffmpeg on purpose
&lt;/h2&gt;

&lt;p&gt;The multi-threaded ffmpeg build is faster. It also needs &lt;code&gt;SharedArrayBuffer&lt;/code&gt;, and that means sending COOP and COEP headers for the whole site. Cross-origin isolation blocks third-party scripts and iframes that don't opt in, and I wanted to keep that door open.&lt;/p&gt;

&lt;p&gt;So it's single-threaded. Slower, and I'm fine with that.&lt;/p&gt;

&lt;p&gt;One thing I expected to be a problem wasn't. I ran the same 1080p, 60-second compression twice on an M-series Mac in Chrome: once with the tab visible (114.9 s) and once with it hidden the whole time (112.9 s). ffmpeg runs in a worker, and Chrome doesn't throttle it when you switch tabs.&lt;/p&gt;

&lt;p&gt;pdf.js was a different story.&lt;/p&gt;

&lt;h2&gt;
  
  
  pdf.js freezes the moment you switch tabs
&lt;/h2&gt;

&lt;p&gt;Compressing a PDF means rendering each page to a canvas. The first version worked fine as long as you watched it. Switch to another tab mid-job, come back, and it was sitting on the same page it had been on when you left.&lt;/p&gt;

&lt;p&gt;The cause is one option. With the default &lt;code&gt;intent: 'display'&lt;/code&gt;, pdf.js moves through a page in small pieces scheduled with &lt;code&gt;requestAnimationFrame&lt;/code&gt;. Chrome pauses rAF in background tabs. So the render just stops, forever, with no error.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;render&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;canvasContext&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;viewport&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;intent&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;print&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;}).&lt;/span&gt;&lt;span class="nx"&gt;promise&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;'print'&lt;/code&gt; schedules its work with &lt;code&gt;setTimeout&lt;/code&gt;, which keeps running in the background. And turning a page into pixels is closer to printing than displaying anyway. That one word appears in three places in my code now, each with a comment so I don't "fix" it later.&lt;/p&gt;

&lt;h2&gt;
  
  
  Big files hit limits I didn't know existed
&lt;/h2&gt;

&lt;p&gt;There's no upload, so there's no server-side size cap. The limits are the user's browser and memory, and they show up in strange places.&lt;/p&gt;

&lt;p&gt;Chrome won't read more than 2046 MiB of a file into a single &lt;code&gt;ArrayBuffer&lt;/code&gt;. ffmpeg.wasm also keeps its input and output in an in-memory file system inside a 32-bit wasm address space. In my tests, a 1.67 GB output worked (heap peaked around 3.34 GB), and a 2.3 GB one threw &lt;code&gt;RangeError: Array buffer allocation failed&lt;/code&gt;. Because it's a normal exception, I catch it and tell the user to close some tabs or cut the file, instead of letting the page crash.&lt;/p&gt;

&lt;p&gt;For cutting video, I stopped using ffmpeg entirely. A quick cut or a "keep the original audio" export doesn't need decoding at all. The code reads just the &lt;code&gt;moov&lt;/code&gt; box (the MP4 index) with mp4box.js, working out which byte ranges hold the samples it needs. It builds the new file as a &lt;code&gt;Blob&lt;/code&gt; stitched together from &lt;code&gt;file.slice()&lt;/code&gt; pieces. The whole video is never loaded into memory at once, so neither limit applies.&lt;/p&gt;

&lt;p&gt;It starts small, too. Most MP4s keep &lt;code&gt;moov&lt;/code&gt; near the front, so it reads 1 MB, then 4, 16 and 64 before deciding. If it has read 256 MB and still found no index, the file probably isn't a clean MP4, and it goes back to ffmpeg.&lt;/p&gt;

&lt;h2&gt;
  
  
  PDF to Word is educated guessing
&lt;/h2&gt;

&lt;p&gt;This one isn't a bug. It's a limitation I had to make peace with.&lt;/p&gt;

&lt;p&gt;A PDF has no paragraphs, no headings and no reading order. It has glyphs drawn at coordinates. Every "PDF to Word" converter, mine included, is rebuilding structure from positions: which lines sit close enough to be a paragraph, where the columns are, whether that bigger bold line is a heading.&lt;/p&gt;

&lt;p&gt;Some things I just drop. Rotated and vertical text (book spines, watermarks, chart axis labels) has no sensible place in a flowing Word document, and forcing it in scrambles the body text. So the converter skips it and counts how much it skipped.&lt;/p&gt;

&lt;p&gt;I'd rather say that plainly than pretend the conversion is exact.&lt;/p&gt;

&lt;h2&gt;
  
  
  The one part that isn't local
&lt;/h2&gt;

&lt;p&gt;There are three AI translators on the site (PDF, image, video), and the models are far too big to run in a tab. So those do send data out. Still not the file, though. The browser pulls the text out of the PDF or image, or the audio track out of the video, and sends only that. The page says so right above the upload box.&lt;/p&gt;

&lt;p&gt;That's also the only paid part. Model calls cost real money. The 23 local tools cost me nothing to run, since your computer does the work.&lt;/p&gt;

&lt;h2&gt;
  
  
  If you want to poke at it
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://anyfilekit.com/compress-video/" rel="noopener noreferrer"&gt;Compress video&lt;/a&gt; (ffmpeg.wasm, single-threaded)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://anyfilekit.com/pdf-to-word/" rel="noopener noreferrer"&gt;PDF to Word&lt;/a&gt; (the guessing described above)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://anyfilekit.com/merge-pdf/" rel="noopener noreferrer"&gt;Merge PDF&lt;/a&gt; (pdf-lib)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://anyfilekit.com/ocr/" rel="noopener noreferrer"&gt;Image to text&lt;/a&gt; (tesseract.js)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Open DevTools on the Network tab while you use them. Beyond the page's own scripts and the wasm files, you shouldn't see anything go out.&lt;/p&gt;

&lt;p&gt;Disclosure: I built AnyFileKit, and it's a one-person project. I'm curious what other people have run into doing heavy processing client-side. Have you found a better way around the 2 GB read limit than slicing?&lt;/p&gt;

</description>
      <category>webassembly</category>
      <category>javascript</category>
      <category>webdev</category>
      <category>privacy</category>
    </item>
  </channel>
</rss>
