<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Rafay Rajput</title>
    <description>The latest articles on DEV Community by Rafay Rajput (@rafay_rajput_).</description>
    <link>https://dev.to/rafay_rajput_</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4174004%2Fbd4a6a1a-5faf-41af-9f8f-ce69b88ad985.png</url>
      <title>DEV Community: Rafay Rajput</title>
      <link>https://dev.to/rafay_rajput_</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/rafay_rajput_"/>
    <language>en</language>
    <item>
      <title>PDF</title>
      <dc:creator>Rafay Rajput</dc:creator>
      <pubDate>Sat, 10 Oct 2026 14:51:58 +0000</pubDate>
      <link>https://dev.to/rafay_rajput_/pdf-feb</link>
      <guid>https://dev.to/rafay_rajput_/pdf-feb</guid>
      <description>&lt;p&gt;&lt;strong&gt;Tags:&lt;/strong&gt; &lt;code&gt;pdf&lt;/code&gt;, &lt;code&gt;javascript&lt;/code&gt;, &lt;code&gt;performance&lt;/code&gt;, &lt;code&gt;webdev&lt;/code&gt;, &lt;code&gt;privacy&lt;/code&gt;&lt;br&gt;
&lt;strong&gt;Title:&lt;/strong&gt; &lt;em&gt;The PDF compression floor: why "compress to exactly 200KB" usually can't be done&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;Most PDF tools implement &lt;em&gt;generic&lt;/em&gt; compression: apply a preset level, report whatever comes out. That's correct for most use cases and useless when there's a hard upload cap.&lt;/p&gt;

&lt;p&gt;I hit this building an exact-size compressor. The interesting part is the constraint, not the code.&lt;/p&gt;

&lt;h2&gt;
  
  
  The floor is real and it's arithmetic
&lt;/h2&gt;

&lt;p&gt;Four mechanisms reduce PDF size:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Mechanism&lt;/th&gt;
&lt;th&gt;Lossy?&lt;/th&gt;
&lt;th&gt;Typical gain&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Image downsampling&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Halving dimensions = 75% fewer pixels&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;JPEG quality reduction&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Large, but degrades below a floor&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Font/resource deduplication&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;No&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Frequently 20–40% on text documents&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Object stream optimisation&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;No&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Modest, free, always worth doing first&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The two lossless ones should run first. On text-heavy documents they often get you most of the way with zero quality cost.&lt;/p&gt;

&lt;p&gt;The two lossy ones interact. Downsample to hit 400KB, then reduce JPEG quality to reach 200KB, and you've made a signature unreadable.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why it fails on real documents
&lt;/h2&gt;

&lt;p&gt;Arithmetic for a 10-page scan at 300dpi:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;300dpi A4 ≈ 2480 × 3508 px ≈ 8.7 megapixels per page&lt;/li&gt;
&lt;li&gt;At ~0.3 bytes/pixel compressed, that's roughly 2.6MB per page&lt;/li&gt;
&lt;li&gt;10 pages ≈ 26MB&lt;/li&gt;
&lt;li&gt;To reach 200KB total: &lt;strong&gt;20KB per page&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;At the same compression ratio: ~65,000 pixels ≈ &lt;strong&gt;255 × 255&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;At 255×255 an A4 page, body text is ~4 pixels tall. Unreadable.&lt;/p&gt;

&lt;p&gt;This is why "compress PDF to 200KB" is often not achievable, and why a tool that claims otherwise is either degrading the document or lying about the result.&lt;/p&gt;

&lt;h2&gt;
  
  
  The honest behaviour
&lt;/h2&gt;

&lt;p&gt;A correct implementation needs three outcomes, not one:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Target met&lt;/strong&gt; — iterate quality/resolution until the budget is satisfied&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Floor reached&lt;/strong&gt; — report &lt;em&gt;how close&lt;/em&gt;, and say the target is unreachable, with the reason&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Not worth degrading&lt;/strong&gt; — stop when legibility breaks, and explain the tradeoff&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Outcome 2 is the one everyone skips and the one that actually matters to a user standing in front of a rejected upload.&lt;/p&gt;

&lt;h2&gt;
  
  
  Architecture note
&lt;/h2&gt;

&lt;p&gt;All of this is arithmetic, and it can run entirely client-side. &lt;code&gt;pdf-lib&lt;/code&gt; plus a canvas pass handles resampling; JPEG re-encoding is &lt;code&gt;canvas.toBlob&lt;/code&gt; with a quality parameter.&lt;/p&gt;

&lt;p&gt;The commercially interesting part: &lt;strong&gt;none of it requires an upload.&lt;/strong&gt; The same work a server would do happens in the browser, which means a document with financial or identity details never leaves the device.&lt;/p&gt;

&lt;p&gt;I wrote up the full method, including per-tool behaviour and the verification step: &lt;a href="https://pdff.online/guides/compress-pdf-to-exact-size" rel="noopener noreferrer"&gt;https://pdff.online/guides/compress-pdf-to-exact-size&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Dev.to / Hashnode:&lt;/strong&gt; same text works. Cross-post to both, canonical to the pdff guide.&lt;/p&gt;

</description>
      <category>javascript</category>
      <category>performance</category>
      <category>software</category>
      <category>webdev</category>
    </item>
  </channel>
</rss>
