<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Dixit Kallepalli</title>
    <description>The latest articles on DEV Community by Dixit Kallepalli (@deekshit989).</description>
    <link>https://dev.to/deekshit989</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4167557%2Ff9188f11-3576-49bb-b866-0393969e115e.png</url>
      <title>DEV Community: Dixit Kallepalli</title>
      <link>https://dev.to/deekshit989</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/deekshit989"/>
    <language>en</language>
    <item>
      <title>Real PDF redaction in the browser (a black box is not enough)</title>
      <dc:creator>Dixit Kallepalli</dc:creator>
      <pubDate>Wed, 07 Oct 2026 00:53:16 +0000</pubDate>
      <link>https://dev.to/deekshit989/real-pdf-redaction-in-the-browser-a-black-box-is-not-enough-d0e</link>
      <guid>https://dev.to/deekshit989/real-pdf-redaction-in-the-browser-a-black-box-is-not-enough-d0e</guid>
      <description>&lt;p&gt;Most "redact PDF" tools draw a black rectangle on top of the text and save. The text is still&lt;br&gt;
there. Select it, copy it, or open the file in any PDF library, and it comes straight back.&lt;br&gt;
Courts, journalists and HR teams have all leaked documents this way.&lt;/p&gt;

&lt;p&gt;I build &lt;a href="https://drible.co" rel="noopener noreferrer"&gt;Drible&lt;/a&gt;, a free set of file tools where most of the work happens in&lt;br&gt;
the browser. When I added redaction to the &lt;a href="https://drible.co/pdf/edit" rel="noopener noreferrer"&gt;PDF editor&lt;/a&gt;, I wanted it&lt;br&gt;
to be real redaction, done without uploading the file. Here is how it works and what I learned.&lt;/p&gt;
&lt;h2&gt;
  
  
  Why the black box fails
&lt;/h2&gt;

&lt;p&gt;A PDF page is a content stream: a list of drawing operators like "move here, show these glyphs,&lt;br&gt;
fill this path". Drawing a black rectangle just appends one more operator. The glyphs underneath&lt;br&gt;
are untouched, so:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;text extraction (and Ctrl+A, Ctrl+C) still returns them&lt;/li&gt;
&lt;li&gt;search still finds them&lt;/li&gt;
&lt;li&gt;the original objects can still be read from the file even if nothing on the page draws them&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So removing the visible text is not enough. You have to remove every object that could still&lt;br&gt;
hold it.&lt;/p&gt;
&lt;h2&gt;
  
  
  The approach: rebuild the page from pixels
&lt;/h2&gt;

&lt;p&gt;Rewriting a content stream to delete exactly the glyphs under a box is hard in general (text can&lt;br&gt;
be split across operators, use custom encodings, sit inside form XObjects, and so on). For a&lt;br&gt;
browser tool I chose a simpler guarantee: &lt;strong&gt;a page with redactions is replaced by an image of&lt;br&gt;
itself, with the boxes burned in.&lt;/strong&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Render the original page with PDF.js to a canvas at about 200 DPI (longest side capped at
5000 px so phones don't run out of memory).&lt;/li&gt;
&lt;li&gt;Fill each redaction box on the canvas in solid black.&lt;/li&gt;
&lt;li&gt;Encode the canvas as a JPEG and embed it on a fresh page of the same size with pdf-lib.&lt;/li&gt;
&lt;li&gt;Remove the old page.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The new page contains only an image, so there is nothing to select, search or extract under the&lt;br&gt;
boxes. The trade-off is that the rest of that page also becomes an image. Pages without&lt;br&gt;
redactions are left untouched, so their text stays selectable.&lt;/p&gt;
&lt;h2&gt;
  
  
  The part people miss: the old page is still in the file
&lt;/h2&gt;

&lt;p&gt;Removing a page in pdf-lib takes it out of the page tree, but its objects (content streams,&lt;br&gt;
fonts, images, annotations) can still be sitting in the file. A PDF reader won't show them,&lt;br&gt;
but anyone who opens the file in a text editor or a PDF library can. Two more steps close that:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Scrub the old page.&lt;/strong&gt; Before it is dropped, its content, resources and annotations are deleted&lt;br&gt;
from the page dictionary. Form fields whose widgets lived on that page are removed too, because a&lt;br&gt;
field's value can hold the very text you redacted.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Garbage-collect unreachable objects.&lt;/strong&gt; Walk the object graph from the trailer (the document&lt;br&gt;
catalog and the info dictionary), mark everything reachable, and delete everything else:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="cm"&gt;/** Delete every indirect object not reachable from the trailer. */&lt;/span&gt;
&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;collectGarbage&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;pdf&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;PDFDocument&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="k"&gt;void&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;ctx&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;pdf&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;context&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;seen&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nb"&gt;Set&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;stack&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;trailerInfo&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;Root&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;trailerInfo&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;Info&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;filter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;Boolean&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;while &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;stack&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;o&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;stack&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;pop&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;o&lt;/span&gt; &lt;span class="k"&gt;instanceof&lt;/span&gt; &lt;span class="nx"&gt;PDFRef&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;o&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;toString&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
      &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;seen&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;has&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;key&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="k"&gt;continue&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
      &lt;span class="nx"&gt;seen&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;add&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;key&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
      &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;target&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;lookup&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;o&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
      &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;target&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="nx"&gt;stack&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;push&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;target&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;o&lt;/span&gt; &lt;span class="k"&gt;instanceof&lt;/span&gt; &lt;span class="nx"&gt;PDFDict&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="k"&gt;for &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;[,&lt;/span&gt; &lt;span class="nx"&gt;v&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="k"&gt;of&lt;/span&gt; &lt;span class="nx"&gt;o&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;entries&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt; &lt;span class="nx"&gt;stack&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;push&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;v&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;o&lt;/span&gt; &lt;span class="k"&gt;instanceof&lt;/span&gt; &lt;span class="nx"&gt;PDFArray&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="k"&gt;for &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;v&lt;/span&gt; &lt;span class="k"&gt;of&lt;/span&gt; &lt;span class="nx"&gt;o&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;asArray&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt; &lt;span class="nx"&gt;stack&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;push&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;v&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;o&lt;/span&gt; &lt;span class="k"&gt;instanceof&lt;/span&gt; &lt;span class="nx"&gt;PDFStream&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="nx"&gt;stack&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;push&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;o&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="k"&gt;for &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;ref&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="k"&gt;of&lt;/span&gt; &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;enumerateIndirectObjects&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;seen&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;has&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;ref&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;toString&lt;/span&gt;&lt;span class="p"&gt;()))&lt;/span&gt; &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="k"&gt;delete&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;ref&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;pdf-lib then writes a clean file without the orphans. A quick way to test any redaction tool:&lt;br&gt;
redact a unique word, save, then run &lt;code&gt;pdftotext out.pdf - | grep yourword&lt;/code&gt; or open the file in a&lt;br&gt;
text editor and search for it. If it shows up, the redaction did not work.&lt;/p&gt;

&lt;h2&gt;
  
  
  Everything else stays real PDF
&lt;/h2&gt;

&lt;p&gt;Redaction is the only feature that turns a page into an image. Everything else the editor adds is&lt;br&gt;
written into the original file as real PDF content:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;text is real text in a standard PDF font, so it stays selectable and searchable&lt;/li&gt;
&lt;li&gt;shapes and freehand ink are vector paths&lt;/li&gt;
&lt;li&gt;images and signatures are embedded once and reused&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Because the original file is modified rather than re-created, its existing text, links and form&lt;br&gt;
fields survive.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why do it in the browser at all
&lt;/h2&gt;

&lt;p&gt;A file you redact is, by definition, a file with something sensitive in it. Uploading it to a&lt;br&gt;
server to remove the sensitive part is backwards. Doing it client-side means the PDF is opened,&lt;br&gt;
edited and saved on the user's device; the server only ever sends the JavaScript.&lt;/p&gt;

&lt;p&gt;48 of Drible's 53 tools work like this. The 5 that need desktop software (PDF compression with&lt;br&gt;
Ghostscript and the Office conversions) do upload, say so on the page, and delete the file as&lt;br&gt;
soon as the result is sent back. I wrote up &lt;a href="https://drible.co/guides/which-pdf-tools-upload-your-files" rel="noopener noreferrer"&gt;which popular PDF sites upload your files and how to&lt;br&gt;
check&lt;/a&gt; if you want the comparison.&lt;/p&gt;

&lt;p&gt;If you try the &lt;a href="https://drible.co/pdf/edit" rel="noopener noreferrer"&gt;editor&lt;/a&gt; and manage to recover redacted text from a&lt;br&gt;
file it saved, I'd really like to hear about it.&lt;/p&gt;

</description>
      <category>pdf</category>
      <category>javascript</category>
      <category>privacy</category>
      <category>webdev</category>
    </item>
  </channel>
</rss>
