<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Saffi Ullah</title>
    <description>The latest articles on DEV Community by Saffi Ullah (@saffi_ullah_706aebeff15ca).</description>
    <link>https://dev.to/saffi_ullah_706aebeff15ca</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4162831%2F5c8d55a2-d2d4-4372-9ac8-69af04719083.png</url>
      <title>DEV Community: Saffi Ullah</title>
      <link>https://dev.to/saffi_ullah_706aebeff15ca</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/saffi_ullah_706aebeff15ca"/>
    <language>en</language>
    <item>
      <title>Why Editing Existing PDF Text Is Much Harder Than It Looks</title>
      <dc:creator>Saffi Ullah</dc:creator>
      <pubDate>Mon, 05 Oct 2026 05:11:10 +0000</pubDate>
      <link>https://dev.to/saffi_ullah_706aebeff15ca/why-editing-existing-pdf-text-is-much-harder-than-it-looks-3568</link>
      <guid>https://dev.to/saffi_ullah_706aebeff15ca/why-editing-existing-pdf-text-is-much-harder-than-it-looks-3568</guid>
      <description>&lt;h1&gt;
  
  
  Why Editing Existing PDF Text Is Much Harder Than It Looks
&lt;/h1&gt;

&lt;p&gt;Editing text sounds simple.&lt;/p&gt;

&lt;p&gt;You open a document, find a sentence, change a few words, and save it.&lt;/p&gt;

&lt;p&gt;That's how most of us think about editing a document.&lt;/p&gt;

&lt;p&gt;But PDFs don't really work that way.&lt;/p&gt;

&lt;p&gt;When you change existing text in a PDF, you aren't necessarily changing a simple piece of text stored inside a paragraph. You're working with fonts, glyphs, character mappings, positioning, content streams, and other parts of the PDF that work together to produce what you see on the screen.&lt;/p&gt;

&lt;p&gt;That's why a PDF editor can sometimes extract text perfectly but still struggle to change that same text without affecting the document's appearance.&lt;/p&gt;

&lt;p&gt;The deeper I got into PDF editing, the more I realized that &lt;strong&gt;editing PDF text is much closer to modifying a rendering system than editing a normal text document.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  A PDF Isn't a Word Document
&lt;/h2&gt;

&lt;p&gt;When you work with a Word document, you can think about the document in fairly familiar terms:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Document
├── Paragraph
│   ├── Text
│   └── Formatting
├── Paragraph
└── Table
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The application understands things such as paragraphs, text runs, styles, and tables.&lt;/p&gt;

&lt;p&gt;A PDF works at a much lower level.&lt;/p&gt;

&lt;p&gt;A simplified way to think about a PDF page is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;PDF Page
   │
   ├── Content stream
   │      ├── Text operations
   │      ├── Graphics operations
   │      └── Positioning
   │
   └── Resources
          ├── Fonts
          ├── Images
          └── Other objects
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The PDF contains instructions that tell a renderer what to draw on the page.&lt;/p&gt;

&lt;p&gt;So a sentence that looks like this:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The annual report was published in 2026.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;doesn't necessarily exist inside the PDF as one nice, editable string.&lt;/p&gt;

&lt;p&gt;It may be split across several text operations, with separate information describing the font, position, encoding, and other details.&lt;/p&gt;

&lt;p&gt;That is where things start getting complicated.&lt;/p&gt;

&lt;h2&gt;
  
  
  Text Extraction Is Not the Same as Text Editing
&lt;/h2&gt;

&lt;p&gt;This is one of the biggest things I learned while working with PDFs.&lt;/p&gt;

&lt;p&gt;A PDF library might extract:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;The annual report was published in 2026.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;without any problem.&lt;/p&gt;

&lt;p&gt;But that doesn't mean it can safely change it.&lt;/p&gt;

&lt;p&gt;Text extraction asks:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;What text can I get from this PDF?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Text editing asks:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;How can I change that text while keeping the PDF looking and behaving correctly?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Those are two very different problems.&lt;/p&gt;

&lt;p&gt;You can have a PDF where text extraction works perfectly but modifying that text causes the font, spacing, position, or layout to change.&lt;/p&gt;

&lt;h2&gt;
  
  
  Characters Aren't Always as Simple as They Look
&lt;/h2&gt;

&lt;p&gt;Another problem is the relationship between characters and glyphs.&lt;/p&gt;

&lt;p&gt;When we see the letter &lt;code&gt;A&lt;/code&gt;, we naturally think the PDF contains the character &lt;code&gt;A&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;But internally, the process can be more complicated.&lt;/p&gt;

&lt;p&gt;A simplified version looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Character code
      ↓
Encoding / CMap
      ↓
Glyph
      ↓
Font
      ↓
Rendered shape
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The PDF specification describes text in terms of character codes that are interpreted using font information to select the glyphs that are actually drawn.&lt;/p&gt;

&lt;p&gt;So the thing you see on the screen isn't necessarily a direct representation of the character you think you're editing.&lt;/p&gt;

&lt;p&gt;This becomes particularly important when you replace existing text.&lt;/p&gt;

&lt;h2&gt;
  
  
  Then There Are Fonts
&lt;/h2&gt;

&lt;p&gt;Fonts are probably one of the biggest headaches in PDF editing.&lt;/p&gt;

&lt;p&gt;A PDF can contain embedded fonts, and those fonts can also be subsets.&lt;/p&gt;

&lt;p&gt;Imagine the original font contains thousands of glyphs:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;A B C D E F ...
0 1 2 3 4 ...
α β γ ...
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;But the PDF only needs a small portion of them.&lt;/p&gt;

&lt;p&gt;It may embed a subset containing only the glyphs that are actually used.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Original font
     ↓
Thousands of glyphs
     ↓
Only required glyphs embedded
     ↓
Smaller PDF
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now imagine you're editing the PDF and introduce a character whose glyph wasn't included in the original subset.&lt;/p&gt;

&lt;p&gt;The editor has a problem.&lt;/p&gt;

&lt;p&gt;It may need to find another suitable font, embed additional font information, or decide that the edit can't safely be performed.&lt;/p&gt;

&lt;p&gt;That's very different from replacing a string in a normal text file.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Font Substitution Can Change the Layout
&lt;/h2&gt;

&lt;p&gt;Let's say a PDF originally contains:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Invoice Total: $1,250.00&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The text was positioned using a particular font and its metrics.&lt;/p&gt;

&lt;p&gt;If an editor replaces that font with something that looks similar, the result might still look slightly different.&lt;/p&gt;

&lt;p&gt;Different fonts can have different:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;character widths&lt;/li&gt;
&lt;li&gt;spacing&lt;/li&gt;
&lt;li&gt;glyph shapes&lt;/li&gt;
&lt;li&gt;ascent and descent&lt;/li&gt;
&lt;li&gt;metrics&lt;/li&gt;
&lt;li&gt;kerning&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So a replacement font can cause text to move or become wider or narrower.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Original:

Invoice Total: $1,250.00


After font substitution:

Invoice Total: $1,250.00
                         ↑
                    different width
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That small difference can be enough to cause text to overlap another element or extend outside its original area.&lt;/p&gt;

&lt;p&gt;This is why finding a font that "looks close enough" isn't necessarily good enough for a serious PDF editor.&lt;/p&gt;

&lt;h2&gt;
  
  
  Positioning Makes Things Even Harder
&lt;/h2&gt;

&lt;p&gt;PDF text is also positioned very precisely.&lt;/p&gt;

&lt;p&gt;The PDF keeps track of things such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;font&lt;/li&gt;
&lt;li&gt;font size&lt;/li&gt;
&lt;li&gt;character spacing&lt;/li&gt;
&lt;li&gt;word spacing&lt;/li&gt;
&lt;li&gt;text position&lt;/li&gt;
&lt;li&gt;transformation matrices&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Consider:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Name: Muhammad Ali&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Now you want to change it to:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Name: Muhammad Ali Khan&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The replacement text is longer.&lt;/p&gt;

&lt;p&gt;What should the editor do?&lt;/p&gt;

&lt;p&gt;Keep the same font size?&lt;/p&gt;

&lt;p&gt;The text might extend into another part of the page.&lt;/p&gt;

&lt;p&gt;Make the font smaller?&lt;/p&gt;

&lt;p&gt;Now it doesn't match the original.&lt;/p&gt;

&lt;p&gt;Change the spacing?&lt;/p&gt;

&lt;p&gt;That can create another visual problem.&lt;/p&gt;

&lt;p&gt;Move other content?&lt;/p&gt;

&lt;p&gt;That becomes even more complicated.&lt;/p&gt;

&lt;p&gt;A PDF generally isn't designed to automatically reflow its content like a word processor.&lt;/p&gt;

&lt;p&gt;That's one of the reasons editing existing PDF text is so difficult.&lt;/p&gt;

&lt;h2&gt;
  
  
  The "It Looks Fine" Problem
&lt;/h2&gt;

&lt;p&gt;There's another problem that isn't immediately obvious.&lt;/p&gt;

&lt;p&gt;A modified PDF can look completely correct on your screen and still have problems internally.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Modified PDF
     │
     ├── Looks correct
     │
     └── But might contain:
          ├── incorrect font mapping
          ├── missing glyphs
          ├── broken text extraction
          └── changed document structure
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This matters because PDFs aren't only meant to be viewed.&lt;/p&gt;

&lt;p&gt;People also:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;search them&lt;/li&gt;
&lt;li&gt;copy text from them&lt;/li&gt;
&lt;li&gt;print them&lt;/li&gt;
&lt;li&gt;index them&lt;/li&gt;
&lt;li&gt;convert them&lt;/li&gt;
&lt;li&gt;process them with other software&lt;/li&gt;
&lt;li&gt;use accessibility tools with them&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So a good PDF editor shouldn't only ask:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Does it look right?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It should also ask:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Did we preserve the document correctly?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Why Covering Text With a White Box Isn't Real Editing
&lt;/h2&gt;

&lt;p&gt;One simple trick is to cover the original text with a white rectangle and then place new text on top.&lt;/p&gt;

&lt;p&gt;Visually, it might look like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Original text
     ↓
████████████
     ↓
New text
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For some visual workflows, that might be acceptable.&lt;/p&gt;

&lt;p&gt;But it isn't actually changing the original text.&lt;/p&gt;

&lt;p&gt;The original content may still exist underneath the rectangle.&lt;/p&gt;

&lt;p&gt;That can cause problems with:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;searching&lt;/li&gt;
&lt;li&gt;copying&lt;/li&gt;
&lt;li&gt;text extraction&lt;/li&gt;
&lt;li&gt;accessibility&lt;/li&gt;
&lt;li&gt;document structure&lt;/li&gt;
&lt;li&gt;sensitive information&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This becomes particularly important with redaction.&lt;/p&gt;

&lt;p&gt;Putting a black rectangle over sensitive information isn't necessarily the same as permanently removing that information from the PDF.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Real Engineering Challenge
&lt;/h2&gt;

&lt;p&gt;Once you put all of these problems together, the process starts looking very different from simple text replacement.&lt;/p&gt;

&lt;p&gt;A simplified workflow might look like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Find text
   ↓
Understand its PDF representation
   ↓
Identify fonts and glyph mappings
   ↓
Understand positioning
   ↓
Determine whether the edit is safe
   ↓
Modify the PDF
   ↓
Save it
   ↓
Validate it
   ↓
Render and inspect the result
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important part is that saving the PDF isn't necessarily the end.&lt;/p&gt;

&lt;p&gt;You need to make sure the resulting document still works.&lt;/p&gt;

&lt;p&gt;Does it open correctly?&lt;/p&gt;

&lt;p&gt;Does the text still extract correctly?&lt;/p&gt;

&lt;p&gt;Does it still render correctly?&lt;/p&gt;

&lt;p&gt;Are the required glyphs available?&lt;/p&gt;

&lt;p&gt;Did the edit change anything it shouldn't have?&lt;/p&gt;

&lt;p&gt;Those questions are just as important as the edit itself.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Building Around This Problem Taught Me
&lt;/h2&gt;

&lt;p&gt;When I first looked at PDF editing, the problem seemed straightforward:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Find text
   ↓
Replace text
   ↓
Save PDF
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The reality is much closer to:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Find text
   ↓
Understand its representation
   ↓
Understand fonts and glyphs
   ↓
Understand positioning
   ↓
Modify the right PDF objects
   ↓
Preserve required resources
   ↓
Save
   ↓
Validate
   ↓
Render and inspect
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's a huge difference.&lt;/p&gt;

&lt;p&gt;The difficult part isn't putting new text on a page.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The difficult part is changing existing content without breaking the relationships that made the original PDF render correctly.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  A Better Way to Think About PDF Editing
&lt;/h2&gt;

&lt;p&gt;If you're building PDF software, I think one of the most useful mental models is this:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;A PDF isn't primarily a collection of editable paragraphs. It's a structured description of a rendered page.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Once you start thinking about PDFs this way, many strange behaviors make more sense.&lt;/p&gt;

&lt;p&gt;Why did the font change?&lt;/p&gt;

&lt;p&gt;Maybe the original font couldn't safely represent the new text.&lt;/p&gt;

&lt;p&gt;Why did the text move?&lt;/p&gt;

&lt;p&gt;Maybe the replacement interacts differently with the original font metrics or positioning.&lt;/p&gt;

&lt;p&gt;Why can text extraction work while editing fails?&lt;/p&gt;

&lt;p&gt;Because extracting text and modifying the underlying PDF are completely different problems.&lt;/p&gt;

&lt;p&gt;Why isn't drawing a white rectangle over text the same as editing it?&lt;/p&gt;

&lt;p&gt;Because visual appearance and underlying document content are two different things.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;Editing existing PDF text is difficult because the text you see is only the final result of several layers working together.&lt;/p&gt;

&lt;p&gt;A simplified view is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Character codes
      ↓
Encoding / CMap
      ↓
Glyphs
      ↓
Fonts
      ↓
Text operations
      ↓
Positioning
      ↓
Content streams
      ↓
PDF objects
      ↓
Rendered page
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Once you understand that pipeline, it becomes easier to see why a PDF editor might work perfectly on one document and struggle with another.&lt;/p&gt;

&lt;p&gt;The hardest part isn't writing new text.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The hard part is changing existing content while preserving the visual and structural properties of the original document.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Disclosure:&lt;br&gt;
 I work on OnlinePDFEdits, a PDF software project. The PDF engineering problems discussed in this article come from my experience working on document-processing and PDF editing software. This article is intended as a technical discussion of PDF editing challenges and is not a product review or advertisement.&lt;/p&gt;

</description>
      <category>python</category>
      <category>productivity</category>
      <category>pdf</category>
      <category>ai</category>
    </item>
  </channel>
</rss>
