Text extracted from a PDF has no structure. Even if the author built the document in Word with Heading 1 and Heading 2, the PDF only keeps "draw these characters, at this size, at this position". Tools that let an LLM read long documents (RAG) cut the text into chunks and search them. Without headings, the only way to cut is by a fixed number of characters.
This post is about how the "Export text for AI" feature of PDF Privacy Checker (a Windows app that finds hidden text in PDFs) adds H1–H3 headings based on font size, weight and numbering. The previous post used the same character coordinates to tell soft wraps from real line breaks. This one continues from there.
It is not only a description of the rules. I measured whether headings actually help retrieval and whether wrong headings are worse than none. The measurement found a bug in the version I had already shipped.
TL;DR
- A PDF has no "this is a heading" marker (unless it is a tagged PDF). I compare each line's character height with the body text to find candidates, then require support from bold, numbering or extra space above. When unsure, the line stays body text
- Most of the rules are about what is not a heading. Each one removes something that real PDFs made me mistake for a heading: amounts, inline emphasis, subtitles, repeated print samples, badges, logos
- On 4 synthetic documents and 128 questions, cutting at headings improved retrieval hits (with a small budget: Japanese 84% → 97%, English 78% → 91%). The gain comes from questions phrased with the section title's words
- Even with half the headings wrong, retrieval in this setup (BM25, 4 synthetic documents) was never worse than no headings at all. My expectation before building it was wrong. The harm to summarization was not measured
- The measurement found a bug in v1.15.0: heading levels were decided page by page, so every heading after page 1 moved up one level. Fixed in v1.15.1
- On documents I did not look at while writing the rules (my own 10 articles printed from Chrome), 83 of 96 large headings were found but only 9 of 81 small ones. The small headings are 1.14× the body size, and pdfium reports the bold font's weight as 340. The rules lost on both counts (not fixed yet)
A PDF has no headings
PDFs can carry document structure (tagged PDF). But most PDFs you receive at work either have no tags or tags you cannot trust: PDFs printed from business software, scans with an OCR layer, exports from old tools.
So the structure has to be inferred from how the page looks. pdfium (the engine behind Chrome's PDF viewer) gives this per character:
| What | pdfium function | Used for |
|---|---|---|
| Character box (coordinates) |
FPDFText_GetCharBox etc. |
line height, width, left edge, spacing |
| Weight | FPDFText_GetFontWeight |
is it bold |
| Font name | FPDFText_GetFontInfo |
fallback when the weight is 0 (does the name contain Bold) |
| The character | FPDFText_GetUnicode |
numbering, trailing punctuation, amounts |
The previous post rebuilt the visual lines from the character boxes. Heading detection runs on each of those visual lines.
The output looks like this (page 1 of a synthetic handbook made for the measurement):
# Employee Handbook 2026
## 1. Travel
### 1.1 Domestic Trips
Trips of 100 km or more one way qualify. The daily allowance is $40 per day. …
### 1.2 International Trips
The rules: size proposes, support decides
A visual line becomes a candidate only if all of these hold:
- 2 to 60 characters, containing letters (kana, kanji or Latin)
- It does not start with a bullet mark (
•,・,※…) - It does not end with punctuation (
。、!?:…) - It was not joined as a soft wrap by the previous post's rules (it is not in the middle of a paragraph)
- It is shorter than 75% of the page's text width (a large cover title is the exception)
Then I look at the line's character height divided by the page's body character height (median):
| Ratio | Treatment |
|---|---|
| 1.6 or more | large heading candidate |
| 1.3 or more | medium heading candidate |
| 1.15 or more | small heading candidate, but only with bold, numbering or extra space above (0.8× the body height or more) |
| below 1.15 | candidate only if bold, 30 characters or fewer, and backed by numbering, space above, or the previous line ending a sentence |
Numbering does not decide whether a line is a heading. It decides how deep it is:
| Pattern | Depth |
|---|---|
第1章 / 第1節 / 第1条 (Japanese chapter / section / article) |
1 / 2 / 3 |
1. / 1 + space / full-width 1.
|
1 |
1.1 / 1.1.1
|
2 / 3 |
(1) |
3 |
A digit glued directly to a word (5-minute drill, 3 reasons, 2026 plan in Japanese, where there is no space) is not treated as a section number. In a real practice-menu PDF, "5分…" was read as "5." and the title lost its place.
These numbers do not come from the PDF specification. I started with guesses and fixed only what broke on real and synthetic documents. Where it broke is below.
Levels are decided relative to the document
The first version mapped the ratio straight to a level (1.6+ → H1, 1.3+ → H2). Business documents broke it. In a notice, the subject line is the largest line and section headings are barely larger than the body. In a report, the title is large and sections are bold at body size. "1.6× means H1" means different things in different documents.
Now levels are relative to the document:
- Title: among the first candidates, the largest one with a ratio of 1.3 or more. H1
-
Numbered: the numbering depth, one level lower if there is a title (
1.under an H1 title is H2) - Everything else: the size tiers (ratios rounded to 0.05), largest first, become H2, H3. H1 first if there is no title
- Bold-only candidates take the lowest tier
Simplified, the core looks like this:
# cand: candidate lines -> {"ratio": size vs body, "depth": numbering depth or None}
# doc : carried across the whole document (title ratio and size tiers)
if doc.get("title") is None and cand:
top = max(c["ratio"] for c in cand.values())
if top >= 1.3:
doc["title"] = next(c["ratio"] for vi, c in sorted(cand.items())
if c["ratio"] >= top - 0.05 and not c["depth"])
title = doc.get("title")
tiers = sorted(set(doc.get("tiers", [])) |
{round(c["ratio"] / 0.05) * 0.05 for c in cand.values()
if not c["depth"] and c["ratio"] >= 1.15
and (title is None or c["ratio"] < title - 0.05)},
reverse=True)
doc["tiers"] = tiers
base = 2 if title is not None else 1
for vi, c in cand.items():
if title is not None and not c["depth"] and c["ratio"] >= title - 0.05:
level = 1 # same size as the title
elif c["depth"]:
level = min(3, c["depth"] + (1 if title is not None else 0))
elif c["ratio"] >= 1.15:
level = min(3, base + tiers.index(round(c["ratio"] / 0.05) * 0.05))
else:
level = min(3, base + len(tiers)) # bold only = lowest tier
The important part is that doc is carried across pages. Not doing that was the v1.15.0 bug that shows up later.
Approaches I rejected
| Approach | Why not |
|---|---|
| Numbering alone makes a heading | numbered steps like "1 Count the cash in the register…" would become headings. Numbering only sets the depth |
| Bold alone makes a heading | inline emphasis is bold too. Require shortness plus numbering or space above |
Bold in the font name |
a name is only a declaration. Use pdfium's weight, and the name only when the weight is 0 |
| Space above alone | cannot be told apart from a paragraph break |
| Drop lines with many digits | removes amount lines, but also headings like "Steps (5 minutes per round)". Amounts are matched by pattern instead (currency sign + digits, thousands separators) |
What real PDFs made me mistake for headings
More than half of the rules exist to say "not a heading". Each came from a real PDF on my machine.
| Mistaken for a heading | What happened | Fix |
|---|---|---|
| "Quoted total ¥132,221" on a quote | a big amount line became H1 | lines containing an amount pattern are never headings |
| a large inline link in "open it and bookmark it ." | inline emphasis became a heading | if the next line continues with a particle or a lowercase letter from the same left edge, it is body text |
| a small app name right above the title | a subtitle became a heading above the title | if the next line is a larger candidate and the spacing is normal, it is body text |
| a company name ×3 on a business-card print proof | repeated sample text became headings | the same string twice on a page is body text |
| badges "OK" / 注意 (caution) / 警告 (warning) on the same PDF | short labels became headings | a size-only H3 needs 3+ characters; 2 characters only if full-width and bold |
| the logo "S O F T W A R E" on the same PDF | letter-spaced decoration became a heading | lines where 70%+ of the tokens are single characters are body text |
During the rewrite I also built a synthetic test of six business document types (a Japanese notice, report and minutes, and an English memo, report and letter). At first only 7 of 22 expected headings were right; after the fixes, all 22 were. Two things came out of it:
- A non-embedded Helvetica-Bold (one of the PDF standard fonts) comes back from pdfium with weight 0. I added the font name's
Boldas a fallback - A rule "not a heading if the previous line is mid-sentence" was dropping the subject line right after the sender block of a notice. I replaced it with "not a heading if the next line continues with a particle or a lowercase letter"
That 22/22 was reached while looking at the test. Further down, I measure on documents I did not look at.
Do headings actually help retrieval?
Now the measurements. The reason for adding headings was "the AI can see the structure", but I had never checked that it helps.
Setup
- Documents: 4 synthetic documents. A Japanese company rulebook (numbered), a Japanese store staff manual (no numbering, bold sizes for levels), an English Employee Handbook (numbered, Helvetica) and an English User Guide (no numbering). Formatting close to Word defaults, 2–3 pages each
- Questions: 2 per section, 128 in total. One uses the section title's words ("What is the daily allowance for international trips?"), one uses the answer paragraph's words ("When a destination has safety advisories, who approves the request and when?")
- Traps: 10 patterns that real PDFs made me mistake for headings (amount lines, large inline emphasis, bold notes) are mixed into the body. The correct answer is "not a heading"
- Search: BM25 (classic ranking by word overlap; Japanese is split into 2-character pieces). No LLM involved, so the numbers are the same on every run
- Collection: the two documents of the same language share one index
- Hit: take chunks from the top of the ranking until the character budget for the AI is full. A hit means a chunk containing the answer string made it in
- Compared: the pre-heading version's output (v1.14.1) cut into fixed-size chunks (the best-performing size), against the output cut at headings
The synthetic documents, ground truth, the Markdown the app exported at each stage, and the scoring scripts are on GitHub. The app itself is closed source, but the scoring runs with the Python standard library alone. Since v1.15.1 the export is free, so you can also regenerate the Markdown with the app.
git clone https://github.com/okinawasoftwarelab/pdf-heading-retrieval-bench
cd pdf-heading-retrieval-bench
python bench_run.py # retrieval hit rates
python bench_mutate.py # harm of wrong headings
python bench_heads.py # heading accuracy per stage
Checked with Python 3.12. No third-party packages are needed.
Results (small budget: 120 characters for Japanese, 250 for English). A hit means a chunk containing the answer string made it into the budget; it is not a score of the LLM's answer quality.
| All | Title-word questions | Body-word questions | ||
|---|---|---|---|---|
| Japanese | no headings (60-char chunks) | 84% | 75% | 94% |
| v1.15.0 | 97% | 100% | 94% | |
| v1.15.1 | 97% | 100% | 94% | |
| correct headings | 97% | 100% | 94% | |
| correct headings (cut only) | 100% | 100% | 100% | |
| English | no headings (125-char chunks) | 78% | 69% | 88% |
| v1.15.0 | 91% | 88% | 94% | |
| v1.15.1 | 89% | 84% | 94% | |
| correct headings | 89% | 84% | 94% | |
| correct headings (cut only) | 94% | 91% | 97% |
Except for "cut only", each chunk starts with its heading path (Employee Handbook 2026 > 1. Travel > 1.2 International Trips).
What it shows:
- The difference is in title-word questions. Rulebooks and manuals do not repeat the section title in the body. The paragraph answering "daily allowance for international trips" does not say "international trips"; only the heading line does. Cutting at headings puts that line at the top of the chunk
- What helped was where the text was cut, not extra information. The no-heading version already contains the heading text ("1.2 International Trips") as a plain line. The conditions differ only in where the chunks are cut and whether the path is added. "Cut only" did best, so the gain comes from putting the heading line at the top of its chunk
- A larger budget shrinks the gap. At 500 characters, Japanese goes 98% → 100%; at 1,000, English goes 92% → 98%. Headings matter when the AI can only be given a little
- Adding the heading path was slightly worse than just cutting at headings. The chapter name goes into every chunk of the chapter, so it does not help tell them apart and only eats budget
- In English, v1.15.0 with its errors (91%) scored a little higher than v1.15.1, which matches the correct headings (89%). As the next section shows, a difference this size moves up or down with how the headings are wrong
Are wrong headings worse than none?
Before building this, I believed wrong headings were worse than no headings. That belief is why the rules fall back to "body text when unsure". I measured this too.
Setup: mix errors into the correct headings at rate p. Three kinds: insert the first few words of a paragraph as a heading (extra; the shape of emphasis or a note becoming a heading), turn a non-title heading back into body text (missed), give a non-title heading a different level (wrong level). Average of 30 random seeds.
| Error | p=10% | 25% | 50% | no headings | |
|---|---|---|---|---|---|
| Japanese | extra | 95.5% | 93.8% | 91.0% | 84% |
| missed | 95.2% | 93.0% | 88.1% | ||
| wrong level | 96.8% | 97.6% | 97.1% | ||
| English | extra | 88.2% | 87.0% | 82.8% | 78% |
| missed | 86.8% | 84.4% | 79.9% | ||
| wrong level | 88.6% | 89.7% | 90.3% |
- With half the headings wrong, retrieval was still never below no headings. For retrieval in this setup, my expectation was wrong
- Missing hurts more than extra. A section that is not cut off merges with its neighbour, and its title words are lost
- Wrong levels barely matter for retrieval. The cut positions do not change
I still kept "body text when unsure". When an LLM reads the text to summarize it, headings are read as structure. If the bold note "Amounts over the limit cannot be approved afterwards" becomes a heading, it reads like the start of a new section. That harm was not measured here. Looking at retrieval alone, these numbers say adding a few too many is better than missing some.
The measurement found a bug in v1.15.0
I also counted right and wrong headings. The 4 documents have 86 correct headings.
| Version | Correct | Missed | Extra (traps) | Wrong level |
|---|---|---|---|---|
| first heading rules | 20 | 32 | 6 (6) | 34 |
| after the business-card fixes | 17 | 35 | 0 | 34 |
| after the business-document fixes = v1.15.0 | 61 | 0 | 3 (3) | 25 |
| v1.15.1 | 86 | 0 | 0 | 0 |
All 25 wrong levels had the same cause. Levels were decided per page, so from page 2 on "this page has no title", and everything moved up one level. Page 2 of the English Employee Handbook:
v1.15.0 v1.15.1
# 3. Leave ## 3. Leave
## 3.1 Paid Vacation ### 3.1 Paid Vacation
## 3.2 Bereavement Leave ### 3.2 Bereavement Leave
# 4. Information Security ## 4. Information Security
Page 1 was right, so tests built from one-page documents never caught it. v1.15.1 carries the title ratio and the size tiers across the whole document (doc in the code above). On later pages, a line the size of the title is still H1.
The 3 extras were traps:
- the bold note "Amounts over the limit cannot be approved afterwards". If a bold-only candidate is immediately followed by another heading, it is a label with no content and goes back to body text
- the two-character bold 重要 ("Important"), caught by the same rule
- a large inline emphasis in the middle of a Japanese sentence (roughly "until the investigation ends / do not discuss it with anyone outside the team / (end of sentence)."). The next line starts with the sentence ending, not a particle, so the existing rule missed it. New rule: if the previous line is mid-sentence and the next line is a short sentence ending (20 characters or fewer), it is body text
The 86/86 for v1.15.1 was reached while looking at these 4 documents. It is like solving a problem after seeing the answer, so it is not evidence that the rules are good. The next section measures on documents I did not look at while writing the rules.
Measured on documents I did not look at
As in the previous post, I turned my own 10 articles (5 Japanese, 5 English) into minimal HTML and printed them to A4 PDF from Chrome. Body text is 10.5pt, large headings 14pt, small headings 12pt, and the font is Yu Gothic. The Markdown sources give the correct headings automatically: 96 large and 81 small, 177 in total.
| Correct | Found | Rate | |
|---|---|---|---|
| large headings (14pt) | 96 | 83 | 86% |
| small headings (12pt) | 81 | 9 | 11% |
| extra headings | 12 | 9 lines inside code blocks, 2 table rows, 1 English sentence ending in a period |
The levels are not consistent either. Of the 83 large headings found, 74 became H1 and 9 became H2. All 9 small headings found became H1, the same level as the large ones.
Why the small headings were missed (the first two are headings from the Japanese articles):
'v1で分かったこと' ratio 1.143 weight 340 space above 1.11
'差分の取り方' ratio 1.143 weight 340 space above 1.16
'The render diff' ratio 1.143 weight 340 space above 1.11
They slipped through both nets by a hair:
- Size: 12pt ÷ 10.5pt = 1.143, 0.007 short of the 1.15 threshold
-
Weight: pdfium reports the
YuGothic-Boldthat Chrome embedded as weight 340. The bodyYuGothic-Regularis 190. The rule says "600 or more is bold", so it is not bold. The font-name fallback only runs when the weight is 0
This is the same kind of problem as Helvetica-Bold's "weight 0". The absolute weight value cannot be trusted; it depends on the font and on the tool that made the PDF. A relative test, "clearly heavier than the body text", would separate 340 from 190.
The page-by-page levels, dropping "1.6× means H1", and now the weight: the cause was the same every time. Absolute thresholds break on some document. Comparing within the document holds up better. The weight problem is not fixed yet.
Another approach: rank the sizes
In the comments of Giacomo's benchmark post on PDF-to-Markdown converters, I asked how he assigns levels in documents without numbering. His answer tackled the same problem in a slightly different way:
- the body size is the most common size in the document
- sizes at least 15% larger than the body are candidates, and the three largest become H1, H2, H3. A 14pt line is not "an H2" by itself
- a size only counts as a level if it appears on at least three lines across at least three pages, so large figure labels and cover-only author names do not become levels
- bold lines at body size get levels below the size levels, but only if bold is a minority in the document
- nothing that does not recur may sit above the title
Both approaches decide levels relative to the document. His also filters by recurrence and treats bold relatively ("is it a minority"). On my printed articles, 12pt is +14.3% over the body, so his 15% line would miss it too. But a relative bold rule could catch the weight-340 headings (I have not tried it). On the other hand, Japanese business documents are usually one or two pages, so "three pages" rarely applies.
Limits
- Few documents. 4 synthetic documents (one paragraph per section, short) and my own 10 articles. 128 questions
- BM25 only. Not tested with embedding search. The gain came from questions phrased with the section title's words, which is exactly where word overlap helps; embedding search may show a smaller gap. A hit means the answer string is contained in a chunk
- Summaries and answers were not measured. There are no numbers for the harm of an LLM reading wrong headings as structure
- Tagged PDF structure is not used. When tags exist and are correct, they should be more reliable
- Vertical text, multi-column layouts and headings inside tables are out of scope
- Bold detection depends on the font. As the weight-340 case shows, even a weight that is returned may not mean what you expect
Takeaways
- A PDF has no heading marker. Compare each line's character height with the body to find candidates, and require support from bold, numbering or spacing. Most rules are about what is not a heading
- Headings helped retrieval, especially for questions phrased with the section title. With a large budget the gap shrinks
- Wrong headings did little harm to retrieval in this setup; missing ones hurt more than extra ones. The harm to summaries is not measured
- "1.6× means H1", "decide per page", "weight 600 or more": every absolute threshold broke on a real document. Comparing within the document holds up better
- Report numbers from documents you tuned on separately from numbers on documents you did not look at
Because of these measurements, v1.15.1 of PDF Privacy Checker fixes the page-by-page levels and the three traps. The small-heading weight problem is next.
PDF Privacy Checker is on the Microsoft Store (detection and the text export for AI are both free). Files never leave your PC; it works fully offline. The source code is not public.
https://apps.microsoft.com/detail/9PLRJHFTPS53?hl=en-us&gl=US
About this article — The implementation and the writing were done together with Claude (Anthropic's AI). Most of the prose was drafted by Claude and checked by me. Choosing the topic, and deciding to fix and ship what the measurements found, are mine.
About the author
Okinawa Software Lab. I lead in-house digital transformation at a small company in Okinawa, Japan. I build the tools we need ourselves, and I publish PDF apps on the Microsoft Store that follow the same principle: everything happens on your own PC.
- Website: https://okinawasoftwarelab.com/en/
Top comments (0)