We re-read a few hundred fee schedules every week, many of them published as PDFs, and check that a quoted amount is still on the page. One bank's price list kept failing the check. The PDF clearly said "CHF 15,000". pdf.js gave us "CHF 1 5 ,000".
Why it happens
page.getTextContent() returns text items, not lines. A PDF generator is free to draw one visible word as several pieces: a glyph run per kerning pair, per font change, or simply per character. Each item carries its string, a transform matrix and a width. If you join items with a space, a number drawn in three pieces comes back with two spaces in it. If you join with nothing, separate words run together.
The fix: join items that touch
The transform matrix gives everything needed. transform[4] is the x position, transform[5] the y position of the baseline, and transform[3] the vertical scale, which for unrotated text is the font size. item.width is the advance width in the same units. Two items belong to the same word when they share a baseline, share a font size, and the second starts where the first ends:
export function joinPdfItems(items) {
let out = "";
let prev = null;
for (const it of items) {
if (!("str" in it)) continue; // marked-content items carry no text
if (prev) {
const size = Math.abs(it.transform[3]) || 1;
const touching = !prev.hasEOL && it.str !== "" && prev.str !== ""
&& Math.abs(it.transform[3] - prev.transform[3]) < 0.01
&& Math.abs(it.transform[5] - prev.transform[5]) < 0.5
&& Math.abs(it.transform[4] - (prev.transform[4] + prev.width)) < 0.1 * size;
out += prev.hasEOL ? "\n" : touching ? "" : " ";
}
out += it.str;
prev = it;
}
return out;
}
The tolerance is relative to the font size: a tenth of an em. A real word space is usually a quarter to a third of an em, so it stays a space, while the sub-pixel gaps between pieces of one word disappear. hasEOL marks the end of a line, so lines stay lines.
Two more traps from the same job
- A URL ending in .pdf does not always return a PDF. Some sites answer with an HTML page that wraps the document in an iframe. Check the first bytes for "%PDF" and, if you got HTML, read the iframe's src once and fetch that.
- Fonts with CID encodings only map back to readable text through their ToUnicode tables. A small hand-rolled PDF text reader often fails on them; pdf.js handles them, so it is worth using even when a lighter parser works for most files.
Testing it
Keep a few real PDFs that failed before as fixtures and assert on the exact strings you expect ("CHF 15,000", not "CHF 1 5 ,000"). The tolerance values are easy to get subtly wrong, and a fixture catches it the day a provider changes its PDF generator.
Top comments (0)