DEV Community

Daniel Pertu
Daniel Pertu

Posted on

Next's compiler dropped a space in 166 of our sentences, and the source was correct

A visual audit of CogniPrep came back with a bug report I did not believe: a sentence on a blog page read "Word analogiesgive a pair and ask for the word that completes a second pair".

The source of that page read, and still reads, like ordinary prose:

<p>
  <strong>Word analogies</strong> give a pair and ask for the word that completes a second
  pair, from five options. Say the first relationship as a sentence, such as &ldquo;the first
  is a part of the second&rdquo;.
</p>
Enter fullscreen mode Exit fullscreen mode

There is a space after </strong>. The editor shows it, git shows it, the typechecker is happy. The space is simply not in the rendered HTML.

The three conditions

JSX whitespace has one rule everybody knows: a run of text whose leading whitespace contains a newline is trimmed. That rule does not apply here, because the leading whitespace is a single space on the same line as </strong>.

I pushed variants of that paragraph through the SWC that Next 16.2.4 bundles, which is the compiler that actually builds the app:

const swc = require('next/dist/build/swc');

(async () => {
  await swc.loadBindings();
  const out = await swc.transform(`const a = ${body};`, {
    filename: 't.tsx',
    jsc: { parser: { syntax: 'typescript', tsx: true }, transform: { react: { runtime: 'automatic' } } },
  });
  console.log(out.code);
})();
Enter fullscreen mode Exit fullscreen mode

The results, where "dropped" means the emitted string literal begins with give rather than " give":

Text run Leading space
After </strong>, wraps a line, contains &ldquo; dropped
After </strong>, wraps a line, no entity kept
After </strong>, one line, contains &ldquo; kept
After </a>, wraps a line, contains &apos; dropped
After </strong>, wraps a line, contains &#8220; dropped
After </strong>, wraps a line, contains a literal " character kept

So three things have to be true at once. The text run has to start on the same line as a previous sibling element, it has to span at least one newline, and it has to contain an HTML entity somewhere in the run, named or numeric. Change any one of those and the space survives. Write the quote mark as a literal character instead of an entity and the space survives.

Here is the compiled output of the failing case, which is the whole bug in one line:

_jsxs("p", {
  children: [
    _jsx("strong", { children: "Word analogies" }),
    "give a pair and ask for the word that completes a second pair, from five options."
  ]
});
Enter fullscreen mode Exit fullscreen mode

One more data point: @swc/core 1.16.13, installed standalone, keeps the space in that exact case. The behaviour belongs to the version Next ships, not to SWC in general, so "it works on my machine with Babel" and "it works with the latest swc" are both true and both useless to anyone building with Next today.

That is also why nobody caught it for months. The bug is invisible in the source, invisible in review, and invisible to every test that renders a component and asserts on its text, because innerText assertions are usually written by copying what the component actually produced.

The part that made it worth fixing properly

Whether a paragraph breaks the rule depends on where its lines happen to wrap. CogniPrep deliberately has no code formatter, so line breaks are wherever a human put them, and a sentence that is safe today becomes broken the moment somebody edits the middle of it and the line gets longer. That makes this a class of bug rather than a list of four typos, and it needs a scan rather than a search.

A text search cannot express "starts on the same line as the previous sibling". The TypeScript compiler API can, in about twenty lines:

const ts = require('typescript');
const { readFileSync } = require('fs');
const { execSync } = require('child_process');

const files = execSync("git ls-files 'app/**/*.tsx' 'components/**/*.tsx'").toString().trim().split('\n');
let hits = 0;

for (const file of files) {
  const text = readFileSync(file, 'utf8');
  const sf = ts.createSourceFile(file, text, ts.ScriptTarget.Latest, true, ts.ScriptKind.TSX);

  const walk = (node) => {
    if (ts.isJsxText(node)) {
      const raw = node.getFullText();
      const siblings = node.parent.getChildren().find((c) => c.kind === ts.SyntaxKind.SyntaxList)?.getChildren() ?? [];
      const prev = siblings[siblings.indexOf(node) - 1];
      const sameLine = prev && !text.slice(prev.getEnd(), node.getStart()).includes('\n');

      if (/^[ \t]/.test(raw) && raw.includes('\n') && /&[a-zA-Z]+;|&#\d+;/.test(raw) && sameLine) {
        hits++;
        console.log(`${file}:${sf.getLineAndCharacterOfPosition(node.getStart()).line + 1}`);
      }
    }
    node.forEachChild(walk);
  };

  walk(sf);
}

console.log({ scanned: files.length, hits });
Enter fullscreen mode Exit fullscreen mode

Run against the commit before the fix, over 898 .tsx files: 166 text runs in 88 files. Run against the current tree: zero. Many of the 166 were on pages written long before the audit that reported the first one, which is the usual shape of this kind of thing. One bug report, one real class, two orders of magnitude more instances than the report.

The fix is the boring one. Each offending leading space became an explicit expression container, which no whitespace rule can touch:

<strong>Word analogies</strong>{' '}give a pair and ask for the word that completes a second
Enter fullscreen mode Exit fullscreen mode

It is ugly, and I would rather have ugly source than a missing space in front of thousands of readers.

See it

The three pages below are live, and the sentences in them are the ones that were broken. Open any of them and read the first few paragraphs, then view source and search for {' '}-shaped output, or just confirm the words are separated:

If you want to see the bug rather than the fix, drop the matrix above into any Next 16 project, or paste the first snippet in this post into a page, add a &rarr; to the sentence, and look at the server HTML.

What I would tell past me

Two things. First, "the source looks right" is not evidence; the compiled output is. Second, when a visual audit reports one cosmetic defect, check whether the defect has a mechanism. If it does, scan for the mechanism. The ratio here was 1 reported to 166 real.

Top comments (0)