DEV Community

Daniel Pertu
Daniel Pertu

Posted on

93 of the 199 questions in our FAQ markup are not on the page

FAQPage structured data has one hard requirement, and it is not about syntax: the question and answer text in the markup has to be present on the page. It is not a place to put extra keywords. It is a machine-readable copy of something a human can read.

Twenty four of our pages carry FAQPage markup. A one-line check says 93 of the 199 questions in it are not on the page at all.

The check

Structured data validators tell you your JSON parses and your required fields exist. None of them tell you whether the strings inside match the page, because that comparison needs the rendered document. In a browser console it is four lines:

const norm = s => s.replace(/\s+/g, " ").trim().toLowerCase();
const faq = [...document.querySelectorAll('script[type="application/ld+json"]')]
  .map(s => JSON.parse(s.textContent)).find(j => j["@type"] === "FAQPage");
const visible = norm(document.body.innerText);
faq.mainEntity.map(q => q.name).filter(q => !visible.includes(norm(q)));
Enter fullscreen mode Exit fullscreen mode

innerText rather than textContent matters here. textContent includes the contents of <script> tags, so the JSON-LD block finds itself and every question matches. That one substitution is the difference between "all clear" and the results below. I made exactly that mistake on the first run and got a clean sweep.

Across our provider pages:

Page Questions in markup Not visible on the page
aon 7 7
cubiks 4 4
predictive-index 4 4
criteria 11 9
testgroup 10 9
sova 9 8
wonderlic 11 8
acer 10 7
revelian 8 7
korn-ferry 7 6
hogan 11 6
shl 7 4
arctic-shores 8 1
thomas, ixly, kenexa, hirevue, assessio, mckinsey-solve, test-partnership 71 0

199 questions, 93 of them invisible, 17 of 24 pages affected, and 7 pages perfectly clean. That last column is the interesting one, because it means the fix already exists in the same codebase.

The cause is two arrays

On the affected pages, the FAQ exists twice. Once inside the JSON-LD block:

<script type="application/ld+json" dangerouslySetInnerHTML={{ __html: JSON.stringify({
  '@context': 'https://schema.org',
  '@type': 'FAQPage',
  mainEntity: [
    { '@type': 'Question', name: 'Do Aon tests use negative marking?', acceptedAnswer: { ... } },
    ...
  ],
})}} />
Enter fullscreen mode Exit fullscreen mode

And once in the JSX, hundreds of lines further down the same file:

{[
  { q: 'Should I guess if I run out of time on an Aon test?', a: '...' },
  ...
].map(({ q, a }) => (
  <div key={q} className="bg-card rounded-lg border p-6">
    <h3 className="mb-2 font-semibold">{q}</h3>
    <p className="text-muted-foreground text-sm leading-relaxed">{a}</p>
  </div>
))}
Enter fullscreen mode Exit fullscreen mode

Read those two examples again. They are about the same fact. Aon deducts marks for wrong answers, so you should not guess. A human reviewing either one would sign it off. But the markup claims the page asks "Do Aon tests use negative marking?" and the page never says that sentence, so the markup is describing a page that does not exist.

That is the failure mode of duplicated content in general, and it is worse here than usual for two reasons. Paraphrase is invisible to review, because both copies are correct and neither is the one you are reading when you check the other. And there is no feedback: nothing renders the JSON-LD, nothing tests it, and a mismatch produces no error anywhere. The drift can only be observed by a script that reads both.

The counts drifted too, which is the same bug showing itself less subtly: eight questions in the markup and six on the page, or eleven and twelve. Somebody edited one list.

The seven clean pages

The newer pages were built with the array as a module constant, mapped in two places:

const FAQS: { q: string; a: string }[] = [ /* ... */ ];

// once into structured data
mainEntity: FAQS.map(({ q, a }) => ({
  '@type': 'Question',
  name: q,
  acceptedAnswer: { '@type': 'Answer', text: a },
})),

// and once into the page
{FAQS.map(({ q, a }) => (
  <div key={q} className="bg-card rounded-lg border p-6">
    <h3 className="mb-2 font-semibold">{q}</h3>
    <p className="text-muted-foreground text-sm leading-relaxed">{a}</p>
  </div>
))}
Enter fullscreen mode Exit fullscreen mode

Nothing clever. One array, two consumers, and the "content must be present on the page" requirement is now structurally impossible to break. You cannot add a question to the markup without adding it to the page, because they are the same array.

Those seven pages score zero in the right-hand column, and they did not need an audit to get there.

Why it is worth fixing beyond the policy

The obvious framing is risk: structured data whose content is not on the page is a spam signal, and rich results can be withdrawn for it.

The better framing is that structured data is an API you are publishing about your own content, and its consumers now include systems that quote you directly rather than linking to you. A question and answer pair in your markup can be read aloud to somebody who never loads the page. If it says something the page does not say, you have shipped an answer that nobody on your team has ever reviewed in context.

The rule that comes out of this is the same one that applies to any crawler-only copy: a machine-readable copy is a copy, and a copy needs a single source. If your structured data contains a string that appears nowhere else in your codebase, that string is unreviewed content published in your name.

See it: paste the four-line snippet above into the console on cogniprep.app/games/thomas. It returns an empty array: every question in the markup is on the page. Run the same snippet on cogniprep.app/games/aon and it returns seven strings. Then scroll to the Frequently Asked Questions section on that page and compare: the topics are all there, the wording is not, and that is the entire bug.

The snippet works on any site that publishes FAQPage markup, including yours. It is four lines, and it answers a question no validator asks.

Top comments (0)