DEV Community

Steven Browning
Steven Browning

Posted on

My FAQ Schema and My FAQ Had Become Two Different Documents

I was rewriting the copy on one of my GPU comparison pages when I noticed something I probably should have checked much earlier.

The FAQ people could read and the FAQ in the page's structured data were completely different.

Not slightly different.

They didn't share a single question.

One page, two different FAQs

The FAQPage JSON-LD had four questions:

  • What matters most for running local LLMs?
  • Can I run a local LLM without a GPU?
  • How much VRAM do I need for local AI?
  • Is NVIDIA required for local LLMs?

The actual page had three:

  • Is VRAM more important than raw GPU speed?
  • Can two GPUs combine VRAM?
  • Should I buy a prebuilt AI PC?

So I had one FAQ for visitors and a different one sitting in the code for anything reading the structured data.

Both were on the same page. Both looked reasonable on their own. Neither matched the other.

I can see how I got there

On these pages, I had essentially been keeping the FAQ in two places.

One copy was the HTML people read. The other was JSON-LD, usually in the <head>.

They might start out matching, but then I rewrite an answer, change a question or add something to the page.

And forget to update the other copy.

Nothing obvious breaks.

The page still loads. The FAQ still opens. The JSON can still be perfectly valid.

But valid JSON doesn't mean it describes what's actually on the page.

That was the part I had missed.

I fixed the page and thought I was done

I rebuilt the GPU page's schema from the FAQ visitors could actually read.

Then I checked my email header analyzer and found a less obvious version of the same problem.

It had six questions. One question had a different title in the schema, and all six answers were worded differently from the answers on the page.

I fixed those too.

On another page, I added two FAQ entries and checked that those two matched in both places.

They did.

So, naturally, I thought I was making progress.

Then I wrote a scanner to check the whole site. Mostly because I wanted to confirm I hadn't left anything else behind.

That turned out to be optimistic.

The numbers were worse than I expected

The scan found:

What I checked Result
Pages with FAQPage markup 45
Pages where schema and visible text matched under the check 17
Pages with mismatches 28
Questions in the schema 334
Schema questions not found on the page 56
Questions found, but with answer text that didn't match 105

On the mining calculator, the schema had six questions and the page had five. None of the questions matched.

The About page had the same problem.

And the email security checker, which I had considered finished a few hours earlier, still had three original questions worded differently in the two places.

One version said:

Why does the check say it could not verify DKIM?

The other said:

Why does it say DKIM could not be verified?

Those mean pretty much the same thing. But they were another sign that I was maintaining two separate copies instead of one source.

I had checked the entries I'd just added.

I hadn't checked the whole FAQ.

Of course, my first scanner was wrong too

The first version looked for a particular HTML pattern:

<details>
  <summary>A question</summary>
  <p>An answer</p>
</details>
Enter fullscreen mode Exit fullscreen mode

It then told me that a lot of my pages had no visible FAQ at all.

Except I could open those pages and see the questions right there.

Some used headings and paragraphs. Some answers had more than one paragraph.

My scanner wasn't finding them because I had told it what the markup should look like instead of accounting for what was actually there.

So now I had a checker that needed checking.

The next version removed scripts and styles, extracted the page text, and compared that with the questions and answers in the schema.

The basic idea was:

const norm = t => t.toLowerCase().replace(/[^a-z0-9]+/g, ' ').trim();

const visible = norm(visibleText);

for (const q of faq.mainEntity) {
  if (!visible.includes(norm(q.name))) {
    fail(`question not on the page: ${q.name}`);
  } else if (!visible.includes(norm(q.acceptedAnswer.text))) {
    fail(`answer differs from the page: ${q.name}`);
  }
}
Enter fullscreen mode Exit fullscreen mode

That caught the differences I was looking for without depending on one FAQ layout.

It still has limits. Finding a question and an answer somewhere in the page text doesn't prove they're paired correctly. And this simplified normalization is meant for my English-language pages, not every language or every kind of content.

But it gave me a much more useful check than assuming every page used the same HTML.

Was this hurting my SEO?

I don't want to turn this into a bigger SEO story than the evidence supports.

Google stopped showing FAQ rich results on May 7, 2026. Its general structured-data guidelines still require marked-up content to be visible to readers.

Those are separate points.

The mismatches were real, but I don't have evidence that they cost me rankings.

What bothered me was simpler: my page was saying two different things depending on which copy you read.

This time it was FAQ wording. With a calculator, it could be a number shown in the explanation that no longer matches the number used in the calculation.

That's the sort of thing I want to catch before someone else has to point it out.

Fixing the copies isn't enough

I could go through all 28 pages and edit the JSON by hand.

That would make them match today.

It wouldn't stop me from changing an answer next week and creating the same problem again.

So the approach I'm moving toward is:

  • Use the visible FAQ as the source. Generate the schema from it at build time instead of editing both copies.
  • Check the whole site on each build. Include pages I didn't touch that day.
  • Don't add FAQPage markup without a visible FAQ. There shouldn't be a separate set of answers available only in the structured data.

The StashGrid local LLM GPU comparison was where this started. By the time I ran this scan, its seven questions matched in both places. The scan had also given me a list of 28 other pages to work through.

The lesson for me was pretty basic.

Checking that my latest edit is correct isn't the same as checking that the page is correct.

And if I'm keeping the same information in two places, I need something that makes sure it stays the same.

Top comments (1)

Collapse
 
beusebiu profile image
Eusebiu Balan •

I went one step earlier than parsing the visible FAQ. Each page keeps its questions in one array, and the template loops over it twice, once for the visible section and once for the JSON-LD in the head.

Change an answer and both copies change with it, and there's no HTML to scrape back out.