Google's structured data documentation has a lot of fields in it. Underneath all of them is one rule, and it is the only one with teeth: the markup must describe content that is visibly on the page. Not content that was there last month, not content the page is about in spirit. Content a person looking at the page can see.
Breaking it is not a style error. A rich result can be suppressed, or a site penalised, for markup that promises something the page does not show.
The usual way to comply is diligence: write the page, write the JSON-LD, remember to update both. That works until the day it does not, and the day it does not is invisible, because nothing on the page looks wrong. So across our 69 content pages we do not comply by remembering. Every builder takes the same data the page renders, and the two cannot disagree because they are the same variable.
The FAQ, rendered and marked up from one array
export function FaqBlock({ entries, heading, emitStructuredData = true }) {
if (entries.length === 0) return null
return (
<section>
{emitStructuredData && <JsonLd data={faqLd(entries)} />}
{heading && <h2>{heading}</h2>}
<dl>
{entries.map((entry) => (
<div key={entry.question}>
<dt>{entry.question}</dt>
<dd>{entry.answer}</dd>
</div>
))}
</dl>
</section>
)
}
entries feeds the <dl> and it feeds faqLd. There is no path by which a question exists in one and not the other. Deleting a question deletes its Question node, and no reviewer has to notice.
You can count it:
curl -s https://pub-trivia.app/faq | grep -o '"@type":"Question"' | wc -l
Sixteen, and sixteen questions render on the page. They are sixteen for the same reason your left and right hands have the same number of fingers, rather than because somebody checked.
Why there is a flag on it
emitStructuredData is the one piece of ceremony in there, and it is not a convenience.
FAQPage is a page-level node. It makes a claim about the whole document: this page is a set of questions and answers, and here they are. A page carrying two FAQ blocks must publish one node covering both. Two nodes, each claiming to describe the page, is a document that describes itself twice and disagrees with itself both times.
So a page with two blocks combines the arrays and lets a single block emit. The default is true because one block is the normal case.
The HowTo steps are the rendered section
Our guide on hosting a quiz night carries HowTo markup. The steps are not typed into the markup builder:
const RUNNING_ORDER = SECTIONS.find((section) => section.heading === "The running order")!
const HOW_TO = howToLd({
node: requireNode("/guides/how-to-host-a-pub-quiz"),
totalTime: "PT2H",
steps: (RUNNING_ORDER.points ?? []).map((point) => ({
name: point.term,
text: point.detail,
})),
})
SECTIONS is the array the page body renders from. The HowTo reaches into it and finds the section that is a procedure. Rewrite a step in the prose and the structured data changes with it, because there was only ever one copy of the step.
curl -s https://pub-trivia.app/guides/how-to-host-a-pub-quiz \
| grep -o '"@type":"[A-Za-z]*"' | sort | uniq -c
Eight HowToStep nodes, and eight steps in the running order on the page.
The markup we refused to emit
Deriving the markup from the content is half the discipline. The other half is declining to claim things, and that part cannot be automated, so it is written into the API as a decision somebody has to make.
HowTo is only on the guides that genuinely are numbered procedures:
Only used where the page really is a numbered procedure a reader follows. "Pub quiz round ideas" is a list, not a HowTo, and marking it up as one would be describing the page as something it is not.
Article is a prop on the page shell rather than a default:
/**
* Emit `Article` structured data. True for the guides, which are editorial
* writing, and false for the product pages, which are not: a pricing table
* marked up as an article is a claim about the page that is not true.
*/
article?: boolean
And the author, which is the one I think about most:
/**
* An editorial page, for the guides.
*
* The author is the Organization rather than a person. A named human author is
* the stronger signal and we do not have one to name. Inventing "Sarah, Head of
* Quizzes" to satisfy a schema field is exactly the kind of thing structured
* data is meant to stop. If a real person starts putting their name to these,
* this is where it goes.
*/
A Person author ranks better than an Organization author. We know that. The field is a factual claim about who wrote the page, and there is no Sarah. (The Organization node itself, and its @id, are a story of their own.)
One date, used twice, on purpose
export function articleLd(node: ContentNode) {
const iso = `${node.updated}T00:00:00Z`
return {
'@context': 'https://schema.org',
'@type': 'Article',
headline: node.heading,
description: node.description,
author: { '@id': organizationIdFor(SELF_APP) },
publisher: { '@id': organizationIdFor(SELF_APP) },
datePublished: iso,
dateModified: iso,
inLanguage: 'en-GB',
}
}
datePublished and dateModified are both the node's updated date, which looks lazy and is a decision.
These are evergreen pages revised in place, not dated posts. "How to host a pub quiz" is not a thing that happened on a Tuesday. The date it was first drafted is not a fact about it that any reader or crawler needs; the date it was last checked is, and that is the date we have. Keeping a datePublished around so the pair looks more complete would mean maintaining a field whose only function is to be in the output.
headline and description come off the registry node, which is also where the <h1> and the <title> come from. One string, three places, no drift.
The breadcrumb is the breadcrumb
export function breadcrumbLd(path: string) {
const trail = breadcrumbFor(path)
if (trail.length < 2) return null
// ... map the trail into itemListElement
}
breadcrumbFor is the function the visible breadcrumb component calls. The rendered trail and the BreadcrumbList are the same trail, and the early return is there so the homepage, whose trail is one item, publishes no node rather than a breadcrumb of length one.
The shape of the rule
If you write JSON-LD by hand, the question to ask of every field is not "is this valid" but "what does the page have to show for this to be true, and what stops the two drifting apart next quarter". If the answer involves anybody remembering anything, it will drift.
Three places to see this working:
-
pub-trivia.app/faq: 16 questions, 16
Questionnodes, oneFAQPage. -
pub-trivia.app/guides/how-to-host-a-pub-quiz:
Article,HowTowith 8 steps,BreadcrumbListmatching the trail at the top of the page. -
pub-trivia.app/guides/pub-quiz-round-ideas: a list of ideas, with no
HowToanywhere in it, which is the point.
Top comments (0)