DEV Community

Daniel Pertu
Daniel Pertu

Posted on

Six hub pages, four ItemLists, two SoftwareApplications and one hub that marks up nothing

The rule we gave ourselves for structured data is that a page may only mark up what is on the page. No FAQPage without visible questions, no Article on something that is not an article, no invented author. It is a good rule and it is easy to agree with in the abstract.

Here is what it actually produced. Every number below came out of the live HTML on pub-trivia.app, not out of the repository, because the deploy is the only version of this that search engines ever see. You can read the same thing with one command:

curl -s https://pub-trivia.app/tools | grep -o '"@type":"[A-Za-z]*"' | sort | uniq -c
Enter fullscreen mode Exit fullscreen mode

The inventory

The site has six content hubs. This is what each one emits:

Hub Nodes beyond the breadcrumb Entries
/features SoftwareApplication 2 offers
/solutions nothing -
/guides ItemList 14
/quiz-questions ItemList 14
/compare ItemList 7
/tools ItemList 4

Every one of them also carries a BreadcrumbList with two items, and an Organization node that comes from the root layout rather than the page.

The four ItemLists are the easy case and they are correct: each entry count is exactly the number of links the page visibly renders. The tools hub lists four tools and marks up four, in the same order, with the same names. That is not a coincidence, it is the builder:

/** A hub's list of its spokes, in the order the page shows them. */
export function itemListLd(nodes: readonly ContentNode[]) {
    return {
        '@context': 'https://schema.org',
        '@type': 'ItemList',
        itemListElement: nodes.map((node, index) => ({
            '@type': 'ListItem',
            position: index + 1,
            name: node.title,
            url: absoluteUrl(node.path),
        })),
    }
}
Enter fullscreen mode Exit fullscreen mode

Twelve lines, and it takes the registry nodes rather than a hand-written list of names and URLs. The page renders its grid from the same array it passes in here, which is the only reason the two cannot drift. Writing the ItemList out by hand would have produced markup that was right on the day it was written.

The hub that marks up nothing

The solutions hub describes eleven kinds of venue, pubs, bars, restaurants, taprooms, sports bars, hotels, student unions, social clubs, charity fundraisers, freelance quizmasters and corporate events. Each one has a paragraph on the hub and a page of its own behind it, and each one is a link.

It emits a breadcrumb and nothing else.

There is no rule behind that, and no comment explaining it. The hub satisfies the condition the other four satisfy, which is a visible, ordered list of pages, and it is the longest such list on the site. It is just a JsonLd line nobody added. The honest reading of our own "only mark up what is on the page" rule is that it protects you from claiming things that are not there, and does nothing at all about the opposite mistake. Four of the six hubs are consistent with each other and the fifth is consistent with the sixth by accident.

The hub that describes the product instead

The features hub carries a SoftwareApplication rather than an ItemList, which is defensible: the page is a description of what the software does, with the nine feature pages as a secondary navigation rather than as the subject. A crawler reading it should come away with "this is a web application in the business category with two subscription offers", and it does.

What is less defensible is that the pricing page carries one too:

{
  "@type": "SoftwareApplication",
  "name": "PubTrivia",
  "url": "https://pub-trivia.app",
  "applicationCategory": "BusinessApplication",
  "operatingSystem": "Web browser",
  "description": "Live pub quiz software for pubs, bars and venues. ...",
  "publisher": { "@id": "https://pub-trivia.app/#organization" },
  "offers": [
    { "@type": "Offer", "name": "Pro", "price": "30.00", "priceCurrency": "GBP", ... },
    { "@type": "Offer", "name": "Ultimate", "price": "100.00", "priceCurrency": "GBP", ... }
  ]
}
Enter fullscreen mode Exit fullscreen mode

Both pages call the same builder with the same two offers and a different description. Neither node has an @id. So the published graph contains two anonymous SoftwareApplication nodes with identical names, identical URLs and identical offers, differing only in how they describe themselves, and nothing in the markup says they are the same application.

That is not a disaster. A consumer will probably reconcile them on the matching url. But the whole reason the Organization node on every page carries "@id": "https://pub-trivia.app/#organization" is so that a crawler merging data from several pages knows it is one company rather than seventy-two, and the node we were most careful about is sitting on the same pages as the node we were not. Giving the application an @id and having both pages reference it is about four lines of work and the right fix.

The one piece of indirection that is working

publisher is a reference, not an object:

publisher: { '@id': organizationIdFor(SELF_APP) },
Enter fullscreen mode Exit fullscreen mode

The thing it points at is emitted by the root layout, which means it is on every page of the site, so the reference resolves inside the same document on both pages that use it. The Organization node itself is serialised once at module load and only on the production host, because a preview deployment claiming the production @id would put a second node for the same company into whatever index picked it up.

Why the offers are in pounds

The visible price on the pricing page is converted to the visitor's own currency from the country the CDN reports. The Offer is in GBP regardless.

That is deliberate. A crawler has to be told one number, and it should be the same number on every crawl, otherwise the price in a search result depends on which data centre fetched the page that morning. The plans are defined in GBP, so that is the number that goes in the markup, and the conversion stays on the visible side where the person reading it has a currency.

What this audit cost

Nothing but a loop over the sitemap and a grep -o '"@type"'. If you have a content site with generated structured data, it is worth doing once: the failure we found is not markup that lies, which is the failure everybody writes rules against, but markup that is simply absent on one page out of six, which no rule catches and no validator complains about, because a page with no ItemList is not invalid. It is just quieter than its neighbours.

While you are in there, the free tools are the four pages in that first ItemList, and all four run in the browser with no account, so you can check the markup against the page and then go and print a quiz scoresheet with it.

Top comments (0)