Our company ships four apps. A live quiz night for pubs, a practice tool for psychometric hiring tests, a barcode scanner for people with gut conditions, and a watcher for rental listings. If you were handed that list cold you would assume four unrelated indie projects that happened to be built by the same person.
A crawler assumes the same thing, and the assumption is worse than neutral. Four small domains that link to each other, with nothing declaring why, is the shape of a link scheme. The relationship either gets stated in a machine readable way or it gets guessed at, and you do not want it guessed at.
The way you state it is schema.org, and the implementation is the part nobody writes about: the same file, copied by hand into four repositories, with one line different in each.
The node every site emits, identically
export function publisherOrganizationLd() {
return {
'@type': 'Organization',
'@id': PUBLISHER.id,
name: PUBLISHER.name,
subOrganization: SONACODE_APPS.map((app) => ({
'@type': 'Organization',
'@id': organizationIdFor(app),
name: app.name,
url: app.url,
})),
}
}
Every one of the four sites emits that, with the same @id and the same four children in the same order. A crawler that has seen any one of the sites has seen the whole family. Agreement is the entire mechanism: four sites making the same claim about who publishes them is evidence, and four sites each making a slightly different claim is noise.
Then each site emits a fuller node about itself, naming that publisher as its parent:
export function siteOrganizationLd() {
return {
'@context': 'https://schema.org',
'@type': 'Organization',
'@id': organizationIdFor(SELF_APP),
name: SELF_APP.name,
url: SELF_APP.url,
logo: `${SELF_APP.url}/logo.svg`,
description: SELF_APP.blurb,
parentOrganization: publisherOrganizationLd(),
}
}
The @id here and the @id of this app inside the parent's subOrganization list are the same string, so they are the same node, and a consumer merges them. That is what makes the sibling listing on one site and the self description on another resolve to one entity rather than two.
The publisher is an identifier, not a link
export const PUBLISHER = {
name: 'Sonacode Ltd',
id: 'https://cogniprep.app/about#sonacode',
} as const
The company does not have a website in this graph, and that is deliberate. An @id in JSON-LD is required to be an IRI, but it is not required to resolve to anything. Pointing it at a fragment on a page that already exists gives all four sites a stable string to agree on, without creating a fifth site that would then have to be maintained in step with them.
If the company ever gets a site, that URL replaces the identifier in all four repositories in the same sitting. Until then, the thing I most wanted to avoid was a page whose only reason to exist is to satisfy a schema.
One line differs, and the type system guards it
export const SELF: AppKey = 'pubtrivia'
That is the only line that is not identical across the four copies. It is also the single most likely thing to be wrong, because the way a new repository gets this file is that somebody pastes it and edits one word.
AppKey is derived from the array:
export const SONACODE_APPS = [ /* ... */ ] as const satisfies readonly SonacodeApp[]
type AppKey = (typeof SONACODE_APPS)[number]['key']
as const satisfies rather than a : readonly SonacodeApp[] annotation, and the difference is not stylistic. The annotation widens every key to string, AppKey becomes string, and a typo in SELF compiles perfectly happily. What you get then is a site that fails to find itself, falls through to listing itself as one of its own siblings, and emits a graph claiming it is its own sibling. With satisfies, the literal types survive, so a misspelling is a compile error in the one place a misspelling was ever going to happen.
This is the second time that exact trick has paid for itself in this codebase. The first was the content registry, where the same widening would have let a page declare a cluster that does not exist.
Derive the fifth copy of the domain, do not store it
The per app @id used to be a field on each entry. It was always exactly ${url}/#organization, which means it was a second copy of the domain sitting next to the first one, free to drift away from it.
export function organizationIdFor(app: { url: string }): string {
return `${app.url}/#organization`
}
Drift there is unusually nasty because it is invisible. Nothing breaks, no page 404s, no test fails. You get two Organization nodes for one app in whatever Search Console decides to show you three months later, and reconstructing why involves diffing JSON-LD across four deployments.
The same instinct produced this:
const APP_COUNT = COUNT_IN_WORDS[SONACODE_APPS.length] ?? String(SONACODE_APPS.length)
export const PUBLISHER_INTRO = `Sonacode builds ${APP_COUNT} apps, and they look unrelated because they are: ...`
The prose on four about pages says "four apps" because the array has four entries. Spelled out in words, because a numeral that small reads as a typo in a sentence. Hardcoding the word is how you end up with a page confidently saying "three apps" above a list of four.
Identifiers stay production, so the graph ships only in production
Every URL in that file is the production one, in every environment. They are identifiers shared across four deployments, not links to whatever host the current build happens to answer on.
Which creates an obvious hazard: a preview deploy emitting the production @id is a second thing claiming to be the real site. So the layout only emits the graph when the build knows it is the real site:
export const IS_PRODUCTION_SITE = SITE_URL === PRODUCTION_URL
The alternative, rewriting the identifiers per environment, is strictly worse. It makes the preview's graph internally consistent and externally meaningless, and it means the thing you test is not the thing you ship.
The copy is the design, and that needs saying out loud
Duplicating a file across four repositories is the sort of thing that gets flagged in review, and the reflex fix is a shared package. We did not, for a boring reason: four repositories deploy independently, and a version skew in a shared package would mean two sites emitting different versions of a node whose entire value is that it is identical everywhere. The coupling is real either way. The duplicate makes it visible.
What the file carries instead is an instruction, in a comment at the top, to diff it against the other three before changing it and to change all four in one sitting. That is a human protocol rather than a technical guarantee, which is honest about what it is.
One more thing worth stealing is the tone of the shared paragraph. It opens by admitting the apps have nothing to do with each other, that a quiz night for a pub and a food scanner for IBS share no audience at all. A page that implied otherwise would be selling, and a reader who arrived from one of the other three apps can tell. What they actually share is a standard, and that is the only claim worth making.
Check it yourself
The graph is in the page source. Pull the publisher node out of two of the four sites and diff them:
curl -s https://pub-trivia.app/about | grep -o '"@id":"[^"]*"' | sort -u
curl -s https://munchable.app/about | grep -o '"@id":"[^"]*"' | sort -u
Two unrelated products, one list of identifiers, in the same order. The cogniprep.app/about#sonacode entry is the publisher both of them point at, and each site's own #organization id appears in both lists.
Then read pub-trivia.app/about down to the "Also from Sonacode Ltd" section. The paragraph, the app names, the one line summaries and the longer descriptions under them are all rendered from the same array the JSON-LD above was built from, so the thing a person reads and the thing a crawler parses cannot disagree.
Top comments (0)