I run ModernWebSEO, a small web design and SEO studio in Istanbul. Our own site is in four languages: Turkish (the main market, at the root), English (/en), Spanish (/es) and Arabic (/ar, right to left). The sitemap holds 700 URLs.
An SEO studio's site gets judged on the details, by people and increasingly by machines: Googlebot, Bingbot, and the crawlers and fetchers behind ChatGPT, Claude, Perplexity and Gemini. This post walks through how the site handles each layer, with code from the repo. The stack is Next.js 16 (App Router), TypeScript, Tailwind CSS v4, deployed on Vercel.
I'll also point out the places where the code is a compromise or simply wrong. Those are usually the useful parts.
1. Folder-per-locale routing, no i18n library
There is no next-intl and no [locale] dynamic segment. Each locale is a plain folder:
app/
page.tsx -> / (Turkish)
hizmetler/ -> /hizmetler
metodoloji/ -> /metodoloji
en/
layout.tsx
services/ -> /en/services
methodology/ -> /en/methodology
es/
servicios/ -> /es/servicios
metodologia/ -> /es/metodologia
ar/
layout.tsx
services/ -> /ar/services
Why folders? The URLs are translated, not just prefixed. /hizmetler/e-ticaret is /en/services/e-commerce in English and /es/servicios/tienda-online in Spanish. Writing each locale's route tree by hand made that explicit, and each language's pages can have their own layout and copy instead of one template with swapped strings.
The cost shows up in <html lang>. The root layout renders <html lang="tr" dir="ltr"> for every page. Reading the locale from headers() in the root layout would make every route dynamic, and nothing could be served as static from the CDN. So the locale layouts set language and direction on a wrapper instead:
// app/ar/layout.tsx (trimmed)
export default function ArLayout({ children }: { children: React.ReactNode }) {
// Root layout no longer calls headers() (static prerender).
// <html> carries lang="tr" dir="ltr"; everything locale-specific lives here.
return (
<div lang="ar" dir="rtl" className={`${cairoAr.variable} font-[family-name:var(--font-cairo)]`}>
{children}
</div>
)
}
The Arabic text renders right to left with its own font, and the page is still static. But the document-level lang is wrong on non-Turkish pages, and hreflang has to carry the real language signal. If I started over, I'd use a [lang] segment with generateStaticParams so the root <html> could get the right attribute at build time. It's a trade-off I made knowingly, and I'd rather say so than pretend it's ideal.
2. hreflang when the slugs are translated
Translated slugs mean hreflang can't be generated by swapping a prefix. Every page needs a map. For services, the map sits next to the route:
// app/en/services/[slug]/page.tsx (excerpt)
const HREFLANG_BY_SLUG: Record<ServiceSlug, Record<string, string>> = {
'e-commerce': {
tr: '/hizmetler/e-ticaret',
en: '/en/services/e-commerce',
es: '/es/servicios/tienda-online',
ar: '/ar/services/e-commerce',
},
// seo, aeo, mobile-responsive ...
}
Blog posts use a separate slug-map module (blogSlugMapTrToEn, blogSlugMapEnToTr, and so on), and generateMetadata builds the alternates from it. x-default points at the Turkish version, because that's the primary market:
// app/en/blog/[slug]/page.tsx (excerpt)
alternates: {
canonical: `/en/blog/${slug}`,
languages: {
tr: `/blog/${blogSlugMapEnToTr[slug] || slug}`,
en: `/en/blog/${slug}`,
es: `/es/blog/${blogSlugMapEnToEs[slug] || slug}`,
ar: `/ar/blog/${blogSlugMapEnToAr[slug] || slug}`,
'x-default': `/blog/${blogSlugMapEnToTr[slug] || slug}`,
},
},
Each locale gets its own self-referencing canonical. The English page never canonicalizes to the Turkish one; that would tell Google to drop the translation.
The same alternates also go into app/sitemap.ts, using the alternates.languages field that Next's MetadataRoute.Sitemap supports:
// app/sitemap.ts (excerpt)
const staticRouteMap = [
{ tr: 'metodoloji', en: 'en/methodology', es: 'es/metodologia', ar: 'ar/methodology' },
{ tr: 'vaka-analizleri', en: 'en/case-studies', es: 'es/casos-de-estudio', ar: 'ar/case-studies' },
// ...
]
const alternates = {
languages: {
tr: buildUrl(route.tr),
en: buildUrl(route.en),
es: buildUrl(route.es),
ar: buildUrl(route.ar),
'x-default': buildUrl(route.tr),
},
}
Gotcha: two sources of hreflang can disagree. While writing this post I compared page metadata against the sitemap. The Spanish methodology page lists all four languages. The English methodology page lists only tr and en, while the sitemap lists all four for it. hreflang has to be reciprocal, so this is the kind of mismatch Search Console reports as missing return links. If you declare alternates in two places, write a small check that diffs them, or generate both from one map.
A related cleanup: posts that were once published under the wrong language prefix (an English post at /blog/...) got permanent redirects in next.config.ts to their correct locale path, so the old URLs don't sit in the index as duplicates.
3. robots.txt as a route handler, with Content Signals
robots.txt is a route handler at app/robots.txt/route.ts, not a static file, so the rules are built from arrays:
// app/robots.txt/route.ts (excerpt)
const CONTENT_SIGNAL = 'search=yes, ai-input=yes, ai-train=yes'
const commonDisallow = ['/api/', '/admin/', '/private/']
function buildRules(userAgents: string[], disallow: string[]): string {
return userAgents
.map(
(ua) =>
`User-agent: ${ua}\nContent-Signal: ${CONTENT_SIGNAL}\nAllow: /\n${disallow
.map((d) => `Disallow: ${d}`)
.join('\n')}`
)
.join('\n\n')
}
The named list covers the search crawlers and the AI ones: GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, PerplexityBot, Perplexity-User, Google-Extended, Applebot-Extended, CCBot, Amazonbot, Meta-ExternalAgent and others. The Content Signals line says what the content may be used for. The same value is also sent as an HTTP header (from next.config.ts headers) and a <meta name="content-signal"> tag, so a client that ignores robots.txt still sees it.
Allowing training is a business decision. For a studio that sells AI search visibility, blocking the crawlers we tell clients to welcome would be odd.
Honest note: a few entries in that list, like Bard, Gemini and Copilot, are not user-agent tokens any crawler actually sends. They are harmless, but they are noise. The real control for Google's AI use is Google-Extended. If you copy a list like this, check each name against the vendor's documentation.
4. llms.txt, and Markdown for agents that ask for it
public/llms.txt is a hand-maintained Markdown file: a short blockquote saying what the business is, then sections for services, pricing, tools, the 91 profession guide pages and blog posts per language. A longer llms-full.txt sits next to it.
One habit I'd recommend for multilingual sites: say which language a linked page is in. Several service pages only exist in Turkish, so their lines end with "(Service page in Turkish.)". A model summarizing the site for an English speaker then knows not to promise an English page that isn't there.
Discovery is handled in the root layout's <head>:
<meta name="content-signal" content="search=yes, ai-input=yes, ai-train=yes" />
<link rel="alternate" type="text/markdown" href="/llms.txt" title="LLM Context" />
and next.config.ts redirects /.well-known/llms.txt to /llms.txt for clients that look there.
Agents that send Accept: text/markdown (or request a .md path) get Markdown instead of HTML. The middleware rewrites those requests to an API route:
// middleware.ts (excerpt)
const accept = request.headers.get('accept') || ''
const isMarkdownRequest =
(accept.includes('text/markdown') || pathname.endsWith('.md')) &&
!pathname.startsWith('/api/')
if (isMarkdownRequest) {
const markdownUrl = new URL('/api/markdown', request.url)
markdownUrl.searchParams.set('path', pathname)
const res = NextResponse.rewrite(markdownUrl)
res.headers.set('Content-Signal', 'search=yes, ai-input=yes, ai-train=yes')
res.headers.set('Vary', 'Accept')
return res
}
/api/markdown resolves the path: the root returns llms.txt, a blog URL returns the post's Markdown body with a small front matter block (title, author, dates, canonical URL) and its FAQ appended, and profession and case study pages get the same treatment. Because the source content already lives in TypeScript data files, the Markdown comes from the same data as the HTML instead of a second copy.
Vary: Accept is set on both branches so a cache doesn't hand the Markdown version to a browser. Check the real headers with curl -I after deploying; frameworks and CDNs sometimes rewrite Vary.
5. One JSON-LD graph, linked by @id
The site-wide structured data is a single <script type="application/ld+json"> with an @graph. The nodes reference each other by @id instead of repeating themselves:
// components/seo/structured-data.tsx (trimmed)
const graph = [
{
'@type': 'Organization',
'@id': `${SITE_URL}/#organization`,
name: SITE_CONFIG.SITE_NAME,
foundingDate: '2020',
founder: { '@id': `${SITE_URL}/#founder` },
contactPoint: {
'@type': 'ContactPoint',
availableLanguage: ['Turkish', 'English', 'Spanish', 'Arabic'],
},
sameAs: SITE_CONFIG.sameAs,
},
{
'@type': 'Person',
'@id': `${SITE_URL}/#founder`,
worksFor: { '@id': `${SITE_URL}/#organization` },
sameAs: SITE_CONFIG.founder.sameAs,
},
{
'@type': 'WebSite',
'@id': `${SITE_URL}/#website`,
publisher: { '@id': `${SITE_URL}/#organization` },
inLanguage: ['tr-TR', 'en-US', 'es-ES', 'ar-AR'],
},
...extraNodes,
]
sameAs lives in one config file (lib/config.ts) and includes the Wikidata entry, so every profile link is maintained once. That matters for entity recognition: Google and AI tools connect "ModernWebSEO" across platforms through those links.
Page-level schema then sets its own language. The English GEO service page declares inLanguage: 'en-US' on its Service node and adds an FAQPage in English. Schema in the page's language, on the page in that language, is the rule.
Two things I'd fix here:
-
ar-ARreads as "Arabic, Argentina", becauseARis Argentina's country code. Plainar(or a real target country, likear-SA) would be correct. - The global WebPage and Article nodes use
dateModified: new Date().toISOString(), so the date changes on every build. That contradicts the rule I follow in the sitemap (next section). A modified date should change when the content changes.
6. IndexNow after every build, without lying about lastmod
IndexNow lets you tell Bing, Yandex and the other participating engines that a URL changed. Google doesn't use it. The site pings it from npm's postbuild hook:
"scripts": {
"build": "next build",
"postbuild": "tsx scripts/submit-indexnow.ts",
"indexnow:manual": "tsx scripts/submit-indexnow.ts"
}
The script fetches the sitemap (which covers all four locales), keeps only URLs whose lastmod falls in the last 24 hours, and posts them in batches:
// lib/indexnow/client.ts (excerpt)
const API_ENDPOINT = 'https://api.indexnow.org/indexnow'
const BATCH_SIZE = 100
const BATCH_DELAY_MS = 2000
const payload = {
host: config.host,
key: config.apiKey,
keyLocation: `https://${config.host}/${config.apiKey}.txt`,
urlList: batch,
}
// summarized from submitBatch():
// 200/202 = ok, 429 = wait 60s and retry once,
// 403 = key file not reachable, 422 = host mismatch
The key file lives in public/, named after the key. A failure logs and moves on; the script deliberately never calls process.exit(1), because a search engine ping should never break a deploy.
The 24-hour filter only works if lastmod is honest. That is the most important line in the sitemap file:
// app/sitemap.ts (excerpt)
/**
* Section content dates (comment translated from Turkish).
* Build time (new Date()) is NOT used:
* telling Google the whole site changed on every deploy breaks
* its trust in lastmod, and pushes every URL into the postbuild
* IndexNow submission (last-24-hours filter).
*/
const CONTENT_DATES = {
core: '2026-06-19',
localizedServices: '2026-07-31',
hazirSite: '2026-10-09',
// ...
} as const
Blog posts use their own updatedAt or publishedAt. A listing page uses the date of its newest post. When I change a section's content, I change its date by hand. It's manual, and that's the point: the date means something.
Gotcha: postbuild sees the old sitemap. postbuild runs while the new deployment is still building, and the script fetches the sitemap from the live domain. That is the previous deployment's sitemap. A page added in this deploy isn't in it yet. For changes that matter, I run npm run indexnow:manual after the deploy is live.
A checklist you can steal
- Translated slugs need an explicit locale map. Generate page alternates and sitemap alternates from the same map, or diff them in CI.
- Self-referencing canonical per locale. Never canonicalize a translation to the original.
- If you skip a
[lang]segment to stay static, putlanganddiron a wrapper and know what you gave up. - robots.txt: name the AI crawlers you want, and verify every name exists.
- llms.txt: say what the site is in the first lines, and mark pages that only exist in one language.
- Markdown negotiation:
Vary: Accept, then check the real headers. - JSON-LD: one graph,
@idreferences,inLanguageper page, real modified dates. - IndexNow: honest
lastmodfirst, then automate.
All of this maps to the technical and GEO pillars of the 47-step methodology we use on client sites. If you want the background on the AI side, I wrote up what llms.txt is and how to write one, and there's a free robots.txt generator with AI crawler rules on the site. The GEO service page covers how we apply it for clients, and the ModernWebSEO hub links everything else.
You can also look at the live files: llms.txt and robots.txt.
Questions about the hreflang maps or the Markdown route, ask in the comments.

Top comments (0)