I run télé.guide, a free French TV guide. It's a solo project with no database, built with Next.js 15 (App Router) and React 19. It covers 350+ channels: the national free-to-air ones, cable and satellite, and small local stations most guides ignore. It has 8 days of listings, a page for every film that airs, actor pages, and a short "what to watch tonight" article every day.
The site is in French, but the problems aren't. Here are four things I learned building it, including one I haven't fixed yet.
1. Pages never call an external API
Listings come from XMLTV feeds, film data from TMDb, and the daily article from OpenAI. None of them is called while a page renders.
Batch scripts run once a day and write JSON files. Pages read those files and are served with ISR.
XMLTV / TMDb / OpenAI ──▶ batch scripts (cron) ──▶ JSON on disk ──▶ Next.js pages (ISR)
One command runs the whole daily pipeline:
npm run tv:update
# epg:precompute 8 days of listings, each programme matched against TMDb
# guide:editorial short AI-written blurbs for the guide
# ce-soir the daily "tonight on TV" article
# films:update details for every film page
# people:update actor pages and their French bios
A page then does little more than readFileSync plus an in-memory memo, with revalidate between 5 minutes (tonight's grid, film pages) and 1 hour (actor pages).
What this buys me:
- API costs don't grow with traffic. A crawler fetching 2,000 pages costs zero TMDb calls and zero OpenAI tokens.
- An upstream outage breaks a script, not the site. Pages keep serving the last files written.
- Rendering is boring. Reading a local JSON file is the whole data layer.
The rule that took me longest to get right was where each file belongs. I ended up with one question: if I delete this file, will a script rebuild it exactly the same?
- Yes: it goes in
cache/. That covers TMDb responses and the parsed listings (about 100 MB of JSON per day). - No: it goes in
data/, which is backed up and never overwritten.
Two things that look like cache are actually data:
- AI-written text. It cost money and Google has already indexed it. Regenerating it gives a different text, so once published it's content, not cache.
- URL slugs. More on that in section 3.
2. Next.js 15.2+ may put your <title> in <body>
Next.js 15.2 introduced streaming metadata. The page starts streaming before generateMetadata resolves. When it resolves, the tags are appended to <body>. User agents on a built-in list of "HTML-limited bots" (Bingbot, Twitterbot, Slackbot, facebookexternalhit…) still get blocking metadata in <head>.
Googlebot isn't on that list, on purpose. The Next.js docs say they verified that bots which run JavaScript, like Googlebot, read the streamed tags correctly. So this is not a Google ranking bug.
I still turned it off, with one line:
// next.config.mjs
const nextConfig = {
// Next.js 15.2+ streams metadata into <body>; force <head> for every user agent
htmlLimitedBots: /.*/,
};
Why:
-
Plenty of HTML readers don't run JavaScript and aren't on the list. SEO crawlers without JS rendering, link-preview bots from smaller platforms, my own
curlchecks: all of them expect the title, description, canonical and Open Graph tags in<head>. -
For a site that lives on search traffic, I want the raw HTML to be the source of truth. If a page has a metadata problem, I want to see it in
view-source:, not only after a JS render. -
It costs me little. My
generateMetadatareads local JSON (see section 1), so blocking on it barely delays the first byte. If yours waits on a slow API, the trade-off is different: hiding that wait is exactly what streaming is for.
To check your own pages, fetch one with a regular browser user agent and see whether <title> comes before </head>:
curl -s -A "Mozilla/5.0" https://your-site.example/some-page | tr -d '\n' \
| awk '{ t=index($0,"<title"); h=index($0,"</head>"); print (t>0 && t<h) ? "title is in <head>" : "title is NOT in <head>" }'
3. A URL is forever
On a content site, every public URL is a promise. I treat slugs as data, not as something computed on the fly:
-
Every film and every person is keyed by its TMDb ID. The slug (
/films/dracula-1992,/acteurs/tom-hanks) is assigned once, stored in a registry file and never recomputed. If a French title changes, the page keeps its URL. -
When a slug really must change, the old one goes into a
formerSlugslist and answers with a 308 to the new one. -
Redirects take one hop. When I retired the sports section, every old URL got a single 301 from
middleware.ts, straight to its final destination. Old channel slugs get the same treatment, with the alias table computed at build time. -
Dated pages expire on purpose. A listings page for a past date is useless a week later. Past dates answer 410 Gone with
X-Robots-Tag: noindex, so Google drops them instead of crawling them forever. An old route that used to return a 200 copy of the home page for any slug (a textbook soft 404) now returns 410 too. -
Thin pages stay out of the index. An actor page is
noindex, followuntil it has a French biography. About 800 are in the sitemap today; the others wait for their bio.
4. What I haven't fixed: a 6 MB page
Not everything is clean. The evening grid, /programme-tv/soiree, ships 6.2 MB of HTML (about 990 KB gzipped).
I parsed the page to see where it goes. Over 90% of the HTML is the RSC payload, the self.__next_f.push(...) scripts at the bottom (75 of them). The server component passes the day's listings as props to a 'use client' component, so visitors can page through channels, filter films and open a programme's details without waiting for a request:
// app/programme-tv/soiree/page.tsx (simplified)
const initialData = { day: today, channels: channelsSlim, grouped: groupedToday };
return <SoireeClient initialData={initialData} initialDay={today} />;
It's only one day, but one day is 12,243 programmes, and every field of every programme gets serialized:
| Field | Size in the payload |
|---|---|
desc (full description) |
1.77 MB |
icon (image URL) |
1.06 MB |
start + stop
|
0.82 MB |
title |
0.38 MB |
channelId |
0.34 MB |
Full descriptions alone weigh more than the titles, times and channel IDs put together. And channelId is repeated on each programme even though the data is already grouped by channel.
The lesson applies to any App Router project: whatever you pass to a client component ships in the HTML, in full. The fix is a slim view model with only the fields the first screen needs, and the rest loaded on demand. For comparison, the daily article page, which is mostly server-rendered text, weighs 48 KB.
To find your own heavy pages, curl -s URL | wc -c is a good start.
Wrapping up
- Do the expensive work in batch and let pages read files.
- Know where your
<title>ends up. - Treat URLs and published AI text as data, not cache.
- Watch what you pass to client components.
The site is tele.guide (in French). Here's tonight's article if you want to look at the HTML. Questions welcome, especially if you kept streaming metadata on and measured the difference.
Top comments (1)
tr.ee/dev-to