If you build static websites, documentation hubs, or content platforms using Next.js, Astro, or 11ty, there is a very high probability that your sitemap.xml is quietly lying to Googlebot every single day.
It probably looks something like this:
// sitemap.ts
export default function sitemap(): MetadataRoute.Sitemap {
return routes.map((route) => ({
url: `https://example.com${route}`,
lastModified: new Date().toISOString(), // 👈 The "innocent" culprit
}));
}
Or maybe you graduated from new Date() and felt clever by doing this:
// Read file modification time from filesystem
const stat = fs.statSync(pagePath);
const lastModified = stat.mtime.toISOString();
On your local machine (localhost:3000), this feels clean. You change a markdown file, run npm run build, inspect sitemap.xml, and see that only the modified file has today’s timestamp. You pat yourself on the back, push to GitHub, and let your CI/CD runner deploy to Vercel, Cloudflare Pages, or AWS.
Then, six weeks later, you check your Google Search Console (GSC) crawl stats and realize:
- Crawl budget is stalling.
- Search engines are ignoring your
lastmodheaders entirely. - Fresh articles take weeks to get indexed while unchanged evergreen pages get repeatedly crawled for no reason.
Here is the post-mortem of how we stumbled into this rabbit hole while building the production wiki engine for Aion 2 Blog & Database, what search engine crawlers actually do when you lie to them, and the bizarre string-parsing bug in git status --porcelain that almost drove us crazy.
1. The CI/CD Shallow Clone Trap
Why does fs.statSync(file).mtime fail in production?
Because modern CI/CD pipelines do not clone repositories the way humans do.
When GitHub Actions, Vercel, or Cloudflare Pages pulls your repo, it runs an optimized shallow clone:
git clone --depth 1 https://github.com/org/repo.git
In a fresh container, Git doesn’t know or care when src/app/pricing/page.tsx was authored four months ago. The filesystem creates every single file at the exact second the runner ran checkout.
As a result:
- Every time you trigger a build—even if you just bumped a dependency in
package.jsonor fixed a single typo in a blog post—every single file in your repository gets a brand newmtime. - Your sitemap claims all 1,500 URLs on your domain were modified at
2026-10-09T08:14:22Z.
What Search Engines Actually Do With This
Googlebot and Bingbot aren't stupid. They maintain an internal hash and crawl history of your pages.
When Googlebot sees your sitemap claim that 1,500 URLs updated in the same minute:
- It schedules requests to crawl them.
- It fetches the pages, compares the DOM / text hashes, and discovers zero diff.
- It notes that your
lastmodmetadata has a 99.8% false-positive rate. -
The Penalty: Search engines flag your domain for Content Churn (artificial freshness manipulation). Google officially states that if a site provides unreliable
lastmodvalues, its crawler dropslastmodfrom its priority queue and falls back to slow heuristic scraping.
You just burnt your site's crawl budget to notify search engines about changes that never happened.
2. "Just Use Git Log!" (And Why That Also Breaks)
The intuitive developer reaction is: "Fine, don't trust filesystem mtime. Let's query Git directly!"
import { execSync } from 'child_process';
function getFileDate(filePath: string): string {
// Query git for the last commit date of this file
const date = execSync(`git log -1 --format=%cs -- ${filePath}`)
.toString()
.trim();
return date;
}
This works locally. But the moment it hits your CI runner:
-
The Depth 1 Amnesia: In a shallow clone (
--depth 1), Git’s commit graph only contains a single commit: the latest HEAD. Any file that wasn't modified in that exact commit has no reachable commit history!git logeither returns empty or attributes every file to the latest merge commit SHA. -
The "Fetch Everything" Trap: You might think, "I'll just add
fetch-depth: 0in my GitHub Action!" But on a large static site with thousands of commits and assets, a full git checkout balloons your CI build time from 20 seconds to 3+ minutes, driving up runner costs.
3. The Nightmare Bug: Murdered by git status --porcelain Whitespace
To solve both local development speed and CI determinism, we designed a hybrid system:
- Clean files in Git: Use committed commit dates from git.
-
Uncommitted local changes (dirty working tree): Fall back to filesystem
mtimeso developers immediately see their edits reflected in build previews.
To detect uncommitted files without overhead, we queried:
git status --porcelain
And then came one of the subtlest, most maddening bugs I’ve encountered all year.
Our initial parser looked completely standard:
// ❌ DO NOT DO THIS
const output = execSync('git status --porcelain --untracked-files=all')
.toString()
.trim(); // 💥 THE FATAL MISTAKE
for (const rawLine of output.split('\n')) {
const line = rawLine.slice(3); // Attempting to strip " M " or "?? "
dirtyFiles.add(line);
}
Can you spot the defect?
According to Git’s official Porcelain V1 format specification:
- Column 1: Index status (
M,A,D,) - Column 2: Worktree status (
M,D,) - Column 3: A mandatory delimiter space (
)
When a tracked file has unstaged modifications in the working tree, the output line starts with a space in column 1:
M src/app/about/page.tsx
M src/app/classes/page.tsx
?? src/app/new-feature/page.tsx
Notice the leading space on line 1?
When you call .trim() on the entire output string, JavaScript strips whitespace from the very beginning of the block.
That means it eats the leading space off the very first line!
Line 1 transforms from:
" M src/app/about/page.tsx"
into:
"M src/app/about/page.tsx"
Then, line.slice(3) slices:
- Line 1:
"src/app/about/page.tsx"becomes"c/app/about/page.tsx"(Missing the leadingsr)! - Line 2:
"src/app/classes/page.tsx"slices correctly.
The consequence? Exactly one file—whichever dirty file was alphabetically first in the repository—silently failed path matching.
It was never detected as dirty. It fell through to the stale fallback. Every single build, one random page silently received an outdated timestamp, corrupting both the sitemap and our Open Graph metadata.
The Fix
Never trim the raw output block. Split first, strip only carriage returns per line, and preserve positional column offsets:
// ✅ Safe Porcelain Parsing
const rawPorcelain = execSync('git status --porcelain --untracked-files=all', {
encoding: 'utf8',
});
const dirtyFiles = new Set<string>();
// Do NOT call .trim() on rawPorcelain!
for (const rawLine of rawPorcelain.split('\n')) {
const line = rawLine.replace(/\r$/, ''); // Strip Windows CRLF safely
if (!line || line.length < 4) continue;
let target = line.slice(3);
// Handle renames: "R old -> new"
const arrow = target.indexOf(' -> ');
if (arrow !== -1) target = target.slice(arrow + 4);
target = target.replace(/^"|"$/g, '').trim();
if (target) dirtyFiles.add(normalizePath(target));
}
4. The Data Dependency Blindspot
Modern static generation doesn't just render static JSX. Most data-dense applications render dynamic external data into static pages.
For example, on Aion 2 Blog & Guides, our guide hub features a real-time Class Builds & Tier List page. The React component lives in src/app/tier-list/page.tsx, but the actual patch ratings and equipment tier data live in src/data/tierList.json.
If an editor updates tierList.json, what happens to src/app/tier-list/page.tsx?
In Git, page.tsx was not modified.
If your build script only checks the commit timestamp of page.tsx:
- The visible UI displays fresh patch calculations.
- The sitemap
lastModifiedand JSON-LD schema tell Google the page hasn't changed in three months!
The Solution: An AST-Assisted Dependency DAG
Instead of treating routes as isolated islands, we built a light AST scanner during pre-build. The generator inspects import statements within each page.tsx:
function scanDataImports(pageFile: string): string[] {
const src = fs.readFileSync(pageFile, 'utf8');
const imports: string[] = [];
const re = /import\s+(?:[^'"]*?\s+from\s+)?['"]([^'"]+)['"]/g;
let match;
while ((match = re.exec(src)) !== null) {
const importPath = match[1];
if (importPath.startsWith('@/data/') || importPath.startsWith('src/data/')) {
imports.push(resolveToFilePath(importPath));
}
}
return imports;
}
When calculating a route's modification date:
$$\text{RouteDate} = \max(\text{Date}(\text{page.tsx}), \max_{dep \in \text{deps}}(\text{Date}(dep)))$$
If tierList.json updates, the /tier-list/ route automatically inherits the newer timestamp without manual intervention.
5. The Production Architecture: Deterministic Content Freshness
Here is the deterministic pipeline we run before next build on aion2.blog. It generates a committed, zero-drift pageDates.ts artifact:
┌───────────────────────────────┐
│ Pre-build: gen-dates.mjs │
└───────────────┬───────────────┘
│
┌───────────────────────┴───────────────────────┐
▼ ▼
┌───────────────────┐ ┌───────────────────┐
│ Git Repo Detector │ │ AST Dependency │
│ Is shallow clone? │ │ Map page.tsx -> │
└─────────┬─────────┘ │ src/data/*.json │
│ └─────────┬─────────┘
│ │
▼ ▼
┌───────────────────────────────────────────────────────────────────┐
│ Date Resolution Engine │
│ 1. Is file dirty in working tree? -> Local mtime (non-UTC offset) │
│ 2. Is shallow CI runner? -> Inherit committed manifest │
│ 3. Clean full git repo? -> `git log -1 --format=%cs` │
│ 4. Route date = max(PageDate, DataDependencies) │
└─────────────────────────────────┬─────────────────────────────────┘
│
▼
┌─────────────────────────────┐
│ Writes src/data/pageDates.ts│
│ Consumed by sitemap.ts & │
│ Schema.org JSON-LD │
└─────────────────────────────┘
The CI Safe-Guard: Conservative Mode
When running on an ephemeral CI runner with a shallow clone, the script activates Conservative Mode:
- Files verified dirty by
git status --porcelainusemtime. - Unmodified files strictly inherit the historical committed dates stored in
pageDates.ts. -
Under no circumstances does it fall back to runner
mtime.
const isShallow = execSync('git rev-parse --is-shallow-repository').toString().trim() === 'true';
if (isShallow) {
console.log('[gen-dates] Conservative mode active: shallow clone detected. Preserving manifest dates.');
}
This guarantees that rebuilding an unchanged commit produces a bit-for-bit identical sitemap with zero content churn.
6. The Results in Production
After rolling out this deterministic pipeline to Aion 2 Blog:
- Zero False-Positive Sitemaps: Our daily CI deployment updates exactly 0 dates unless genuine content or data changes were committed.
- Instant Indexing of True Updates: Within 48 hours of deploying a genuine guide update (e.g. game patch notes or new dungeon maps), Googlebot re-crawled the exact modified URL without wasting cycles on 40+ static pages.
-
No Timezone Desync: Replacing naive
date.toISOString()with local calendar date builders prevented the common bug where editing at 11:30 PM UTC+8 shifted dates across calendar boundaries.
Summary Checklist for Your Own Static Sites
If you operate a content-heavy static site, check your build setup today:
- [ ] Are you using
new Date()insitemap.ts? Remove it immediately. A sitemap where every page shares the build timestamp is treated as noise by search engines. - [ ] Are you using
fs.statSync().mtimein CI? Verify whether your CI uses shallow clones (--depth 1). If so, yourmtimeis the build timestamp. - [ ] Are you trimming raw
git status --porcelainoutput? Don't strip the leading space from line 1. - [ ] Do your pages import dynamic JSON datasets? Make sure changes to those datasets propagate to the parent page’s
lastmod. - [ ] Is your
sitemap.xmltimestamp matching your on-page JSON-LD? Ensure both consume a single source of truth rather than calculating timestamps independently.
Building fast web apps is fun, but making sure your static build doesn't lie to the web indexers is where real production resilience lives.
Have you ever caught your CI/CD runner secretly breaking your metadata? Let me know in the comments!
Top comments (1)
Some comments may only be visible to logged-in visitors. Sign in to view all comments.