DEV Community

Cover image for A Practical GEO Audit Pipeline with CLI and MCP: What to Check Before AI Search Cites You
L D (一π狐言)
L D (一π狐言)

Posted on

A Practical GEO Audit Pipeline with CLI and MCP: What to Check Before AI Search Cites You

Traditional SEO asks whether a search engine can crawl and rank a page. A GEO audit adds another question: can an AI system retrieve, understand, and trust the page when it is building an answer?

That difference is why I prefer to audit a small set of concrete signals before writing more content. If the crawler path is blocked, if the page intent is vague, or if the entity information is inconsistent, publishing another article usually does not solve the underlying problem.

The audit starts with access

The first check is not keyword coverage. It is whether the relevant crawlers can fetch the page at all.

For public documentation, blogs, and marketing sites, start with:

User-agent: GPTBot
Allow: /

User-agent: OAI-SearchBot
Allow: /

User-agent: ClaudeBot
Allow: /

User-agent: PerplexityBot
Allow: /

User-agent: Google-Extended
Allow: /

Sitemap: https://example.com/sitemap.xml
Enter fullscreen mode Exit fullscreen mode

Treat this as a starting point, not a legal policy. If the site has private areas, block them explicitly instead of using a broad Disallow rule that also removes the pages you want indexed.

After updating robots.txt, verify the live response and check a few representative URLs. A policy that looks correct in the repository but is served from a stale cache is not an audit result.

Check whether the machine-readable layer matches the page

AI systems do better when the page has a clear identity and a consistent description across formats.

The useful checks are:

  • The title and H1 describe the same topic.
  • The page has a canonical URL.
  • The visible copy and JSON-LD describe the same organization, product, or article.
  • Organization, Person, Article, FAQPage, or Product schema is used only when it matches the page.
  • sameAs links connect the same entity across GitHub, LinkedIn, X, and the official website.

This is where many sites lose points. The page tells a human one story, the structured data tells a crawler another, and the AI answer has no reliable entity to cite.

Audit the pages around the main page

Important trust signals often live outside the homepage:

  • /about may contain the organization and author identity.
  • /blog/ and article routes may contain Article schema.
  • /contact, /privacy, and /terms may confirm that the site is a real operating property.
  • /sitemap.xml may expose the actual public page inventory.

A single-page audit can miss all of those signals. A multi-page audit should check the root page, discover linked trust routes, verify their HTTP status, and aggregate the results before scoring.

Capture live HTTP headers

Security headers are not visible in the rendered HTML. They are returned by the server or edge layer:

Strict-Transport-Security
Content-Security-Policy
Enter fullscreen mode Exit fullscreen mode

HSTS and CSP are not magic GEO signals by themselves, but they contribute to the broader trust picture. A tool that reads only the DOM cannot verify them. The fetch pipeline needs to keep the live response headers and pass them into the audit.

Test the pipeline from the CLI

GeoScore ships as an open-source CLI and MCP server. A basic audit can be run with:

npx geoscore https://example.com
Enter fullscreen mode Exit fullscreen mode

For automation:

npx geoscore https://example.com --json
npx geoscore https://example.com --quiet
npx geoscore https://example.com --fix
Enter fullscreen mode Exit fullscreen mode

The JSON mode is useful when the audit needs to feed a dashboard or CI job. The quiet mode returns the score for scripts. The fix mode generates starter files for llms.txt, robots.txt, and JSON-LD, which should then be reviewed by a human before deployment.

Expose the audit to an agent with MCP

The MCP server lets an agent call the same audit primitives without scraping the website UI.

{
  "mcpServers": {
    "geoscore": {
      "command": "npx",
      "args": ["geoscore-mcp"]
    }
  }
}
Enter fullscreen mode Exit fullscreen mode

The audit_url tool accepts a URL and returns the audit result. The generate_fix_files tool produces starter files based on the findings.

This is useful because the agent can keep the audit in the same loop as the code change:

  1. Audit the live URL.
  2. Identify the missing signal.
  3. Edit the page or server configuration.
  4. Rebuild and deploy.
  5. Re-run the audit against the production URL.

That loop is more reliable than asking a model whether a page is "good for SEO."

A repeatable scorecard

For a small site, I would start with six checks:

Check Pass condition
Crawler access Relevant AI crawlers can fetch public pages
Machine-readable identity Title, H1, canonical, and schema describe the same entity
Structured data Only relevant schema types are present and validated
Trust routes About, contact, privacy, and sitemap resolve correctly
Technical trust Live HTTP headers are captured and evaluated
Measurement The audit is repeated after deployment and recorded over time

This will not guarantee an AI citation. No checklist can. It does remove the failures that prevent a page from being eligible in the first place.

Where GeoScore fits

I build GeoScore, a free and open-source GEO audit tool with a browser UI, CLI, and MCP server. It scores a URL across twelve dimensions, including AI crawlability, machine-readable content, structured data, trust signals, and freshness.

You can run the browser audit at:

https://geoscore.help/?utm_source=dev&utm_medium=article&utm_campaign=20260917_geo_audit

The source code is available on GitHub:

https://github.com/qq136692547-cmyk/geo-score

Disclosure: I build GeoScore. The links above are to the open-source project and the hosted audit tool.

Top comments (0)