DEV Community

Jesper Deng
Jesper Deng

Posted on

I Built an AI Search Monitoring Platform Because “Are We in ChatGPT?” Shouldn’t Be a Guess

AI search monitoring is a repeatable way to observe whether a brand appears, is mentioned, and is cited in answers to a stable set of buyer questions.

It is an observation loop, not an official ranking.

I built AI Search Vitals because traditional SEO reports were not answering a question I kept hearing:

When a potential customer asks an AI system for a recommendation, what does it actually say about our brand?

The Problem No One Talks About

Traditional search gives you familiar signals: impressions, clicks, rankings, and backlinks.

AI answers are different.

The same buyer question can produce different results depending on the model, prompt wording, market, language, retrieval mode, source set, and time of day. A screenshot may show what happened once, but it does not explain whether the result is a pattern.

There are four problems I wanted to solve:

  • A brand can be relevant to a category but never appear in the answer.
  • A brand can be mentioned without being cited as a source.
  • A competitor can be recommended while your better explanation is ignored.
  • A single visibility score cannot explain what changed or what to fix next.

The key insight was simple:

The hard part is not generating another score. The hard part is preserving enough evidence to understand the answer behind the score.

What I Decided to Build

AI Search Vitals is an AI search visibility monitoring platform for brands, marketing teams, SEO and GEO practitioners, agencies, and founders.

The workflow starts with a project:

  1. Add a brand, website, market, language, and competitors.
  2. Save the buyer questions you want to monitor.
  3. Select the AI channels enabled for the workspace.
  4. Run an initial baseline.
  5. Review mentions, recommendations, competitors, and citations.
  6. Schedule daily or weekly observations.
  7. Use the results to decide which page, prompt, or content gap to improve.

The product keeps the evidence behind every successful observation:

  • The original buyer prompt
  • The actual model and provider
  • Market and language
  • Search mode and run timestamp
  • Raw AI answer
  • Brand and competitor mentions
  • Ordered recommendation position when available
  • Citation URLs, domains, titles, and snippets
  • Token usage and request metadata

This makes it possible to ask better questions than “Did our score go up?”

For example:

  • Did the brand move from absent to mentioned?
  • Was the mention accurate or misleading?
  • Did the model cite our website or a competitor?
  • Which page became the source?
  • Did the change happen across a prompt group or only once?

AI Search Vitals also includes competitor share of voice, citation exploration, prompt research, monitoring history, CSV export, and GEO audit tools for crawlability, content readiness, and AI Query Expansion.

How It Works Under the Hood

The application runs on Next.js 15, React, and TypeScript, with OpenNext and Cloudflare Workers handling the production runtime.

Workspace, project, prompt, run, observation, mention, and citation data are stored in Cloudflare D1 through Drizzle. OpenRouter adapters keep provider metadata, usage information, citations, and raw response context instead of reducing every response to a plain string.

Scheduled monitoring uses a server-controlled workflow:

  • Cloudflare Cron finds prompts that are due.
  • Each prompt and model target becomes a queue job.
  • The worker executes the observation.
  • The parser extracts mentions, positions, and citations.
  • The observation is stored with an idempotency key.
  • Retryable provider failures can be inspected without duplicating the result.

The goal is not to make AI search look deterministic.

The goal is to make changes comparable.

Why This Is Harder Than It Sounds

1. Prompt stability matters

If the question changes every time, the result is difficult to compare.

That is why AI Search Vitals stores prompt snapshots and encourages teams to separate discovery, evaluation, comparison, and task-based questions. A stable prompt set creates a baseline that can be reviewed over time.

2. Mentions need context

Counting a brand name is not enough.

The system needs to distinguish between a useful recommendation, a passing reference, an incorrect description, and a competitor comparison. A mention without context can lead to the wrong content decision.

3. A mention is not a citation

A model can name a brand without linking to its website.

It can also cite a page for one narrow fact without recommending the whole brand. These are different signals, so they need to be measured separately.

4. Citation formats are inconsistent

Different providers can return citations through annotations, markdown links, metadata, or no structured source list at all.

The parser therefore checks provider annotations first and then uses deterministic URL extraction as a fallback. The result is not perfect, but it is inspectable and tied to the actual answer.

What I Learned Building This

Evidence beats an opaque score

A score can tell you that something changed. Raw answer evidence helps explain why.

The useful unit is not “visibility increased by 8%.” It is:

For this prompt and channel, the answer changed, the brand became more specific, and this page became the cited source.

Measurement boundaries build trust

AI Search Vitals does not claim to reproduce the first-party consumer experience of ChatGPT, Google AI Overviews, or any other platform.

The current product records provider and API proxy observations. These observations are useful for repeatable analysis, but they are not official or universal rankings.

Google AI Overviews still require a manual check in the target market. Estimated intent is directional, not official search volume. A citation is evidence, not a guarantee of future visibility.

GEO and monitoring should be connected

Monitoring tells you what the answer system did.

A GEO audit helps investigate why:

  • Can the page be fetched?
  • Are the headings and metadata clear?
  • Is the entity defined consistently?
  • Does the page contain enough evidence?
  • Are there related questions the content should answer?

That creates a loop:

Observe → inspect evidence → improve the page → observe again.

Try It

Start with a small prompt set instead of monitoring everything at once.

Use five to twelve buyer questions across:

  • Category discovery
  • Product comparison
  • Recommendation
  • Problem solving
  • Brand-specific queries

Read the raw answers before reacting to the metrics. Then compare the same prompts, channel, market, and language after a content change.

At the time of writing, AI Search Vitals includes a seven-day trial without a credit card, so the easiest way to start is with one project and a focused baseline.

You can try AI Search Vitals, read the AI Search Monitoring Guide, or review the AI Citation Tracking Guide.

I am especially interested in hearing how other teams measure AI visibility without turning a changing answer into a fake ranking report.

Top comments (0)