DEV Community

Cover image for Building an AI Search Visibility & Brand Analyzer with Gemini, BigQuery, and Google Search Grounding
Hastimal Jangid
Hastimal Jangid

Posted on

Building an AI Search Visibility & Brand Analyzer with Gemini, BigQuery, and Google Search Grounding

AI Search Journey Lab — Part 3 of 7

Measuring how brands appear across AI search journeys using Gemini, Google Search grounding, deterministic visibility scans, and BigQuery history.

Open source: ai-search-journey-lab

Previous: Building Grounded Local Search with Gemini and Google Maps on Google Cloud


The first two articles answered a user question. This one measures the system itself.

In Part 1, I focused on the search journey:

How does one natural-language request become a grounded decision?

In Part 2, I went deeper into the actual local-search implementation:

How do Gemini, Google Places, Search grounding, deterministic ranking, and Maps work together?

Once that worked, I started asking a different question.

Suppose I run the same kind of search repeatedly.

What happens if instead of only asking:

Which business should the user choose?

I ask:

Which brands keep appearing?
Which brands are cited?
Which competitors show up more often?
Does a brand appear in the initial answer, the fan-out journey, or only in supporting evidence?
How does that change over time?

That was the starting point for V3 — AI Visibility.

This article is about turning an AI search workflow into something measurable.


Why AI visibility needs a different data model

A single search result is transient.

If I want to measure visibility, I need history.

One run is interesting.

Fifty runs are analyzable.

Five hundred runs can start showing patterns.

So V3 introduced a new concern into the project: persistence

That is where BigQuery enters the architecture.


What I wanted to measure

For each visibility run, I wanted to capture things such as:

  • target brand
  • competitors
  • original query
  • fan-out tasks
  • mentions
  • citations
  • candidate coverage
  • source coverage
  • ranking position
  • execution timestamp
  • historical trend

The important shift was this:

V1/V2:
user query → recommendation

V3:
many queries → repeatable measurements
Enter fullscreen mode Exit fullscreen mode

That turns the project from a search demo (which we have planned) into an analytics system. :)


Architecture

Building an AI Search Visibility & Brand Analyzer with Gemini


Step 1: I made visibility scans deterministic

One important design choice was that V3 should not rely on a user manually clicking around and interpreting the results.

I wanted repeatable runs.

That means the runner should accept an explicit input scope and execute the same workflow consistently.

Conceptually:

run = VisibilityRun(
    target_brand="Example Brand",
    competitors=[
        "Competitor A",
        "Competitor B",
    ],
    queries=[
        "best local coffee shops in San Antonio",
        "quiet coffee shop for remote work",
    ],
)
Enter fullscreen mode Exit fullscreen mode

Then the runner executes the scan and records the results.

The goal is reproducibility.

If I run the same scenario tomorrow, I want to compare the resulting data rather than rely on screenshots or memory.

AI Search Visibility


Step 2: Separate mention detection from visibility scoring

A brand appearing in a response is not the same as a brand being strongly visible.

For example:

Target brand:
RankRabbit

Response:
"Other platforms include RankRabbit, Competitor A, and Competitor B."
Enter fullscreen mode Exit fullscreen mode

That is a mention.

But now compare:

1. RankRabbit — recommended first
2. Competitor A
3. Competitor B
Enter fullscreen mode Exit fullscreen mode

Those two appearances should not necessarily have the same visibility weight.

So I separate several concepts:

  • mention presence
  • position
  • citation presence
  • fan-out coverage
  • narrative prominence

That lets the application calculate visibility metrics explicitly instead of collapsing everything into one LLM judgment.


Step 3: Capture evidence before calculating visibility

The workflow should preserve what caused a brand to be counted.

Conceptually:

{
  "brand": "Example Brand",
  "mentioned": true,
  "rank": 2,
  "citation_present": true,
  "fanout_coverage": 0.67,
  "evidence_sources": [
    "source-a",
    "source-b"
  ]
}
Enter fullscreen mode Exit fullscreen mode

That data is much more valuable later than a single score such as:

visibility_score = 74
Enter fullscreen mode Exit fullscreen mode

because it lets me explain where the score came from.


Step 4: Persist visibility history in BigQuery

V3 can run in-memory for a quick demo, but historical analysis needs persistence.

I used a BigQuery dataset for that purpose:

Project:
ai-search-journey-lab

Dataset:
ai_search_journey_v3
Enter fullscreen mode Exit fullscreen mode

Let's explore how bigQuery tables look like.....

bq query \
  --use_legacy_sql=false \
  '
  SELECT
    table_name,
    table_type,
    creation_time
  FROM
    `ai-search-journey-lab.ai_search_journey_v3.INFORMATION_SCHEMA.TABLES`
  ORDER BY
    table_name
  '
Enter fullscreen mode Exit fullscreen mode

Bigquery dataset


Step 5: Query historical visibility with SQL

Once I started persisting V3 runs in BigQuery, I no longer had to rely only on what the Streamlit dashboard displayed.

I could query the visibility history directly from the terminal.
That became one of my favorite parts of V3 because it gave me a second way to validate everything the UI was showing.

For all of the examples below, I use the BigQuery CLI:

Once scan results are persisted, SQL becomes one of the most useful tools in the project.

Big Query visibility scans

Top comments (0)