DEV Community

Cover image for Measuring Share of Voice in LLM Answers With Pandas
Furqan Khalid
Furqan Khalid

Posted on

Measuring Share of Voice in LLM Answers With Pandas

"Are we visible in ChatGPT?" is a weaker question than "when ChatGPT answers questions in our category, who does it recommend, and how often is it us?"

That second question is share of voice (SOV), and it's the metric that makes AI visibility data useful for competitive decisions. This post computes it from stored LLM answers with pandas, plus two views that SOV alone misses: head-to-head win rates and a co-mention matrix.

It assumes you already have a table of answers. If you don't, the first post in this series builds a collector.


Input data

One row per (answer, brand) pair, which is the score table from the earlier schema, flattened:

import pandas as pd

df = pd.read_sql("""
    SELECT r.id AS run_id, r.run_date, r.engine, p.topic, p.intent_stage,
           b.name AS brand, b.is_ours, s.mentioned, s.position, s.cited
    FROM run r
    JOIN prompt p ON p.id = r.prompt_id
    JOIN score s  ON s.run_id = r.id AND s.scorer_version = 'v1'
    JOIN brand b  ON b.id = s.brand_id
    WHERE r.run_date >= date('now', '-28 days')
""", conn)
Enter fullscreen mode Exit fullscreen mode

Every answer is scored against every tracked brand, so a brand that wasn't mentioned still has a row with mentioned = 0. That's what makes the denominators below correct.

Definition 1: Mention share

The simplest SOV: of all brand mentions across all answers, what fraction were yours?

mentions = df[df.mentioned == 1]
sov = (mentions.groupby("brand").size() / len(mentions)).sort_values(ascending=False)
Enter fullscreen mode Exit fullscreen mode

This is easy to explain and easy to distort. A brand mentioned in passing ("also, tools like X exist") counts the same as the brand recommended first. So weight by position.

Definition 2: Position-weighted share

Give more credit to earlier list positions. A reciprocal weight (1, 1/2, 1/3…) is a common choice. It mirrors how attention drops off down a list. Unlisted mentions get a small flat weight:

def weight(row):
    if not row.mentioned:
        return 0.0
    if pd.isna(row.position):
        return 0.25            # mentioned in prose, not in a ranked list
    return 1.0 / row.position

df["w"] = df.apply(weight, axis=1)
weighted_sov = (df.groupby("brand").w.sum() / df.w.sum()).sort_values(ascending=False)
Enter fullscreen mode Exit fullscreen mode

The 0.25 is a judgement call. Pick a value, document it and keep it fixed. Changing weights mid-quarter makes trends meaningless.

Break it down, or it lies

A single SOV number averages over engines and buyer stages that behave very differently. Pivot it:

def share(g):
    return g.groupby("brand").w.sum() / g.w.sum()

by_engine = (df.groupby("engine").apply(share)
               .unstack("brand").round(3))
by_stage  = (df.groupby("intent_stage").apply(share)
               .unstack("brand").round(3))
print(by_engine)
Enter fullscreen mode Exit fullscreen mode
brand       Acme   Beacon  Corvid  Delta
engine
claude      0.12   0.41    0.29    0.18
gemini      0.22   0.30    0.31    0.17
openai      0.09   0.46    0.27    0.18
perplexity  0.31   0.25    0.26    0.18
Enter fullscreen mode Exit fullscreen mode

(Illustrative.) A pattern like this, strong in Perplexity and weak in OpenAI and Claude, usually points to a gap in how well the brand is known across the wider web rather than a problem with its own pages. Perplexity leans heavily on live retrieval, so good pages get picked up quickly. Answers that rely more on what the model already knows reflect your footprint in reviews, forums and publications.

Heatmap of position-weighted share of voice by engine and brand

Head-to-head win rate

SOV tells you your slice of the pie. It doesn't tell you who specifically beats you, and where. For each pair of brands, look only at answers where both appear and count who ranks higher:

from itertools import combinations

ranked = df[df.mentioned == 1].copy()
ranked["pos"] = ranked.position.fillna(99)   # prose-only mentions rank last
wide = ranked.pivot_table(index="run_id", columns="brand", values="pos")

rows = []
for a, b in combinations(wide.columns, 2):
    both = wide[[a, b]].dropna()
    if len(both) < 10:
        continue                     # too few shared answers to say anything
    a_wins = (both[a] < both[b]).sum()
    b_wins = (both[b] < both[a]).sum()
    rows.append({"a": a, "b": b, "shared": len(both),
                 "a_win_rate": a_wins / max(1, a_wins + b_wins)})

h2h = pd.DataFrame(rows).sort_values("shared", ascending=False)
Enter fullscreen mode Exit fullscreen mode

Read it as: "In the 140 answers where both Acme and Beacon appeared, Acme was listed first 31% of the time." That's a much more actionable sentence than "Acme has 18% SOV."

Diverging bars: Acme listed first in 31% of shared answers vs Beacon, 46% vs Corvid, 64% vs Delta

Co-mention matrix: who the model thinks you compete with

Which brands does the model put in the same answer? This is often not the competitor list in your sales deck:

m = (df[df.mentioned == 1]
       .pivot_table(index="run_id", columns="brand", values="mentioned", fill_value=0))
co = m.T @ m                                   # brand × brand co-occurrence counts
diag = pd.Series(co.values.diagonal(), index=co.index)
jaccard = co / (diag.values[:, None] + diag.values[None, :] - co)
Enter fullscreen mode Exit fullscreen mode

High Jaccard similarity between you and a brand means the model treats you as substitutes. If an unexpected brand shows up with high similarity, the model has placed you in a different category than you intended. If that category is wrong, it's a positioning problem worth more attention than any single prompt.

Sanity checks before you share the numbers

  • Report n. Every SOV cell should carry the number of answers behind it. Hide cells with n < 30.
  • Fix the brand list. Adding a fifth competitor mid-month shrinks everyone else's share. Version the brand set like you version prompts.
  • Separate engines before averaging. One engine with twice as many prompts will dominate a pooled number.
  • Check alias coverage. Run df[df.mentioned == 0] answers through a quick LLM check ("which product brands are named here?") once a month to find aliases your regex misses.

From numbers to actions

SOV shows where you're losing. To see why, look at what the engines cite in the answers your competitor wins. The next post covers diffing those cited pages against yours.

If you'd rather get these views without building the pipeline, Vista AI's Competitor Analysis tracks two to five competitors across ChatGPT, Claude, Gemini, Perplexity, Google AI Mode and Copilot. It shows side-by-side rankings and visibility gaps per engine, lists which competitor pages get cited most, and surfaces content gaps, quick-win queries and emerging topics. It also sends alerts when rankings change.

Build or buy, keep one rule: an SOV number without its denominator and its engine breakdown isn't ready to present.

Top comments (0)