DEV Community

Cover image for Build a Minimal AI Rank Tracker in Python
Furqan Khalid
Furqan Khalid

Posted on AI-assisted

Build a Minimal AI Rank Tracker in Python

If you've ever typed "best X for Y" into ChatGPT to see whether your product shows up, you've already done AI rank tracking, just with a sample size of one.

This post builds a small Python tracker that does it properly. It:

  1. sends the same buyer-style prompt to several engines through their APIs,
  2. detects whether your brand is mentioned and where it sits in the answer,
  3. repeats the run N times and reports a mention rate with a confidence interval instead of a yes/no.

The full script is around 150 lines. The interesting part isn't the API calls. It's the measurement.


Why "rank" is fuzzy inside an LLM answer

A Google SERP has slots. An LLM answer has prose, bullet lists and tables. A useful "rank" for a brand inside an answer is really four signals:

Signal Question it answers
mentioned Did the brand appear at all?
position If the answer is a list, which item was it?
cited Did a source URL point to our domain?
framing Was it recommended or just listed? (covered in a later post)

We'll compute the first three here.

ne prompt sent to several engines N times, each answer scored for mention, position and citation, then aggregated with a confidence interval

Setup

pip install openai requests
export OPENAI_API_KEY=...
export PERPLEXITY_API_KEY=...
Enter fullscreen mode Exit fullscreen mode

Config lives at the top of the script. Brand aliases matter more than you'd expect: LLMs write "HubSpot CRM", "Hubspot" and "HubSpot's free CRM" in the same week.

# tracker.py
import os, re, math, json, time
from dataclasses import dataclass, asdict
from urllib.parse import urlparse

import requests
from openai import OpenAI

BRAND = {
    "name": "Acme Analytics",
    "aliases": [r"acme\s*analytics", r"\bacme\b"],
    "domains": ["acmeanalytics.com"],
}
PROMPT = "What's the best product analytics tool for a 10-person B2B SaaS startup?"
RUNS = 10
OPENAI_MODEL = os.getenv("OPENAI_MODEL", "gpt-4.1")  # use whatever current model you're targeting
Enter fullscreen mode Exit fullscreen mode

Calling the engines

Each engine function returns the same shape: the answer text plus the list of source URLs. Normalising early keeps the rest of the code engine-agnostic.

@dataclass
class Answer:
    engine: str
    text: str
    sources: list[str]

oai = OpenAI()

def ask_openai(prompt: str) -> Answer:
    resp = oai.responses.create(
        model=OPENAI_MODEL,
        tools=[{"type": "web_search"}],   # lets the model browse, like ChatGPT search
        input=prompt,
    )
    sources = []
    for item in resp.output:
        if item.type != "message":
            continue
        for part in item.content:
            for ann in (getattr(part, "annotations", None) or []):
                if ann.type == "url_citation":
                    sources.append(ann.url)
    return Answer("openai", resp.output_text, sources)

def ask_perplexity(prompt: str) -> Answer:
    r = requests.post(
        "https://api.perplexity.ai/chat/completions",
        headers={"Authorization": f"Bearer {os.environ['PERPLEXITY_API_KEY']}"},
        json={"model": "sonar", "messages": [{"role": "user", "content": prompt}]},
        timeout=60,
    )
    r.raise_for_status()
    data = r.json()
    sources = data.get("citations") or [s["url"] for s in data.get("search_results", [])]
    return Answer("perplexity", data["choices"][0]["message"]["content"], sources)

ENGINES = [ask_openai, ask_perplexity]
Enter fullscreen mode Exit fullscreen mode

Gemini (with Google Search grounding) and Claude (with its web search tool) slot in the same way. Each has its own citation format, which is covered in the next post in this series.

Detecting mention and position

Mention detection is a regex over the aliases. Position is trickier. Most "best X" answers come back as a numbered or bulleted list, so we split the answer into list items and find the first item that mentions the brand.

LIST_ITEM = re.compile(r"^\s*(?:\d+[.)]|[-*•])\s+(.*)$", re.MULTILINE)

def mentions(text: str) -> bool:
    return any(re.search(p, text, re.IGNORECASE) for p in BRAND["aliases"])

def list_position(text: str) -> int | None:
    items = LIST_ITEM.findall(text)
    for i, item in enumerate(items, start=1):
        if mentions(item):
            return i
    return None

def cited(sources: list[str]) -> bool:
    hosts = {urlparse(u).hostname or "" for u in sources}
    return any(h == d or h.endswith("." + d) for h in hosts for d in BRAND["domains"])
Enter fullscreen mode Exit fullscreen mode

This is deliberately simple. It will miscount when an answer nests sub-bullets under each recommendation, and the \bacme\b alias will false-positive on unrelated "Acme" mentions. For a real tracker, you'd parse the markdown into a tree and use an LLM-as-judge pass to confirm ambiguous matches. Start with regex, look at the failures, then decide.

Why a single run tells you almost nothing

Run the same prompt twice against any of these engines and you'll often get a different list. Sampling temperature, live search results, retrieval ranking and model routing all add variance.

So the unit of measurement isn't "did we appear?" It's "in what fraction of runs did we appear?" With small N, you need an interval around that fraction. The Wilson score interval behaves well at small sample sizes and near 0% or 100%, which is exactly where brand mention rates tend to sit:

def wilson(k: int, n: int, z: float = 1.96) -> tuple[float, float]:
    if n == 0:
        return (0.0, 0.0)
    p = k / n
    denom = 1 + z**2 / n
    centre = (p + z**2 / (2 * n)) / denom
    half = z * math.sqrt(p * (1 - p) / n + z**2 / (4 * n**2)) / denom
    return (max(0.0, centre - half), min(1.0, centre + half))
Enter fullscreen mode Exit fullscreen mode

To make that concrete: 3 mentions out of 10 runs gives a 95% interval of roughly 11% to 60%. That's the honest answer to "how visible are we in ChatGPT for this prompt?" after ten runs. A screenshot showing you at #1 is one draw from that distribution.

95% Wilson intervals for a 30% mention rate at 1, 5, 10, 30, 100 and 300 runs, narrowing as runs increase

Putting it together

def run():
    rows = []
    for engine in ENGINES:
        for i in range(RUNS):
            try:
                a = engine(PROMPT)
            except Exception as e:
                print(f"{engine.__name__} run {i} failed: {e}")
                continue
            rows.append({
                "engine": a.engine,
                "mentioned": mentions(a.text),
                "position": list_position(a.text),
                "cited": cited(a.sources),
                "n_sources": len(a.sources),
            })
            time.sleep(1)  # be polite to rate limits

    for eng in sorted({r["engine"] for r in rows}):
        rs = [r for r in rows if r["engine"] == eng]
        k = sum(r["mentioned"] for r in rs)
        lo, hi = wilson(k, len(rs))
        positions = [r["position"] for r in rs if r["position"]]
        avg_pos = sum(positions) / len(positions) if positions else None
        cite_rate = sum(r["cited"] for r in rs) / len(rs)
        print(f"{eng:11} mention {k}/{len(rs)}  95% CI [{lo:.0%}, {hi:.0%}]  "
              f"avg pos {avg_pos or '-'}  cited {cite_rate:.0%}")

    with open("runs.jsonl", "a") as f:
        for r in rows:
            f.write(json.dumps({"prompt": PROMPT, "ts": time.time(), **r}) + "\n")

if __name__ == "__main__":
    run()
Enter fullscreen mode Exit fullscreen mode

The output looks like this (the format is real, the numbers are illustrative):

openai      mention 3/10  95% CI [11%, 60%]  avg pos 4.3  cited 10%
perplexity  mention 7/10  95% CI [40%, 89%]  avg pos 2.1  cited 50%
Enter fullscreen mode Exit fullscreen mode

The caveat nobody puts in their screenshots

API answers are not the same as app answers. The ChatGPT app has a system prompt, memory, personalisation and its own search behaviour. The API call above has none of that. Perplexity's sonar API is closer to its product, but it still isn't identical.

That's fine as long as you treat the API as a consistent instrument rather than a perfect replica of what users see. You care about the trend: is the mention rate for this prompt moving up after you shipped that comparison page? A stable, repeatable measurement beats an accurate-looking one-off.

Where this script stops scaling

This works for one prompt. A real prompt set is 30 to 100 prompts across 5 or 6 engines, run daily. At 100 prompts × 6 engines × 10 samples, that's 6,000 calls a day before you've stored, deduplicated or charted anything. You'll also need to handle rate limits, model version changes and the regex false positives above.

If you'd rather not run that infrastructure yourself, Vista AI's AI Rank Tracker does it as a service. It tracks prompts daily across ChatGPT, Gemini, Claude, Perplexity, Copilot and Google AI Mode, rolls the results into a 0–100 visibility index, groups prompts by buyer intent and flags "quick win" prompts where you're close to being recommended. It can also import queries from Google Search Console to seed the prompt list.

Either way, keep the core rule from this post: never report a mention rate without a sample size.

Top comments (0)