If you're tracking whether AI assistants cite your site, you'll hit a boring but real problem fast: every API returns citations in a different shape.
- OpenAI puts them in annotations on the message content.
- Perplexity returns a top-level list.
- Gemini returns grounding chunks, and the URLs are redirect links.
- Claude attaches citation objects to individual text blocks.
This post writes one normaliser for all four, then cleans the URLs so you can count citations by domain and by page. The code is Python, but the shapes are the same in any SDK.
The target shape
Whatever the engine, we want this:
from dataclasses import dataclass
@dataclass(frozen=True)
class Citation:
url: str # canonicalised
domain: str # registrable-ish host, no "www."
title: str | None
cited_text: str | None # the passage that cites it, when the API provides one
cited_text is the most useful field and the least consistently available. When you have it, you can see which sentence of the answer each source supports, which tells you why your page was used.
OpenAI (Responses API with web search)
Citations live on output_text content parts as annotations of type url_citation, with character offsets into the text:
def from_openai(resp) -> list[Citation]:
out = []
for item in resp.output:
if item.type != "message":
continue
for part in item.content:
text = getattr(part, "text", "") or ""
for ann in (getattr(part, "annotations", None) or []):
if ann.type != "url_citation":
continue
snippet = text[ann.start_index:ann.end_index] if ann.end_index else None
out.append(make(ann.url, ann.title, snippet))
return out
The offsets let you recover the exact span of text that cites each source.
Perplexity (Sonar API)
The simplest of the four. The response has a top-level citations array of URLs, and newer responses also include search_results with titles and dates. Inline markers like [1] in the answer text index into that array (1-based):
import re
def from_perplexity(data: dict) -> list[Citation]:
results = data.get("search_results") or []
urls = data.get("citations") or [r["url"] for r in results]
titles = {r["url"]: r.get("title") for r in results}
text = data["choices"][0]["message"]["content"]
# map [n] markers to the sentence they appear in
snippets: dict[int, str] = {}
for sentence in re.split(r"(?<=[.!?])\s+", text):
for n in re.findall(r"\[(\d+)\]", sentence):
snippets.setdefault(int(n), re.sub(r"\[\d+\]", "", sentence).strip())
return [make(u, titles.get(u), snippets.get(i)) for i, u in enumerate(urls, start=1)]
The sentence split is crude, but it's good enough to see which claim each source backs up.
Gemini (Google Search grounding)
With the google-genai SDK and the google_search tool, sources come back in grounding_metadata. Two catches:
-
grounding_chunks[i].web.uriis usually a redirect URL onvertexaisearch.cloud.google.com, not the real page. -
grounding_supportsmaps text segments to chunk indices, which gives youcited_text.
import requests
def resolve_redirect(url: str) -> str:
try:
r = requests.head(url, allow_redirects=True, timeout=10)
return r.url
except requests.RequestException:
return url
def from_gemini(resp) -> list[Citation]:
meta = resp.candidates[0].grounding_metadata
if not meta or not meta.grounding_chunks:
return []
chunks = meta.grounding_chunks
texts: dict[int, str] = {}
for sup in (meta.grounding_supports or []):
for idx in sup.grounding_chunk_indices:
texts.setdefault(idx, sup.segment.text)
return [
make(resolve_redirect(c.web.uri), c.web.title, texts.get(i))
for i, c in enumerate(chunks) if c.web
]
web.title often contains the source domain, which is a handy fallback if the redirect can't be resolved. Cache resolved redirects. You'll see the same ones repeatedly, and resolving them is the slowest step in the pipeline.
Claude (web search tool)
With the web search server tool enabled, Claude's response content is a list of blocks. Text blocks can carry a citations list, and each web_search_result_location citation has url, title and cited_text:
def from_claude(msg) -> list[Citation]:
out = []
for block in msg.content:
if block.type != "text":
continue
for c in (getattr(block, "citations", None) or []):
if c.type == "web_search_result_location":
out.append(make(c.url, c.title, c.cited_text))
return out
Claude also returns web_search_tool_result blocks listing everything it retrieved. That's a different, larger set than what it cited. Track both: being retrieved but never cited means the engine found your page and decided another source was better.
Cleaning URLs so they count correctly
Raw citation URLs are messy. The same page shows up with tracking parameters, fragments, www. and without, and with trailing slashes. Without canonicalising, one page looks like five.
from urllib.parse import urlsplit, urlunsplit, parse_qsl, urlencode
TRACKING = {"fbclid", "gclid", "mc_cid", "mc_eid", "ref", "ref_src"}
def canonical(url: str) -> str:
s = urlsplit(url.strip())
host = (s.hostname or "").lower().removeprefix("www.")
query = [(k, v) for k, v in parse_qsl(s.query, keep_blank_values=True)
if not k.lower().startswith("utm_") and k.lower() not in TRACKING]
path = s.path.rstrip("/") or "/"
return urlunsplit(("https", host, path, urlencode(sorted(query)), ""))
def make(url, title, cited_text) -> Citation:
c = canonical(url)
return Citation(c, urlsplit(c).hostname or "", title, cited_text)
Notably, ChatGPT appends utm_source=chatgpt.com to many outbound links. Strip it for counting, but keep it in mind for analytics, because it's how that traffic shows up in GA4.
For "domain", removeprefix("www.") is a shortcut. If you need correct registrable domains (blog.example.co.uk → example.co.uk), use the tldextract package, which knows the public suffix list.
Counting it
With everything in one shape, the useful aggregates are a few lines:
from collections import Counter
def summarise(all_citations: list[Citation], our_domain: str):
by_domain = Counter(c.domain for c in all_citations)
our_pages = Counter(c.url for c in all_citations
if c.domain == our_domain or c.domain.endswith("." + our_domain))
share = sum(our_pages.values()) / max(1, len(all_citations))
return by_domain.most_common(15), our_pages.most_common(10), share
The top-15 domain list is usually the most revealing output. It tends to show that AI answers in your category lean heavily on a handful of third-party sites (review platforms, forums, a few publishers) that you don't control, but can earn a presence on.
Mentioned vs. cited
Keep these two signals separate in your data:
- Mentioned: your brand name appears in the answer text.
- Cited: a source URL points at your domain.
You can be mentioned without being cited (the model knows you, but sourced the claim elsewhere) and cited without being mentioned (your blog post supported a general claim). They need different fixes, so collapsing them into one "visibility" boolean throws away the most actionable information you have.
Doing this continuously
The code above handles one response at a time. Running it across a prompt set, daily, for every engine, and storing the cited passage alongside each URL, is what a citation tracker does.
Vista AI's AI Citation Tracking is one option if you don't want to maintain it. It captures the full prompt, response and source for every mention, tracks brand-name variations, scores sentiment, and flags inaccuracies so you can go from "the AI said something wrong about us" to the page it got it from.
If you build your own, the most important decision is to store cited_text. A count tells you that you were cited. The passage tells you why.


Top comments (0)