DEV Community

Cover image for Parsing Citations from OpenAI, Perplexity, Gemini and Claude Into One Format
Furqan Khalid
Furqan Khalid

Posted on

Parsing Citations from OpenAI, Perplexity, Gemini and Claude Into One Format

If you're tracking whether AI assistants cite your site, you'll hit a boring but real problem fast: every API returns citations in a different shape.

  • OpenAI puts them in annotations on the message content.
  • Perplexity returns a top-level list.
  • Gemini returns grounding chunks, and the URLs are redirect links.
  • Claude attaches citation objects to individual text blocks.

This post writes one normaliser for all four, then cleans the URLs so you can count citations by domain and by page. The code is Python, but the shapes are the same in any SDK.


Diagram of where OpenAI, Perplexity, Gemini and Claude return citations, all normalised into one Citation record

The target shape

Whatever the engine, we want this:

from dataclasses import dataclass

@dataclass(frozen=True)
class Citation:
    url: str            # canonicalised
    domain: str         # registrable-ish host, no "www."
    title: str | None
    cited_text: str | None   # the passage that cites it, when the API provides one
Enter fullscreen mode Exit fullscreen mode

cited_text is the most useful field and the least consistently available. When you have it, you can see which sentence of the answer each source supports, which tells you why your page was used.

OpenAI (Responses API with web search)

Citations live on output_text content parts as annotations of type url_citation, with character offsets into the text:

def from_openai(resp) -> list[Citation]:
    out = []
    for item in resp.output:
        if item.type != "message":
            continue
        for part in item.content:
            text = getattr(part, "text", "") or ""
            for ann in (getattr(part, "annotations", None) or []):
                if ann.type != "url_citation":
                    continue
                snippet = text[ann.start_index:ann.end_index] if ann.end_index else None
                out.append(make(ann.url, ann.title, snippet))
    return out
Enter fullscreen mode Exit fullscreen mode

The offsets let you recover the exact span of text that cites each source.

Perplexity (Sonar API)

The simplest of the four. The response has a top-level citations array of URLs, and newer responses also include search_results with titles and dates. Inline markers like [1] in the answer text index into that array (1-based):

import re

def from_perplexity(data: dict) -> list[Citation]:
    results = data.get("search_results") or []
    urls = data.get("citations") or [r["url"] for r in results]
    titles = {r["url"]: r.get("title") for r in results}
    text = data["choices"][0]["message"]["content"]

    # map [n] markers to the sentence they appear in
    snippets: dict[int, str] = {}
    for sentence in re.split(r"(?<=[.!?])\s+", text):
        for n in re.findall(r"\[(\d+)\]", sentence):
            snippets.setdefault(int(n), re.sub(r"\[\d+\]", "", sentence).strip())

    return [make(u, titles.get(u), snippets.get(i)) for i, u in enumerate(urls, start=1)]
Enter fullscreen mode Exit fullscreen mode

The sentence split is crude, but it's good enough to see which claim each source backs up.

Gemini (Google Search grounding)

With the google-genai SDK and the google_search tool, sources come back in grounding_metadata. Two catches:

  1. grounding_chunks[i].web.uri is usually a redirect URL on vertexaisearch.cloud.google.com, not the real page.
  2. grounding_supports maps text segments to chunk indices, which gives you cited_text.
import requests

def resolve_redirect(url: str) -> str:
    try:
        r = requests.head(url, allow_redirects=True, timeout=10)
        return r.url
    except requests.RequestException:
        return url

def from_gemini(resp) -> list[Citation]:
    meta = resp.candidates[0].grounding_metadata
    if not meta or not meta.grounding_chunks:
        return []
    chunks = meta.grounding_chunks
    texts: dict[int, str] = {}
    for sup in (meta.grounding_supports or []):
        for idx in sup.grounding_chunk_indices:
            texts.setdefault(idx, sup.segment.text)
    return [
        make(resolve_redirect(c.web.uri), c.web.title, texts.get(i))
        for i, c in enumerate(chunks) if c.web
    ]
Enter fullscreen mode Exit fullscreen mode

web.title often contains the source domain, which is a handy fallback if the redirect can't be resolved. Cache resolved redirects. You'll see the same ones repeatedly, and resolving them is the slowest step in the pipeline.

Claude (web search tool)

With the web search server tool enabled, Claude's response content is a list of blocks. Text blocks can carry a citations list, and each web_search_result_location citation has url, title and cited_text:

def from_claude(msg) -> list[Citation]:
    out = []
    for block in msg.content:
        if block.type != "text":
            continue
        for c in (getattr(block, "citations", None) or []):
            if c.type == "web_search_result_location":
                out.append(make(c.url, c.title, c.cited_text))
    return out
Enter fullscreen mode Exit fullscreen mode

Claude also returns web_search_tool_result blocks listing everything it retrieved. That's a different, larger set than what it cited. Track both: being retrieved but never cited means the engine found your page and decided another source was better.

Cleaning URLs so they count correctly

Raw citation URLs are messy. The same page shows up with tracking parameters, fragments, www. and without, and with trailing slashes. Without canonicalising, one page looks like five.

from urllib.parse import urlsplit, urlunsplit, parse_qsl, urlencode

TRACKING = {"fbclid", "gclid", "mc_cid", "mc_eid", "ref", "ref_src"}

def canonical(url: str) -> str:
    s = urlsplit(url.strip())
    host = (s.hostname or "").lower().removeprefix("www.")
    query = [(k, v) for k, v in parse_qsl(s.query, keep_blank_values=True)
             if not k.lower().startswith("utm_") and k.lower() not in TRACKING]
    path = s.path.rstrip("/") or "/"
    return urlunsplit(("https", host, path, urlencode(sorted(query)), ""))

def make(url, title, cited_text) -> Citation:
    c = canonical(url)
    return Citation(c, urlsplit(c).hostname or "", title, cited_text)
Enter fullscreen mode Exit fullscreen mode

Notably, ChatGPT appends utm_source=chatgpt.com to many outbound links. Strip it for counting, but keep it in mind for analytics, because it's how that traffic shows up in GA4.

For "domain", removeprefix("www.") is a shortcut. If you need correct registrable domains (blog.example.co.uk → example.co.uk), use the tldextract package, which knows the public suffix list.

Counting it

With everything in one shape, the useful aggregates are a few lines:

from collections import Counter

def summarise(all_citations: list[Citation], our_domain: str):
    by_domain = Counter(c.domain for c in all_citations)
    our_pages = Counter(c.url for c in all_citations
                        if c.domain == our_domain or c.domain.endswith("." + our_domain))
    share = sum(our_pages.values()) / max(1, len(all_citations))
    return by_domain.most_common(15), our_pages.most_common(10), share
Enter fullscreen mode Exit fullscreen mode

The top-15 domain list is usually the most revealing output. It tends to show that AI answers in your category lean heavily on a handful of third-party sites (review platforms, forums, a few publishers) that you don't control, but can earn a presence on.

Bar chart of citation share by domain: g2.com 19%, reddit.com 14%, competitor sites, and your own domain at 4%

Mentioned vs. cited

Keep these two signals separate in your data:

  • Mentioned: your brand name appears in the answer text.
  • Cited: a source URL points at your domain.

You can be mentioned without being cited (the model knows you, but sourced the claim elsewhere) and cited without being mentioned (your blog post supported a general claim). They need different fixes, so collapsing them into one "visibility" boolean throws away the most actionable information you have.

Doing this continuously

The code above handles one response at a time. Running it across a prompt set, daily, for every engine, and storing the cited passage alongside each URL, is what a citation tracker does.

Vista AI's AI Citation Tracking is one option if you don't want to maintain it. It captures the full prompt, response and source for every mention, tracks brand-name variations, scores sentiment, and flags inaccuracies so you can go from "the AI said something wrong about us" to the page it got it from.

If you build your own, the most important decision is to store cited_text. A count tells you that you were cited. The passage tells you why.

Top comments (0)