Your meta description does nothing for ChatGPT. Google half ignores it too, rewriting the thing more often than not if you've ever checked. In the free tier, the one carrying something like 90% of ChatGPT's traffic, that text never gets read. Not by the model, not by whatever decides what shows up in the answer.
A study on 58,000 pages ChatGPT actually pulled just published the real mechanics. What anchors the snippet you see is your H1. Everything after it gets cut at 202 characters, hard stop. 1 in 7 pages doesn't even have an H1. The biggest space thief before that H1: alt text on the first image, up to 50 characters gone before the model reads a single word of your content.
So the question that actually matters isn't "how do I rank on ChatGPT."
It's meaner than that. Does fixing this checklist change what ChatGPT actually cites from me, or just what it shows while waiting for someone to recrawl?
Stop Optimizing Your Meta Description for ChatGPT
You've spent years tuning that field. Character count, keyword up front, a little hook to bump click-through on Google. None of that work reaches the model that increasingly decides whether your page gets mentioned at all.
Here's the split, and it's not subtle. In free and instant mode (the mode most people default to) ChatGPT pulls from a house index built by crawling your page once. That index snippet ignores your meta description completely. It anchors on your H1 and reads roughly 150 useful characters after it before stopping, landing around 202 characters total once you count the H1 itself.
In thinking mode, on a paid account, ChatGPT does something closer to a live Google scrape. That version picks up your meta description about 1 time in 3. Worth keeping the field alive for Google and for thinking mode. Just don't expect it to save your free-tier snippet, because it won't.
I went deeper on the SEO defects hiding from Search Console, a different layer of the same problem, if you want the companion piece.
The H1 Checklist That Actually Controls Your Snippet
6 items, run against sites that scored "perfect" on every audit tool and still fed ChatGPT garbage.
Your H1 exists and says something real. 1 in 7 pages in the study had none. No H1, no anchor, and the model grabs whatever text sits first on the page (a breadcrumb, a date, sometimes literally "Home"). The snippet stops being yours to control once that happens.
Nothing eats the space in front of your H1. The worst offender by far: alt text on the first image, sometimes 50 characters gone before the model reaches your actual heading. Move the image below the fold or shorten the alt text. Either works, pick one.
A real sentence lives in the roughly 150 characters right after your H1. Not a tagline, not a CTA, a sentence that states the point of the page, because that's the window ChatGPT keeps.
noindex does nothing here. Confirmed independently by Jerome Salomon at Oncrawl, cross-checked against the study's own data. Tell Google to skip a page and ChatGPT's reading pipeline caches it anyway.
noindex tells Google to leave. The cache never got the memo.
JSON-LD schema gets stripped before the on-demand read. Whatever facts you buried in structured markup, hoping the model would pick them up, it won't see them there. Put the fact in visible text or don't expect it cited.
Pages over 4MB get rejected outright. No partial read, no graceful fallback. The crawler just moves on.
One thing worth being precise about instead of cutting short: the meta description isn't dead everywhere. Dead in the mode carrying the bulk of actual usage, alive about a third of the time in thinking mode. Keep it for Google. Stop treating it as your ChatGPT lever, because it never was one.
The title tag itself deserves its own pass. I covered the 4-move framework for AI crawlers elsewhere.
Whether fixing this checklist changes what gets cited is a separate question. Hold onto that one, it comes back.
The Copy-Paste Audit Prompt
Checking each of those 6 points by hand on every page you own takes forever. A prompt sorts it faster, so I run this one.
You are auditing a webpage for ChatGPT retrieval readiness.
I will give you a URL or a raw HTML block. Score it against these checks:
1. H1 present? (yes/no, quote the exact text if yes)
2. What text/markup appears BEFORE the H1 in source order? Flag anything over 20 characters, especially image alt text.
3. Length in characters of the first paragraph or sentence immediately following the H1.
4. Page weight in KB if determinable from the HTML.
5. Does the meta description duplicate or contradict the H1 content? (relevant for thinking-mode fallback only)
Return a scorecard, then list the exact edits needed, in priority order, to fix anything that fails.
Paste a URL or the raw HTML, get a scorecard back. Feels a bit like handing Skynet a to-do list, except this one actually reports back instead of going quiet for 3 weeks.
Random, but it's stuck with me: the theme-detection script I run before any audit flags an H1 literally reading "Home" on about 1 in 20 sites. Not a typo, not minimalism, just the WordPress default nobody ever swapped out since setup day 🤦♂️
What SearchBot Actually Saw
The prompt above audits your HTML. It says nothing about what OpenAI's crawler actually saw, which is a different problem entirely. You can fix every item on the checklist and still be flying blind on whether SearchBot came back for the updated version.
For that you need real data, not guesses from source code. I connect an agent to Search Console, Analytics, and the CMS through ChatSEO (full disclosure, that's an affiliate link: https://tolt.link/chatseo) and have it confirm crawl dates and index status directly instead of inferring from a static file. Same tool as the companion piece, different job this time. There it was for the initial audit. Here it's for checking whether a fix actually landed.
Skip that step and you're auditing a page, not a citation. Two different animals, and confusing them is how people fix their H1 and then wonder why nothing changed 3 weeks later.
2 Stores, 2 Robots, 2 Ages

Picture 2 stock rooms. The first holds a photo of your storefront taken the day a robot walked by. The second holds a live look, rebuilt on request, but only if enough time passed since the last look.
The index is the first room. It's the snippet, anchored on your H1, frozen at whatever state your page was in when SearchBot last crawled it. The study found no observed expiration on that freeze. 13% of snippets ran more than a month stale relative to the live page. Fix your H1 today and that fix sits in a queue until the robot happens to walk by again.
The cache is the second room, and it works completely differently. When ChatGPT reads a full page on demand (the thing it does for those 760 pages out of 58,000 it actually opens) it converts the whole page to Markdown and stores that conversion. Fresh for 30 minutes. After that, a new request should trigger a new conversion. Except the study observed retention beyond 90 days on some pages, which doesn't square with a 30-minute freshness window unless the cache gets reused across requests instead of rebuilt each time.
That's exactly what the fingerprint test showed, and it's the strangest result in the whole study, worth its own paragraph instead of a footnote buried at the end. A paid account requests a page. 22 minutes later, a second account (different session, different country, no shared login) requests the same URL. It gets back the identical Markdown conversion, word for word, down to formatting quirks that would be close to impossible to regenerate twice by chance. Zero hits landed on the origin server for that second request, confirmed against the origin's own logs. 1 crawl, 2 accounts served. The header on that page explicitly forbade caching. Cached anyway. It's the ChatGPT equivalent of deja vu in the Matrix, same file served twice, no glitch explanation offered.
Maybe I'm reading too much into 3 observations that happen to line up, but a shared cache that ignores no-store isn't a bug you patch with your own robots.txt. It's infrastructure you don't control and can't opt out of.
So that H1 you just fixed. Who does it actually serve, and starting when?
Why the Free Tier Actually Matters Most
Free and instant mode runs on the house index (labrador, in the study's own naming) for close to 100% of stable, non-time-sensitive questions. No toggle to click, no subscription to remember, just the default most people never leave.
Thinking mode, on a paid account, leans on the live Google scrape (bright, same naming) roughly 3 out of 4 times. Different pipeline, different rules, different snippet mechanics entirely.
So the H1 and 202-character checklist mostly serves the free-tier user. Which, going by the study's own traffic split, is most of ChatGPT's actual usage, not the power user paying for thinking mode and getting a whole different retrieval path underneath.
What This Study Doesn't Prove
Worth being straight about the study itself before trusting a word of the checklist above. Third-party research, published by an SEO/GEO agency (Resoneo) with an obvious incentive to look sharp on exactly this topic. Methodology's public though, and cross-checked against an independent researcher (Jerome Salomon at Oncrawl), which is more transparency than most vendor studies bother with.
1 measurement channel, result_source, disappeared from OpenAI's own responses on July 21, mid-study. Routing between labrador and bright isn't deterministic either. Replay the same prompt and about a third of the time it switches engines on you, which makes any single-run comparison close to useless.
And the big one, the one that actually answers the question I opened with. This study proves what composes the snippet ChatGPT shows once it's decided to show you something. It does not prove that fixing your H1 raises the odds of getting shown in the first place. Those are 2 different metrics, and the second one hasn't been measured directly, not in this study, not anywhere I've found. It's a caching layer without public documentation, not a HAL 9000 conspiracy.
What we know, we know precisely. The H1 anchors the snippet. The 202 characters after it do the rest. The meta description does nothing in free tier, comes back about 1 time in 3 in thinking mode, not more. Fix that and your snippet changes. Verified, not a hunch.
What we don't know stays half known too. Cited or not cited, that jump wasn't measured here, not by anyone in this study.
So I fix the checklist anyway. Not because I've got proof it moves citations. Because the day SearchBot circles back, I'd rather it find a clean H1 than half a page.
Foresight.
Sources
Resoneo, "What ChatGPT pulls, what it shows, what it cites," think.resoneo.com/chatgpt-retrieval, July 2026.
This post may contain affiliate links. If you click them, I might earn a small commission (costs you nothing, and helps me keep shipping quality articles every day for your reading pleasure).
Top comments (0)