Two days ago I published a census of every AI crawler that hit my site over 34 days. The headline was that ChatGPT now fetches my pages for live us...
For further actions, you may consider blocking this person and/or reporting abuse
This is genuinely interesting, and I love that you left the correction in the article because i was already going to brag on this insight to my team. The updated result is more interesting than the original claim and thankfully more correct. Bing rank starts to look less like the destination and more like an entry ticket; very interesting to see that the presentation and content matter more then mere ranking , maybe it's looking for non obvious clues
I wonder whether a useful next experiment would be a matched-pair test: take two pages targeting similar intent, give one a clear comparison table and concise evidence blocks, leave the other unchanged, then watch retrievals and citations over several weeks. That could help separate “the format is causal” from “comparison queries simply get cited more.” Very new territory, but that distinction feels worth chasing.
Thanks — and the matched-pair design aims at exactly the confound I'm stuck on.
The trouble is my own sample can't settle it. The page carrying 62 of the 81 citations is a comparison, and comparison questions are also the ones people bring to an AI assistant in the first place. So "the format is causal" and "comparison intent simply gets cited more" are perfectly collinear in my data — no amount of re-reading the existing logs pulls them apart.
Two things I'd add to your design if I run it. First, publish both pages together and change one of them later, using the unchanged page as its own before/after; otherwise publication age quietly becomes the variable you're measuring. Second, score it on citations rather than impressions, because the whole finding was that rank and citation came apart at the page level: the position-2.81 page has zero citations, while a page that never cracked the top 14 by impressions earned seven.
If I get the pair up I'll post the numbers either way, including the boring outcome where the format does nothing.
This is exactly the kind of boring experimental discipline that makes the result interesting. The comparison-intent confound is a knot you can't untie by staring harder at the same data. Publishing the pair together and then changing one page later is smart—it turns the unchanged page into a little time machine.
I'd be especially curious whether the citation effect appears first in fresh answers before it shows up in search impressions, or whether both move together. If you run it, even a null result would map the terrain.
Don't know yet, and that ordering is the measurement I most want. The awkward part is that the two signals aren't sampled the same way, so "which moved first" is partly an artifact of how often I can look.
Search impressions arrive daily and pre-aggregated — I get them whether or not I was paying attention. Citations only exist at the moment someone asks. There's no impression log for them, so the only way to observe one is to ask on a schedule and record what comes back, which makes my probe interval the floor on how fast I can detect any change at all. If I probe weekly and impressions update daily, impressions will look like they moved first every single time, regardless of what actually happened.
So the answerable version is narrower than the question: does the citation change show up in the first scheduled probe after the edit, while impressions are still flat for several days after? That requires a probe cadence tight enough to distinguish "moved first" from "I looked sooner" — daily probes against a fixed query set, timestamped against the edit, and the edit made on a day nothing else on the site changes.
Agreed on the null result, and the pair is the reason it's worth running rather than a nice-to-have. If both pages move together, the effect wasn't the edit — it was something seasonal or index-wide. That's exactly the outcome I'd otherwise write up as a success, which is the failure mode the unchanged page exists to catch.
Love how carefully you're separating “moved first” from “I happened to look first”—that distinction is where a neat experiment becomes a real one. The sampling clocks are basically two cameras recording at different frame rates, so chronology needs its own calibration.
Would a staggered cadence help: daily citation probes around the edit window, then taper to weekly once impressions respond—or use a few fixed query families so one noisy prompt doesn't impersonate a trend? Your unchanged twin page still feels like the hero here; it keeps a good story from outrunning the evidence.
Staggered cadence is the right shape, with one catch: the taper has to be tied to the edit clock, not to when impressions respond. If I drop from daily to weekly the moment impressions move, the sampling rate becomes a function of the outcome — a citation change that arrives late gets detected late for a reason that has nothing to do with the page. So: daily for a fixed window after the edit, decided in advance (I'd take three weeks, since Bing's refresh on my pages has run one to three), then weekly.
Query families — yes, and they fix a second problem I hadn't named. A single prompt can't separate "the page got cited" from "this particular wording happened to hit it". Three or four families (direct comparison, a price question, "best X for Y", plain informational), each with a few phrasings, scored as the fraction of phrasings that cite the page. One noisy prompt then moves a fraction, not a binary. And the twin gets the identical families on the identical schedule, or it's a control in name only.
The part I still don't have a design for is the answer surface itself. The engine changes under both pages at once — retrieval, model, whatever Microsoft ships that week — and the twin cancels that only if both pages are equally exposed to it. Two comparison pages probably are; a comparison paired with a guide isn't. That's the clause of "similar intent" I'd be strictest about when picking the pair.
The 4.2:1 crawl ratio is a useful operational signal, but I’d be careful turning the 87% overlap into “citations follow the Bing index.” In a separate 100-query ChatGPT shopping test we ran, 75% of the top products matched Google Shopping’s top 3. Different surface and intent, but that contrast suggests the retrieval supply chain may vary by mode. For the follow-up, I’d split the funnel page by page: Bing index/rank → ChatGPT-User fetch → actual citation. Have you found pages that rank in Bing’s top 10 but never trigger ChatGPT-User? Those misses may be more diagnostic than the aggregate bot ratio.
Following up, because I was wrong about my own setup and it turned out to matter.
I said I had never registered in Bing Webmaster Tools. I had — the account was sitting there with 91 days of data I had never once opened. So I can actually answer your question now.
Caveat first: BWT labels that report "Microsoft Copilots and Partners," so it measures the Microsoft AI surface, not ChatGPT specifically. Adjacent to my claim, not identical to it. Same site, same 91-day window (19 May – 17 Aug): Bing organic gave 650 impressions and 4 clicks. AI citations over the same window: 81.
The page-level funnel you asked for:
You called it. My best-ranking page in Bing — position 2.81 — has zero citations. So do 3.28, 4.00 and 4.32. The page taking 62 of the 81 citations ranks worse (4.48) than four pages that were never cited once. And
business-automation-guidepulled 7 citations without appearing in my top-14 Bing pages by impressions at all.So on this site, at page level, rank does not predict citation. The misses were more diagnostic than the aggregate ratio, exactly as you said. What separates the winner looks like format rather than position: it is the only head-to-head "X vs Y vs Z" comparison table in the set.
One thing your funnel split surfaced that I had not considered: citations were flat zero until 23 June, then 25 in the first 30 days and 56 in the last 30. Bing impressions did not move like that at all. Whatever gates citation is not stepping in time with the index.
Fair push, and it makes me separate two things I ran together.
The 87% isn't my measurement — it's Seer Interactive's, over 500+ SearchGPT citations. My own data is just the crawl ratio and the per-page fetch counts. So your shopping result doesn't really contradict it, it bounds it: different surface, different intent, plausibly a different retrieval path. I generalized a SearchGPT-shaped finding into "AI citations" and should have marked that seam.
On your direct question — no, I can't answer it yet, for a reason that's awkward given item 2 of my own checklist: I have no Bing rank data. I never registered in Bing Webmaster Tools. I have the fetch column of your funnel and nothing for the index/rank column in the middle, which is exactly where the misses would show up.
What the fetch side does say is that it's brutally concentrated. One Chatwoot-vs-Intercom comparison took 101 of 1,394 ChatGPT-User hits in 36 days, and most of the site got zero. That shape argues for your point rather than mine: an aggregate bot ratio averages over a large tail of pages that never enter the funnel at all, so it can look healthy while the per-page reality is a handful of winners.
Registering in BWT is now the blocking step for the follow-up. Page-by-page, misses first.