DEV Community

Cover image for I was going to write that a cited page never plateaus. Then I checked.
brainbootdev
brainbootdev

Posted on Originally published at lattice.hashnode.dev

I was going to write that a cited page never plateaus. Then I checked.

I had the sentence written. A cited page does not plateau, it gets cited harder. It followed neatly from the previous piece, it sounded like a law, and the numbers in front of me appeared to support it. Between 16 and 23 August our cited surface went from 950 pages carrying 610,189 Copilot citations to 1,001 pages carrying 737,197. The pages we already had were clearly absorbing most of the gain.

Then I tracked the individual pages instead of the totals, and the sentence turned out to be false.

Four in ten of our cited pages did not grow at all

Of the 946 URLs present in both captures, 60.7% grew, 35.4% did not move, and 3.5% declined. Four more fell out of the cited set entirely. Plateauing is not the exception in this data. It is what happened to more than a third of the pages in a single week.

The first lesson is about the shape of the original claim rather than about the citations. I had a decomposition that looked like evidence: 97% of the week's gain came from pages that were already cited, and only 3% from newly cited ones. That reads like proof that existing pages compound. It is close to arithmetically forced. A page enters the cited set at the reporting threshold, near zero by definition. Our 55 new entrants arrived at a median of 9 citations each. They could not have contributed much of a 127,008 citation gain in their first week no matter how the engine behaved. That statistic would look roughly the same for almost any growing corpus, which means it says almost nothing about ours.

The rates are flat. The volumes are a cliff.

So I split the 946 pages into ten equal bands by size and measured each band's growth rate against its own starting total.

band growth rate share of all new citations
1 (smallest) 24.7% 0.1%
2 65.4% 0.4%
3 29.6% 0.4%
4 35.0% 0.9%
5 32.6% 1.6%
6 26.7% 2.5%
7 27.7% 5.2%
8 27.4% 9.8%
9 18.9% 16.5%
10 (largest) 18.8% 62.5%

Nine of the ten bands grew between 18.8% and 35.0% in the week. The tenth, a small band, grew 65.4% on a median starting count of 8 citations, which is noise rather than a signal. The largest pages grew the slowest of any band, at 18.8%.

The top tenth of pages took 62.5% of every new citation the network gained. The bottom half took 3.3%.

Both of those columns describe the same ten rows of the same table in the same week. That is the finding, and it is why the chart carries two panels rather than one. An average hides it. A single bar chart hides it. You have to draw the rate and the volume side by side before the shape is obvious.

And the order barely moved

Rank correlation between the two captures was 0.989. Forty six of the top fifty pages held their position. 74.5% of pages did not change decile at all.

Put those together and the interpretation runs opposite to the intuition. The engine is not picking winners week to week. It grew nearly every size band at a similar rate. The concentration does not come from large pages being favoured. It comes from a distribution that already spanned four orders of magnitude, compounding at roughly one rate. Multiplicative growth on a skewed base produces a brutally skewed gain even when the multiplier is fair.

That inverts the practical reading of the earlier piece. "Fifty pages carry half the network" sounds like an instruction to grow more pages into the top fifty. In this window that happened to almost nobody. What a page was worth was mostly set before the week began, and the week mostly multiplied it. If that holds over longer horizons, the lever is what a page is when it enters, not how it is nurtured afterwards.

What this does not say

I do not know yet whether it holds over longer horizons. Seven days of rank stability is weak evidence about six months, and I cannot separate the engine's behaviour from our own publishing behaviour inside one window. We hold a third capture from 4 August. The same test across nineteen days is the obvious next move, and I will publish whatever it says, including if it contradicts this.

The measurement, stated plainly:

  • Microsoft Copilot citations from the Bing Webmaster Tools AI Performance API, read per URL across 14 properties.
  • These are page level figures. The property level total for the same network, measured 2 September, is 940,174. Both are correct and they measure different things. Page level runs below property level because pages under Bing's reporting threshold appear in the property total but never as rows. I will not quote one as the other.
  • One engine. Bing is the only one that publishes per property citation counts. This is not a measure of ChatGPT, Claude, Perplexity or Gemini.
  • Citations are not clicks, and they are not crawler fetches. Three separate things.
  • Small band percentages are noisy, for the reason given above.

What I can say is narrower and, I think, more useful than the sentence I nearly published. In this network, over this week, growth in AI citation was multiplicative and close to rank preserving. The average page grew 12%. More than a third grew not at all. And the top tenth, growing slower than eight of the nine bands beneath it, still took nearly two thirds of everything gained.

Most claims being made about AI citation right now cannot be checked by the people making them. Bing is the only engine that will tell you which of your pages it cited and how often. That is one engine out of several, it is a floor rather than a total, and it is still the only place this test can be run at all.


We know which fifty pages. If you want your product named inside one of them, that is a conversation we are happy to have: deepsynthesis.org/lattice

Top comments (0)