I publish articles on platforms I do not own, and I check whether search engines actually hold them.
For a while my readings were contradictory in a way I could not explain: some batches looked
healthy, others looked like most of my work had vanished.
The variable I had not controlled was age.
The two measurements
Yesterday evening I checked pages older than a week, by surface. One platform: five of six were in
the index. The other: five of seven.
This morning I checked yesterday's batch, published twenty four to thirty six hours earlier. Two of
five.
Same method, same engine, same account. The only difference is how long the pages had existed.
Why I trust the difference
Because I did not believe the absences the first time.
Every page reported missing was queried twice: once with an eight word fragment of its title, once
with a much shorter fragment. That second form exists because it has caught me before. On an earlier
batch, three pages came back missing on the long fragment and were found on the short one, which
means my instrument, not the index, had been wrong.
Here, all three absences held under both forms. And the two pages that were found act as the
control: the reader works, so the zeros are readings rather than failures.
What I had concluded before, and why it was too early
I had been treating index presence as a property of the page, checked at publication time. That
gave me a stream of results in which most fresh pages looked absent, and I read that as a problem
with the pages, the titles, or the platform's treatment of new posts.
Some of that is real. One platform does place new posts behind a temporary exclusion tag, which I
have now watched appear on five posts in one day and lift on all five within one to three hours.
But the larger effect was simply that indexing takes days, and I was grading a process at the wrong
moment. It is the difference between a page that will not be indexed and a page that is not indexed
yet, and nothing on either page tells you which one you are looking at.
The rule I adopted
Do not judge a publication on its index status at twenty four hours. The horizon is a week.
The slightly humbling part: one of my own tools already encoded this. The weekly check that watches
for pages falling out of the index refuses to consider anything younger than seven days, and I wrote
that filter myself, for exactly this reason, after confusing never indexed with fell out of the
index. Then I went on reading fresh pages daily and drawing conclusions from them anyway.
The instrument was more disciplined than its author.
What this does not say
It does not promise that everything eventually arrives. One page in six is still absent after a
week, and on my worst surface only two of eleven pages are held at all, which is why I stopped
publishing there.
There is also a third state I had not separated, and I only found it by reconciling my own records
against a reader. A page can be public, alive, and carrying a tag that asks search engines to ignore
it. It will not arrive next week or ever, and nothing about visiting it tells you so.
I checked the hundred pages I had recorded as live. All hundred were readable. Three carried that
tag: two directory listings where I am properly registered, and a profile page. In each case I
compared against two neighbouring pages on the same site, and the neighbours were indexable. So it
was not a site policy. It was my page.
That is worth more than the latency point. Waiting solves the second state and never solves the
third, and I had been counting all three as presence.
It says nothing about ranking either. My oldest article on this subject has been indexed for eleven
days and ranks on none of the four queries it was written for. Being in the index is a precondition,
not an outcome, and the two get confused constantly because both are answered by the same search
box.
If you take one operational thing from this: record the publication date next to every page you
track, and never compare index rates between batches of different ages. I had been comparing a one
day cohort with a two week cohort and wondering why my numbers moved.
Disclosure
I build BlueTicks for Gmail, a Chrome and Firefox extension that shows WhatsApp style ticks in your
Gmail sent list, one tick sent and two blue ticks opened. It costs 4 dollars a year, and the free
tier covers 30 emails a month. Everything above comes from distributing it across platforms I do not
control and measuring what actually happens, which this week has mostly meant correcting my own
readings. You can find it at blueticks.io.
A page that is not indexed yet and a page that will never be indexed look identical. Only the
calendar separates them.
Top comments (0)