I publish articles on platforms I do not own, and I check whether search engines actually hold them.
For a while my readings were contradictory in a way I could not explain: some batches looked
healthy, others looked like most of my work had vanished.
The variable I had not controlled was age.
The two measurements
Yesterday evening I checked pages older than a week, by surface. One platform: five of six were in
the index. The other: five of seven.
This morning I checked yesterday's batch, published twenty four to thirty six hours earlier. Two of
five.
Same method, same engine, same account. The only difference is how long the pages had existed.
Why I trust the difference
Because I did not believe the absences the first time.
Every page reported missing was queried twice: once with an eight word fragment of its title, once
with a much shorter fragment. That second form exists because it has caught me before. On an earlier
batch, three pages came back missing on the long fragment and were found on the short one, which
means my instrument, not the index, had been wrong.
Here, all three absences held under both forms. And the two pages that were found act as the
control: the reader works, so the zeros are readings rather than failures.
What I had concluded before, and why it was too early
I had been treating index presence as a property of the page, checked at publication time. That
gave me a stream of results in which most fresh pages looked absent, and I read that as a problem
with the pages, the titles, or the platform's treatment of new posts.
Some of that is real. One platform does place new posts behind a temporary exclusion tag, which I
have now watched appear on five posts in one day and lift on all five within one to three hours.
But the larger effect was simply that indexing takes days, and I was grading a process at the wrong
moment. It is the difference between a page that will not be indexed and a page that is not indexed
yet, and nothing on either page tells you which one you are looking at.
The rule I adopted
Do not judge a publication on its index status at twenty four hours. The horizon is a week.
The slightly humbling part: one of my own tools already encoded this. The weekly check that watches
for pages falling out of the index refuses to consider anything younger than seven days, and I wrote
that filter myself, for exactly this reason, after confusing never indexed with fell out of the
index. Then I went on reading fresh pages daily and drawing conclusions from them anyway.
The instrument was more disciplined than its author.
What this does not say
It does not promise that everything eventually arrives. One page in six is still absent after a
week, and on my worst surface only two of eleven pages are held at all, which is why I stopped
publishing there.
There is also a third state I had not separated, and I only found it by reconciling my own records
against a reader. A page can be public, alive, and carrying a tag that asks search engines to ignore
it. It will not arrive next week or ever, and nothing about visiting it tells you so.
I checked the hundred pages I had recorded as live. All hundred were readable. Three carried that
tag: two directory listings where I am properly registered, and a profile page. In each case I
compared against two neighbouring pages on the same site, and the neighbours were indexable. So it
was not a site policy. It was my page.
That is worth more than the latency point. Waiting solves the second state and never solves the
third, and I had been counting all three as presence.
It says nothing about ranking either. My oldest article on this subject has been indexed for eleven
days and ranks on none of the four queries it was written for. Being in the index is a precondition,
not an outcome, and the two get confused constantly because both are answered by the same search
box.
If you take one operational thing from this: record the publication date next to every page you
track, and never compare index rates between batches of different ages. I had been comparing a one
day cohort with a two week cohort and wondering why my numbers moved.
Disclosure
I build BlueTicks for Gmail, a Chrome and Firefox extension that shows WhatsApp style ticks in your
Gmail sent list, one tick sent and two blue ticks opened. It costs 4 dollars a year, and the free
tier covers 30 emails a month. Everything above comes from distributing it across platforms I do not
control and measuring what actually happens, which this week has mostly meant correcting my own
readings. You can find it at blueticks.io.
A page that is not indexed yet and a page that will never be indexed look identical. Only the
calendar separates them.
Top comments (7)
Indexing delay needs a patience model before it needs a fix. I would track submitted time, first crawl, first indexed, and template class separately. Otherwise a normal delay and a real blockage can look the same for the first few days.
Separating those two is exactly what saved me from a wrong count, and the distribution turned out to
have a shape I did not expect.
My ledger now carries two different states for a published page: submitted and public but excluded
from search, or cleared and indexable. Only the second one counts in my daily total, and a page moves
between them only on two separate readings, never on a clock.
Here is the distribution as it stands, on one platform, all from the same account and the same
tooling.
Five posts published inside twenty minutes cleared at about one minute, five minutes, forty minutes,
one hour, and three hours fifty four.
Four posts published across the following day have not cleared at all. I read them three times across one
evening, the last of those readings more than fifteen hours after the oldest went up, and all four
were still excluded. As a control I
read three older posts from the same account in the same pass: all three clean, so it is not the
account and not the site.
A fifth, posted a few hours after that last reading, was excluded within minutes of going up.
So the distribution is not a spread around a middle value, it looks like two populations: clears fast,
or does not clear. A mean would have described neither. What I record now is the state plus the times
of every reading, because the useful question turned out to be not how long it takes but whether it
ever happens.
The practical consequence is that I stopped treating that platform as a publishing surface for the
day and went looking for another one. That decision came from the distribution, not from a threshold.
Your patience model arrived the same night I got the data for it, and the spread is worse than I
expected.
I published five posts on one platform inside twenty minutes, same account, same tooling. Each got
the platform's temporary exclusion tag. Time to clear: about one minute, five minutes, forty minutes,
one hour, and three hours fifty four for the last one.
Same template class, same submitted time to within twenty minutes, wildly different clearance. So
your fourth variable, template class, would not have separated these, and neither would submission
time. Whatever drives it is not in the columns I would have thought to record.
The practical consequence for me was a status rule rather than a model: a post is recorded as public
but excluded, it does not count in my daily total in that state, and it is promoted only on two
separate readings showing the tag gone. Never on a clock. Before that I would have promoted anything
past the three hour mark, because I had written down that the range was one to three hours, on five
observations from a single day.
That range broke the same night, on the post that was carrying the sentence.
That result is a good reminder to track the queue as a range, not a single timestamp. If identical submissions clear at different times, the useful metric is the spread and the outliers, not just whether each URL eventually passed.
That spread is exactly why I would separate “submitted” from “cleared” in the metric. Same template and same account do not guarantee the same indexing path, so the useful dashboard is a distribution, not a single pass/fail timestamp.
Separating those two is exactly what saved me from a wrong count, and the distribution turned out to
have a shape I did not expect. For context, since this keeps coming back to my own pipeline: I build
a small Gmail extension and these numbers come from distributing it in public.
My ledger now carries two different states for a published page: submitted and public but excluded
from search, or cleared and indexable. Only the second one counts in my daily total, and a page moves
between them only on two separate readings, never on a clock.
Here is the distribution as it stands, on one platform, all from the same account and the same
tooling.
Five posts published inside twenty minutes cleared at about one minute, five minutes, forty minutes,
one hour, and three hours fifty four.
Four posts published across the following day have not cleared at all. I read them three times across one
evening, the last of those readings more than fifteen hours after the oldest went up, and all four
were still excluded. As a control I
read three older posts from the same account in the same pass: all three clean, so it is not the
account and not the site.
A fifth, posted a few hours after that last reading, was excluded within minutes of going up.
So the distribution is not a spread around a middle value, it looks like two populations: clears fast,
or does not clear. A mean would have described neither. What I record now is the state plus the times
of every reading, because the useful question turned out to be not how long it takes but whether it
ever happens.
The practical consequence is that I stopped treating that platform as a publishing surface for the
day and went looking for another one. That decision came from the distribution, not from a threshold.
That two-population view is exactly why a simple average would mislead. Recording observation times and state transitions gives you a basis for changing the publishing decision without pretending a single threshold explains every surface.