A visibility check I run every day asks the Google Knowledge Graph Search API whether an entity exists, and for a while I stored the answer as a count of results. Today that column said 4. Here is what the four actually were.
43 kg:/g/11njjcc2q5 Marin T. Kael Author
43 kg:/g/11z874h8fr Marin T. Kael (no description)
0.0374 kg:/m/01sp5k Andean condor Bird
The fourth sat below the cut of what I log. Two results are the entity I asked about. One is a vulture.
resultScore is not a scale you can count over
The endpoint does not return only matches. It returns a ranked list, and the tail of that list is whatever the index had lying around. Between the second result and the third, the score falls from 43 to 0.0374. That is three orders of magnitude inside one response, which tells you the values are not a probability and not a percentage of anything. They are a ranking signal with no documented upper bound, and they are not comparable between two different queries.
So a stored count answers a question nobody asked. Four results does not mean four candidate identities. It means one identity, a duplicate of it, and padding.
If you want a number, put a floor on the score and store how many cleared it. Then say in your own report that the floor is arbitrary, because it is. An arbitrary threshold you can see beats a count that hides one.
The thing the count was hiding
The same name resolved to two distinct machine ids. One carries the description Author, the other carries none. A count of 4 cannot show that, and neither can a boolean saying the entity is present.
Duplicates matter because downstream consumers pick one. If a knowledge panel, an answer engine or your own structured data all reach for the same name and land on different ids, they are describing different things that happen to share a label. The one without a description is the weaker of the two, and I would rather know it exists than know there were four results.
What I store now, per query: the top id, its score, the number of results above the floor, and the date the top id was first seen. That last column is the useful one. An id that quietly swaps is invisible to every other field.
Two more queries, both zero
The same run asks two work level questions, for a book title and a series name, both including the author name. Both return zero results. So the person resolves and the works do not.
That is a perfectly reasonable state for the index to be in, and it is also the most useful line in the whole stage, because it separates two things a single yes or no would fuse. Entity recognised is not the same as entity connected to its output. If you only log whether the graph knows the name, you will report success on the day you have the least to celebrate.
The general shape
Every one of these was the same mistake in a different column. A response was reduced to its length instead of its content, and the reduction happened before anything read it. The fix is not a better metric, it is storing the identifier and the score next to the count so that the count can be checked against them later. A number you cannot audit against the raw response is a number you are choosing to trust forever.
I now keep the top three raw results in the snapshot for exactly this reason. It costs a few hundred bytes a day, and it is how I found the condor.
Top comments (0)