Disclosure: I build measurement tooling in this area. Nothing here is a claim about any vendor's product — every error described is mine.
For 19 days in July I ran the same 20 questions through one model every morning and counted which companies came back. Then I looked at what the counter was actually doing.
I threw away 40% of the rows. Only part of that was provably wrong, and separating the two turned out to be the whole exercise.
On naming: I don't name any of the companies involved, or the vendor whose API I used. The bugs were mine, not theirs, and attaching my errors to their names would be a second mistake in the opposite direction. Where a figure is a direct count I say so; where it's derived by arithmetic, or comes from a note I wrote at the time rather than a fresh query, I say that too.
The setup
- 19 consecutive days, 8–26 July 2026, Sydney time, no gaps
- 20 fixed questions, unchanged for the whole run
- One commercial LLM API, one model, held constant
- A watchlist of 71 name entries — I called them companies at the time, but several are open-source projects, which turns out to matter
- Matching, verbatim from the code that ran:
const responseText = rawResponse.toLowerCase();
if (responseText.includes(company.toLowerCase())) {
- Two separate 500s, which I'll come back to: a 500-token cap on the API request, and only the first 500 characters of each response written to the database
That matching rule is where I would have pointed if you'd asked me what went wrong. It turned out to be one of three causes, and not the most interesting one.
20 questions × 19 days = 380 answers. Those 380 produced 1,814 rows — one row per name matched per answer, so about 4.8 rows each. Daily row counts ranged from 79 to 117, a 1.48x spread across days where the questions, the model and the schedule were all identical.
Except it wasn't 380. It was 376. I'll come back to that too.
What "I removed 40%" actually means
I removed 723 rows, 39.9% of the total. For months I described that as a false positive rate. It isn't one. It's three different things with three different causes, and only the first is an error I can demonstrate.
1. Substring matches — 492 rows (27.1%)
Seven names on the watchlist collide with ordinary English — either as whole words, or as strings sitting inside longer words. A substring match has no way to tell a company from a sentence.
Individually, those seven produced between 31 and 152 phantom rows each.
Cause: code. This is the part I can call wrong without qualification: the matched text is not a mention by any reading.
Word boundaries and case sensitivity would have caught the clearest of these. I have not re-run the corrected rule against the original answers, because those answers no longer exist in full — so I'm describing a fix, not a measured improvement.
2. An entry whose count can't be interpreted — 213 rows (11.7%)
One entry accounted for 213 rows by itself. I've since checked every name's count: it is the largest of all 50 names that appeared, ahead of the second by a wide margin, and more than a tenth of the entire dataset.
Two things are true about that entry at once. It is common technical vocabulary in this field — it turns up in ordinary writing about the subject, independently of whether any company is being discussed. And the vendor whose model generated all 380 answers was itself on my watchlist. A model naming its own vendor is expected behaviour, not a bug.
So those 213 rows are some mixture of genuine third-party mention, ordinary vocabulary, and a system talking about itself — and I stored only a 500-character prefix of each answer, so I cannot recover the proportions.
I want to be careful here, because I got this wrong in an earlier draft: I wrote that no matching rule could have saved this entry. That's too strong. Case sensitivity would separate some of it — I've since written exactly that rule. But it would not separate self-reference from genuine mention, and nothing in my stored data can.
Cause: design. I built a list of "companies in this category" without asking which entries were capable of producing an interpretable count — and without noticing that one of them was the thing doing the measuring.
This is why I call it removed rather than wrong. Some unknown share of those 213 were real.
3. Names I no longer have — 18 rows (1.0%)
Twelve names were excluded in the original cleanup. I can reconstruct eight from a note written at the time; those eight total 705 rows. The other four are gone.
I know they account for 18 rows because the arithmetic closes — but the 18 rests on a filtered count of 1,091 that I recorded on 4 August 2026 and have not re-derived from the raw table. I can't re-derive it: the query would need the four names I no longer have. So the 1.0% is real arithmetic on a number I'm trusting my past self about.
I can't recover the four from the totals either. With every name's count now in front of me, 235 different four-name combinations sum to 18.
Cause: record-keeping. The cleanup happened outside anything that kept a log.
These 18 rows aren't false positives at all. They were removed because I'd lost track of them.
Add those up and the honest version is: 492 rows I can show were wrong (27.1%), 213 I can't classify either way (11.7%), and 18 I dropped for bookkeeping (1.0%). After removing all 723: 1,091 rows, 38 names, 57.4 per day.
The part I can't close: false negatives
Everything above is about counting things that weren't there. The harder question is what I failed to count.
21 of the 71 entries on the watchlist produced zero rows across all 19 days.
I know why one of them did. The watchlist entry was missing a punctuation character that appears in the name — I have the file, and the entry is written without it. The string never matched anything. Not one row, across nineteen days, out of 376 answers.
From inside the data, an entry with a typo and an entry genuinely absent from every answer look identical. Both are zero.
I found that one by accident. I have not checked the other twenty.
So the summary of the whole exercise is: I can account for 27% of my rows being wrong, and I have no idea what my false negative rate is.
Four ways to get a zero
A zero in this dataset has at least four possible causes, and from stored data you usually can't tell them apart.
1. The name genuinely wasn't mentioned.
2. The name didn't match — punctuation, spacing, a rebrand, an abbreviation the model used instead. Confirmed at least once, as above.
3. The answer was cut off before it got there. The API request capped responses at 500 tokens, so long answers stopped mid-sentence and anything named after the cut-off could not be counted. I did not store the stop reason, so I can't tell you how many answers actually hit the cap — only that the cap applied to all of them.
4. The call failed, or came back empty. The pipeline had no retry policy and no guard for an empty response. A failed call and an answer containing none of my 71 names produce exactly the same thing: nothing.
Which brings me back to 376. Twenty questions across nineteen days should be 380 question-days. The database has 376. Four answers produced no rows at all, and I cannot tell you which of the four causes above applies to any of them. That's 1.05% of the run, and it's the cleanest illustration I have of the problem: the gap is visible, the reason isn't.
This matters to me because "we don't appear in AI answers" is a claim I was preparing to make. It requires ruling out causes 2, 3 and 4, and I had ruled out none of them.
Two 500s, doing different damage
I spent a while treating these as one limitation. They aren't.
The API cap (500 tokens) truncates the answer before matching runs. It limits what could be counted.
The storage truncation (500 characters) happens after matching, when the row is written. It limits what can be re-examined. I can re-read the first 500 characters of each answer, so I can check matches that happened early — but any match past that point is unverifiable from storage, and 500 characters is a fraction of a 500-token answer.
One caps the measurement. The other caps the audit.
Two more things that were wrong, which I'd have missed
While preparing this I re-read the stored schema rather than just the counts, and found two problems I hadn't reported to myself:
A column that never meant anything. Each row carries a rank_position. I'd assumed it was position within an answer. It was assigned as the row's index in a results array spanning all 20 questions of that day's run — so it encodes the order rows happened to be appended, nothing more. The column isn't noisy. It's meaningless, and it has been in the table since day one.
A date column off by a day. Rows store a run date computed in UTC, while every analysis query I've written converts to Sydney time. The job runs at 01:00 Sydney — I checked, and every row in the table was written in that hour — which is the previous day in UTC. The two columns disagree by one day for all 1,814 rows. I've been reading the converted one, so the figures here hold; had I grouped by the stored column I'd have gotten a different answer to the same question.
Neither of these shows up as a wrong number in a dashboard. Both would have survived any amount of staring at the output.
Limits of this writeup
- Four of the twelve excluded names aren't preserved. 18 rows, 1.0%. I'm reporting arithmetic, not a list — and 235 combinations fit it.
- The 1,091 figure is from a note dated 4 August 2026, not a fresh query, and it can't be re-derived without the four missing names. The 18, the 723, the 39.9%, the 38 and the 57.4 all depend on it.
- Only the first 500 characters of each response were stored. The counts are recoverable; the matching decisions mostly aren't.
- The per-row mention count was hardcoded to 1. Every figure here counts records, not how many times a name appeared inside an answer. Those are different measurements and I've seen them conflated, including by me.
- The pipeline definition is preserved — an exported workflow file whose last-modified date is the day before the run began, and the code quoted above comes from it. I can't prove from inside the file that it's byte-identical to what executed on each of the 19 days; the file carries no internal timestamp, and I have three similarly-named exports.
- One model, one set of 20 questions, 19 days, one category. I haven't tested whether any of this generalises, and I'm not claiming it does.
Reproduction
SELECT COUNT(DISTINCT DATE(created_at AT TIME ZONE 'Australia/Sydney')) AS days,
COUNT(DISTINCT query) AS queries,
COUNT(DISTINCT company_name) AS companies,
COUNT(*) AS rows
FROM geo_research
WHERE DATE(created_at AT TIME ZONE 'Australia/Sydney')
BETWEEN '2026-07-08' AND '2026-07-26';
-- 19, 20, 50, 1814
The 376:
SELECT COUNT(*) FROM (
SELECT DATE(created_at AT TIME ZONE 'Australia/Sydney') AS d, query
FROM geo_research
WHERE DATE(created_at AT TIME ZONE 'Australia/Sydney')
BETWEEN '2026-07-08' AND '2026-07-26'
GROUP BY 1,2) t;
-- 376, not 380
That companies = 50 rather than 71 is the false-negative problem stated as a query result: 21 entries never produced a single row.
The per-name breakdown behind the 492 / 213 / 18 split is the same table grouped by name, which I've left out for the reason at the top. Every piece of arithmetic is in the piece.
What I'd do differently
- Word boundaries and case sensitivity before anything else
- Ask, for every entry on a watchlist, whether it can produce an interpretable count — starting with whether the measuring tool is on the list
- Store the whole response, not a prefix. Storage is cheaper than doubt
- Store the stop reason and the failure, so a zero can be traced
- Write down every exclusion as you exclude it
- Read the schema, not just the totals
- Treat a zero as a question, not an answer
The last one is the only one I'd call a lesson. The rest is just being careful.
Top comments (0)