DEV Community

Marin T. Kael
Marin T. Kael

Posted on Originally published at marin-t-kael.de

I pre-registered six reliability tests for my AI search measurement. None passed.

Since May I have been asking three answer engines with web search (OpenAI Search, Gemini and Claude on claude.ai) the same 16 questions about a new author. Each answer is scored by fixed rules from -3 to +3. In October the first book comes out, and the obvious next question is whether that launch changes what the engines say.

Before answering that, the measurement itself had to be checked. In May I pre-registered six instrument hypotheses (Q0-INST) with fixed thresholds. The Validation Report Q3/2026 tests them against the data frozen on 1 October, window 11 May to 30 September.

None is confirmed. Three cannot be tested because the planned data collection never took place. Three are not confirmed.

What was tested

Hypothesis Threshold Result
INST-01 test-retest of API surfaces r ≥ 0.9 with 24 h repeat probes not testable, repeat probes never collected
INST-02 multiple snapshots for model probes r ≥ 0.7 with five snapshots and median not testable, single measurement r = 0.746 (OpenAI), 0.612 (Gemini)
INST-03 CUSUM detects a model update detection not testable, no version field for OpenAI and Gemini
INST-04 internal consistency α 0.5 to below 0.7 for model surfaces not confirmed, Gemini α = 0.407
INST-05 Wikidata as a stable anchor coverage > 0.85, stable not confirmed, both registered items deleted
INST-06 Wikidata vs Knowledge Graph κ ≥ 0.8 not confirmed, κ = 0.0

Next-day agreement

Test-retest agreement per question

The same question gets the same score on the next measurement day in 86.6 percent of pairs for OpenAI, 87.1 for Gemini and 94.3 for Claude. Those numbers flatter the instrument. Ten of the 16 questions never mention the author, so almost every answer gets the same score. Without the questions whose score never changes, agreement drops to 78.6, 70.6 and 85.0 percent. The direct question about the person is the least stable: 45.7 percent (OpenAI), 43.8 (Gemini), 73.8 (Claude).

The correlation of two single measurements one day apart is 0.746, 0.612 and 0.811. After subtracting each question's mean, 0.326, 0.367 and 0.759 remain. Most of the apparent reliability comes from the questions being different from each other, not from a day being measured consistently.

One score is not a scale

Cronbach's α across the 16 questions is 0.553 for OpenAI, 0.407 for Gemini and 0.61 for Claude. Depending on the engine, 6 to 10 questions have no variance at all and contribute nothing. Only the three direct questions at Claude reach α above 0.7 as a group (0.784). If you track "AI visibility" as one number summed over a question set, check this first. Here, the categories have to be analysed separately.

Drift without a cause

Daily value per provider with CUSUM alarms

Against the May reference, a tabular CUSUM (k = 0.5, h = 5 SD) raises only upward alarms. That is the visibility increase from the earlier report. With the reference reset after each alarm, five downward alarms appear: OpenAI on 27 August and 30 September, Gemini on 14 August, Claude on 17 August and 19 September.

Each alarm was decomposed by question and checked against the gap register, a log of outages, cadence changes and interventions coded by whether they can shift the level. None of the five is explained. The interval means sat 1.58 to 4.01 percentage points below their reference.

The Claude alarm on 19 September is the clearest. The daily value fell to -2.08 percent after months between 9 and 19, driven by the questions about the person and the book. A sample of the answer texts shows that web search found no page about the author on those days; the models named a well-known namesake or said they found nothing, and the rubric scored part of that as a hallucination (-3). So the drop mixes a change in search results with unknown cause and a scoring rule that does not separate "not found" from "confused with someone else".

Anchors that disappear

The two Wikidata items registered as anchors in May were found deleted on 25 June. Successor items exist since 2 June and were stable from then on, but they do not cover the first three weeks. Google's Knowledge Graph returns the author on 140 of 141 days and the book on none of 139. κ between the two surfaces is 0.0, partly by construction, since Wikidata shows a hit on every day.

Coverage

Gemini was measured completely on 102 of 113 scheduled days, OpenAI on 53 of 113 (mostly an exhausted quota), Claude on 55 of 106. For Claude, 17 missing days have no register entry, including ten days from 22 June to 1 July without any run. Their cause cannot be determined afterwards.

What this means for the launch

The effect window runs from 8 October to 7 November. The report sets these rules for it:

  • Effects are stated per engine and question group, and attributed to the launch only if the shift appears at more than one engine.
  • Shifts of 1.58 to 4.01 points happened without a known cause, so effects of that size cannot be separated from instrument drift.
  • No sum score over all 16 questions, no cross-engine headline on days when one engine is missing, no claims about model updates at OpenAI and Gemini while the answers carry no version field.

Four fixes for the measurement follow: run the repeat probes Q0-INST specified, store model and version with every answer, review the scoring rule for "not found but names a namesake" and publish the result as a new codebook version, and secure the OpenAI quota so the effect window has no gaps.

Limits

n = 1 at the author level. Agreement on consecutive days does not separate real change from noise. The event coding for the alarm check was fixed before the formal check but with knowledge of the alarm days, and was not reviewed independently. The answer-text review for the 19 September alarm is a sample. All questions but one are in German.

Data and code

Deutsche Fassung: marin-t-kael.de/research/berichte/q3-2026-validierung

Top comments (0)