DEV Community

Lana Plouffe
Lana Plouffe

Posted on

I pre-registered a prediction that my own finding would fail on this market. The product list held; the advice did not.

This week I published a result I liked: adding a buyer's actual situation to a software buying question changes which products an AI engine names. It reproduced on a second market. Two markets, two hits, no misses.

That is exactly the shape of a finding that is really an artefact of the person writing the questions — which was me. So before running anything else, I picked the market where my own explanation says the effect should fail, wrote the prediction and the numbers that would count as a miss into a file, and committed it before the first API call.

This is that run. Part of the prediction held and part of it failed, and the part that failed is the more interesting half.

The claim being tested

The explanation I had been giving was specific: buyer context does not mainly reorder the incumbents, it admits a class the generic question cannot reach — the self-hosted line. On CRM and on call center software, that is what happened. Both markets have a large, mature line of software you install on your own equipment, and both times the context questions pulled that block onto the list.

A specific explanation predicts its own absence. If that is the mechanism, then a market with no such class behind its incumbents should barely move.

Amazon advertising software is that market. Every bid change in the category runs through Amazon's own advertising API, inside Amazon's cloud, executed by a company Amazon approved as a partner. There is no product you can run on your own hardware, because the thing being automated is not yours to host. I could not name a single self-hosted product in the category to track — and that inability is the prediction, not a gap in the preparation.

What I committed before running it

Read out of the file, which was committed before the first API call:

PRIMARY — SELFHOST regex, raw answer text, per run.
    predict  <= 15/132 on both runs
    MISS     >= 60/132 on either run   (halfway to the ~125 the other two markets hit)

SECONDARY — roster size vs out/amazonppc/runA (58 vendors).
    board's own run-to-run variation: 58 -> 66, +14%
    crm +68%, callcenter +69%
    predict  < +30%
    MISS     >= +50%

TERTIARY — top-10 overlap on the vendor FOLD KEY vs out/amazonppc/runA.
    board's own noise floor (runA vs runB): 9/10
    crm 4/10, callcenter 6/10 and 7/10
    predict  >= 8/10
    MISS     <= 6/10
    Declared in advance: 7/10 is a grey zone on this endpoint and will be reported as one.
    The PRIMARY endpoint decides the headline.

INHOUSE regex: MISS if >= 60/132 on either run.
Enter fullscreen mode Exit fullscreen mode

The primary instrument is a regex over the raw answer text, so it does not depend on what my extractor chooses to count. Its value comes from being the same instrument on all three markets, unchanged.

Control row first

A comparison between two different question sets means nothing until you know what two runs of the same questions do. So, 44 questions, ChatGPT, Gemini and Perplexity, 132 answers per run:

control top-10 overlap
the board vs its own second run 9 of 10
the context run vs its own second run 9 of 10

That is the noise floor. Anything near 10 of 10 is a repeat; a drop below it is the question.

Result 1: the product list did not move. The prediction held here.

market products named, board → with context top-10 overlap
CRM 65 → 109 4 of 10
call center 75 → 127 6 of 10
Amazon advertising 58 → 60 9 of 10

Roster growth of 3.4%, against 68% and 69% on the two markets where the effect fired. The second run gives 65 products, 12.1%, and 10 of 10. 37 of the board's 58 products are on both lists.

Per engine, never averaged — top-10 overlap between the board and its context twin:

engine top-10 overlap
ChatGPT 8 of 10
Gemini 7 of 10
Perplexity 9 of 10

On the two markets where the class exists, the same comparison was 4 of 10 and 6 of 10. Here it is indistinguishable from asking the same questions twice.

Result 2: the answers moved a lot. The prediction failed here.

Same runs. Instead of counting which products were named, count how many answers talk about running it yourself. Same regex, all three markets:

market board with context
CRM 0 of 132 123 of 132
call center 9 of 132 127 of 132
Amazon advertising 0 of 132 49 of 132

I predicted at most 15. It came back 49 and 46.

The second instrument, which tracks the one escape hatch this market does have — being told to do it yourself against Amazon's own console and API — is a clean miss against the number I wrote down:

board run A board run B context run A context run B
answers mentioning the do-it-yourself route 22 21 102 105
— of which ChatGPT 12 8 41 42
— of which Gemini 4 5 41 42
— of which Perplexity 6 8 20 21

My miss line was 60. It came in at 102 and 105 against a board baseline of 22 and 21.

So what actually happened

The engines were given a buyer whose contract says no outside company may hold their trading data. In a market with a self-hosted line, they answer by naming self-hosted products. In this market there are none to name — and they did not invent any. The list of products stayed put, 9 of the top 10 unchanged.

What changed is the advice around the list. The engines started telling the buyer to run it themselves anyway: use Amazon's own console, work the bulk sheets, build against the advertising API. That is not a product recommendation, so it barely touches the ranking, and it is not nothing either — it is the engine telling a buyer in this category not to buy anything.

The 23 products on the context list that are not on the board are mostly that same answer wearing a product's clothes. The largest newcomers:

product answers naming it
Amazon Ads Console 9
Amazon Ads Campaign Manager 7
Kenshoo 4
Marin 4

The only "class" this market had available to admit was Amazon itself.

Scoring myself against what I wrote down

Being honest about this is the entire reason for writing the prediction down first.

  • Roster growth — held. Predicted under 30%, got 3.4%. Miss line 50%.
  • Top-10 overlap — held. Predicted 8 or better, got 9 and 10. Miss line 6.
  • Do-it-yourself language — missed, and by the number I wrote down first, not one I chose afterwards. Miss line 60, got 102.
  • Self-hosting language — my prediction was wrong and my miss line was not reached. I said at most 15 and set the miss at 60; it came in at 49 and 46. I did not declare a verdict for the band in between, which is a flaw in how I wrote the pre-registration rather than a result. I am reporting it as undeclared instead of picking whichever side flatters me.

The mechanism claim I was making is too strong and the corrected version is narrower: buyer context reliably changes what the engine says about how to buy, and only changes which products it names when the market contains a class of product that the generic question was filtering out. Those are two different effects. I had been reporting them as one, because on the first two markets I looked at they happened to move together.

What I can't claim from this

  • This context run is not published as a board, and I am not promoting it to one on the strength of a run designed to attack my own finding. The boards linked below are live with their raw answers; the Amazon context run is not, so those figures are the ones here you cannot go and check for yourself today. I would rather say which is which than blur them.
  • Three engines, named on each board's own ranking.json. No Claude, no Grok — no API key for either, a limit and not a choice.
  • Named is not recommended. I count that a product appeared in an answer to a buying question. "Consider X" and a bare list item count the same. Share of shelf, not endorsement.
  • The rewrite is mine. I wrote the 44 context twins under a rule — describe the buyer's situation, never a product's features — with a grep enforcing that no question contains a vendor name from the board or any of the words the test is about. A different person writing them would get different numbers.
  • One of the newcomers is named by the board's own question. Question 39 on the live board already says "Amazon Campaign Manager", and the context twin inherits that wording verbatim, so it appears identically in both arms and the board's own second run names it 4 times. It is in the table because it is in the data, not because the context introduced it.
  • One market, chosen by me, for a reason I stated in advance. Picking the market where your effect should fail is better than not picking one, and it is still one market.
  • No time series. Each board carries a second run as a repeatability check, not a second date. A product at 1 mention is a level, not a decline.

The raw data

Each board links its own raw files at the bottom: answers-runA.jsonl is the untouched engine responses, mentions-runA.jsonl every extraction, ranking.json the table. Click a product name and you get the questions that named it; click a question and you see the verbatim text each engine returned.

If your product is on one of these boards and the line looks wrong to you, tell me — I would rather fix a board than defend one.

Disclosure, up front rather than buried: connexion.me is mine, the boards and the raw runs are free, and there is a paid monitoring subscription linked from each board page. Nothing here is behind it.

Top comments (0)