DEV Community

Peter Hallander
Peter Hallander

Posted on

We ran the same ChatGPT prompt three times in a row. Here is what changed.

We wanted to know how much a single ChatGPT answer can be trusted as a measurement. So we ran the same prompt through ChatGPT more than once, with nothing changed between runs, and recorded what came back each time.

The setup

We used DataForSEO's ChatGPT LLM Scraper on 2026-08-16. For each prompt we ran the exact same query three times, minutes apart, same settings every time. Nothing about the prompt or the account changed between runs. Any difference in the answers is the model, not us.

Email marketing apps for Shopify

The prompt was "best email marketing app for Shopify," run three times.

Across the three runs, 16 distinct domains showed up somewhere in the answers. Only 4 of those 16 appeared in all three runs.

Brand Runs appeared in Position
Klaviyo 3 of 3 1st, every time
Omnisend 3 of 3 2nd, every time
Mailchimp 2 of 3 4th twice, absent once

Klaviyo and Omnisend held their spots exactly. Mailchimp did not. It sat at position 4 in two runs and was missing from the answer entirely in the third. Nothing about the prompt changed between that run and the other two.

If you had only run this prompt once, and it happened to be the run without Mailchimp, you would have concluded Mailchimp does not appear for this query. Run it again and that conclusion is wrong.

A wider check: HTML-to-PDF APIs

To see if this held up outside one prompt, we ran 4 prompts about the HTML-to-PDF API category, 3 runs each, for 12 observations total.

66 distinct domains showed up across those 12 observations. Only 5 domains appeared in every run of the specific prompt they showed up on. 4 of 8 brands we were tracking appeared in some runs of a prompt and were missing from others.

The pattern from the email marketing prompt was not a one-off. The names at the top of an answer tend to hold their position. Everything below that is less settled than a single run makes it look.

What it cost to check

Running all 12 observations for the HTML-to-PDF category cost $0.004 per call, $0.048 in total. Checking whether an answer is stable is not expensive. Not checking is the part that costs something, because it produces a wrong number with the same confidence as a right one.

The conclusion

A tool that runs a prompt once and reports "brand visibility: 67%" is reporting a single sample as if it were a measurement. Based on what we saw here, that number would land differently depending on which of the three runs happened to be sampled.

The top of a ranking is stable. Klaviyo held position 1 in every run. Omnisend held position 2 in every run. That part, one run would have told you correctly.

The marginal brand is not stable. Mailchimp came and went. In the wider check, more than half the brands we tracked were inconsistent from run to run. That instability sits exactly where most brands actually are, not at the top, and it is exactly what a business paying for this kind of tracking wants to know: are we in or out, and how often.

A single run cannot answer that question. It can only tell you what happened once.


Measured on 2026-08-16 using DataForSEO's ChatGPT LLM Scraper. Same prompts, same settings, three runs each, minutes apart.

Top comments (0)