When an AI answer gets a number wrong, it is rarely the multiplication. It is the setup around it: a percentage taken of the wrong base, a yearly rate split into months as if compounding didn't exist, a figure that quietly changes between step two and step five. The arithmetic is clean, the table is formatted, and the conclusion is off.
You can't catch that by glancing at the result, and re-running the same question usually reproduces the same setup. What works is a second pass that rebuilds every calculated number from your inputs, then checks the base and the units separately from the arithmetic. Below is that prompt, a test run scored against values computed in Python, and what the test does and doesn't prove.
What do wrong numbers in an AI answer actually look like?
Five shapes cover most of them:
- Wrong base. A percent change divided by the new value instead of the old one. Profit going from 35,913 to 39,075 is +8.8%, not +8.1%; both look reasonable.
- Simple vs. compounded. "18% a year is 1.5% a month." It isn't: 1.5% compounded monthly is about 19.6% a year. The compounded monthly rate for 18% is about 1.39%.
- Drifting inputs. A rate, a count, or a number of months that changes between steps without anyone saying so.
- Rounding stated one way, used another. "About 2,207 orders" in the text, 2,206.6 in the formula.
- Percent vs. percentage points. A margin going from 42.6% to 39.3% fell 3.3 points, which is not the same as falling 3.3%.
None of these is an arithmetic error. That's why "double-check your math" as an instruction does so little: the model re-does the multiplication and confirms it.
The copy-paste math-check prompt
Paste it into a fresh chat in ChatGPT, Claude, or Gemini, not the one that produced the answer, with your original question and the AI's answer filled in:
Check the numbers in the AI answer below. Do not trust any figure in it.
1. List every number in the answer that was calculated (not copied from my question). For each, write the formula that should produce it, using only figures from my question.
2. Recompute each one yourself, one operation per line. Work it out before you look at the answer's value.
3. Mark each one MATCH, MISMATCH (give both values), or CAN'T TELL (say which assumption it hinges on).
4. Check the inputs: flag any figure the answer used that isn't in my question, or that changed between steps (a rate, a count, a number of months).
5. Check the base and the units: percentages taken of the right base, monthly vs. yearly, per-order vs. total, percent change vs. percentage points, simple vs. compounded.
6. Finish with one line: the wrong number that would change a decision, if there is one.
My original question:
[paste question]
The AI answer:
[paste answer]
Step 1 is the important one. Writing the formula from your inputs, before looking at the answer's value, is what separates a check from a rubber stamp. Step 5 exists because base and unit errors survive step 2: recomputing a wrong formula gives the same wrong number.
What happens on a real AI answer?
I asked a model (Claude Sonnet, no tools, no code execution) a small-business question with eleven inputs: last year's revenue, cost of goods as a percentage, a monthly subscription, ads for nine months, payment fees of 2.9% plus €0.30 per order across 1,870 orders. Then next year with 18% growth and ads running all twelve months. It asked for both years' profit, the percentage change, and the compounded monthly rate equal to 18% a year.
Then I computed every figure in Python.
The model got all of it right. Last year's profit €35,913.30, next year's €39,074.93, +8.8%, 1.39% a month. It even added that 18% ÷ 12 = 1.5% would compound to about 19.6%. The only blemish: it wrote "about 2,207 orders" and then used 2,206.6 in the fee line, a €0.12 difference.
So the first test was a test of false alarms. I ran the check prompt on that answer in a fresh context (Claude Sonnet, no code execution):
- Every calculated figure: MATCH. No false positives.
- The 2,207 vs. 2,206.6 inconsistency, flagged as CAN'T TELL with "no effect on the result," which is correct.
- An assumption the question left open: whether revenue and costs were both stated net or gross of VAT. I hadn't said, and it was right to ask.
A clean answer is useful to have, but it doesn't show the prompt catches anything. So I made a copy of the same answer with two planted errors, the first two shapes above:
- Percent change divided by next year's profit: +8.1% instead of +8.8%.
- Monthly rate given as 18 ÷ 12 = 1.5%, labeled compounded.
A second fresh context ran the same prompt on the planted copy:
| Item | True value | Planted | Checker verdict |
|---|---|---|---|
| Profit change | +8.8% | +8.1% | MISMATCH, and it named the cause: divided by next year's profit instead of last year's |
| Monthly rate | 1.39% | 1.5% | MISMATCH, "a simple division presented as a compounded rate," with 1.015^12 ≈ 1.196 as proof |
| Every other figure | correct | correct | MATCH |
Two planted errors, two caught, zero false positives across both runs. It also checked a claim hidden in prose, that the ad increase absorbs most of the extra gross margin, and confirmed it: ads take 5,250 of 8,411.63, about 62%.
Where does this prompt stop?
Three limits showed up in the test itself:
- The two runs didn't flag the same things. The first raised the VAT question; the second didn't. Run it once and you get one reading, not a guarantee.
- It contradicted itself once. In the planted run it called the 2,206.6 figure "consistent," then noted the same 2,207 vs. 2,206.6 gap in the very next section, under Inputs. Read the whole output, not just the summary line.
- It checks with the same kind of mental arithmetic that made the answer. That held up here, but a checker and an answerer that share a blind spot will agree on the wrong number.
And the bigger limit: this was one question, two runs, one model. It shows the prompt finds the errors it is built to find. It doesn't show it finds everything.
How do you make the numbers trustworthy?
Use the prompt for triage, then settle anything that matters outside the chat:
- Take the formulas from step 1 and run them yourself. A spreadsheet or three lines of Python. You're checking the model's formulas, which is faster than checking its arithmetic.
- Read step 5's verdict on the base first. Wrong base and simple-vs-compounded are the errors that change decisions, and they're the ones a calculator won't catch for you.
- Treat every CAN'T TELL as a question back to yourself. In the test, the VAT basis was a real gap in my question, not in the answer.
- Fix the input, not the output. If a number is wrong because the setup was wrong, correct the question and regenerate. Patching one figure in a table leaves the figures derived from it wrong.
Can't I just tell the AI to use Python?
If your model can run code, yes, and the arithmetic will be exact. It still won't fix the setup. Code computes 3161.63 / 39074.93 perfectly, and +8.1% is still the wrong answer to "how much did profit grow." The check that matters is whether each number is the right formula for the question you asked, and that is what steps 1 and 5 of the prompt do.
Written with AI assistance. The test question, the original answer, and both checker runs were produced as described (Claude Sonnet, no code execution, separate fresh contexts), and every value was checked in Python. The two errors in the second run were planted by me and are disclosed as such. Results are reported faithfully, including the inconsistency and the run-to-run difference.
This is part of the Verify AI Output series, alongside catching AI hallucinations and checking AI citations. I keep my verification prompts, this one included, in one Notion workspace; if you want that library already organized, there's a free preview of the AI-Augmented Notion Workspace to try first.
Top comments (0)