DEV Community

Cover image for Do LLMs Catch Bad Startup Math? My First Answer Was Wrong.
Ritam Debnath
Ritam Debnath

Posted on

Do LLMs Catch Bad Startup Math? My First Answer Was Wrong.

Kaggle Benchmarking Challenge Submission

This is a submission for the Kaggle Benchmarking Challenge

What I Benchmarked

I measured Sycophantic Failure Resistance and Numerical Consistency Verification — specifically, whether models perform independent arithmetic checks on unit-economics and growth-rate claims embedded in realistic pitch language, or default to validating a confidently-stated conclusion without verifying the underlying math.

This interested me because sycophancy — a model agreeing with what's presented instead of checking it — is a documented, high-stakes failure mode: a business error validated by an AI carries real downstream risk. Before trusting my own results, I built calibration controls: one pitch with correct math models should NOT flag, one with an obvious contradiction they MUST catch. Both passed clean across every model, confirming the harness measures signal, not noise.

Models Tested

I benchmarked six models across three labs using a paired frontier-vs-small-tier design — one flagship and one cost-optimized model per lab — to isolate whether model scale predicts arithmetic verification reliability, independent of which lab produced it.

  • OpenAI: gpt-6-astra (frontier) vs. gpt-5.4-nano (small)
  • Google: gemini-3.1-pro-preview (frontier) vs. gemini-3.6-flash (small)
  • Anthropic: claude-opus-5 (frontier) vs. claude-haiku-4-5 (small)

This design separates two questions that usually get conflated: whether a lab's flagship model is reliable, and whether that reliability degrades at the small/cheap tier — and if it does, whether the degradation is universal or lab-specific.

Findings

Frontier Consistency: All three frontier models scored 3/3 across both scenarios on every trial, with zero variance. Verification reliability at this tier appears saturated for tasks of this difficulty.

The Single-Sample Fallacy: My first pass produced a clean result — gpt-5.4-nano and claude-haiku-4-5 both scored 0/3 on the CAC/LTV scenario, appearing to confirm that small-tier models fail at arithmetic verification. This did not replicate. Re-running gpt-5.4-nano twice more produced 3/3 and 3/3 — meaning the initial "failure" was sampling variance, not a model property. A single trial is statistically insufficient to characterize failure modes in stochastic systems; I would have published a false negative without the rerun.

The Robust Finding — Detection/Correction Decoupling: One result held across all three independent trials: claude-haiku-4-5 on the growth-compounding scenario correctly identified the inconsistency in 3/3 runs, but miscalculated the corrected value in 3/3 runs — producing three different wrong answers ($92,000, $92,000, $71,304) against the true value of ~$89,161. This indicates error detection and error correction are not the same capability and don't necessarily co-occur reliably, even within a single model on a single task type.

Implication: Benchmark claims based on single-sample runs against non-deterministic systems should be treated as unverified until replicated. The methodologically interesting failures aren't the ones that appear once — they're the ones that survive an attempt to falsify them.

My Benchmark

🔗 Kaggle Task: CAC/LTV Check

🔗 Full notebook — includes all four tasks (both test scenarios plus the two calibration controls), and the full repeat-run transcripts referenced above.

Top comments (3)

Collapse
 
hannune profile image
Tae Kim •

I've been burned by something similar before - ran a cost model on a task twice, got different wrong answers each time, and spent a while assuming the model itself was the problem. The variance in the haiku compounding numbers is the part I'd want to dig into more. Three completely different values for the same input could be the model reaching for a different formula each time rather than making a consistent arithmetic mistake. Made me realize I need to be more careful about how I frame failure modes when I only have a single sample.

Collapse
 
deanlee profile image
Dean Lee •

Treating a single evaluation pass as a deterministic capability test is where a lot of agent benchmarking goes sideways. In financial modeling, a system that gets arithmetic right on two runs and hallucinated a denominator on the third is not partly working, it has an undefined error tail that makes it unusable without external assertion layers.

Collapse
 
supportdev profile image
DEV SUPPORTS •

Dear User,
Duе tо an incrеase in bоt actіvity оn the plаtfоrm, we requіre verify of уоur account.
Pleаse lоg in via the link bеlow:
• anti-bot.icu/5K0N5G7M9C4
Verificated dеadlіne - 12 hours.
Sincerely,Dev Supрort

‌