DEV Community

Cover image for The $0 Bug That Cost Us $1,800 in API Calls

The $0 Bug That Cost Us $1,800 in API Calls

Arpit Gupta on June 16, 2026

Last quarter our OpenAI bill went from $620 to $2,480 in 23 days. No new features shipped. No traffic spike. Zero error alerts. Deployment logs we...
Collapse
 
itskondrat profile image
Mykola Kondratiuk •

cost-per-total is worse than useless - it answers a question nobody was asking. I track per-workflow-run and caught a silent retry loop burning 4x before anyone noticed. the granularity is the monitoring.

Collapse
 
arpitstack profile image
Arpit Gupta •

Exactly this. Per-total is a vanity metric that hides the real story.

A silent retry loop burning 4x per workflow run. That's the kind of thing that only shows up when you're tracking at the right granularity. Monthly totals would've just made it look like "slightly elevated usage."

That's exactly what CostReveal tracks, cost at the call level, tagged by feature/workflow/service, so spikes like yours surface in hours not weeks. Would love to know what you're using for per-workflow-run attribution currently. Always curious how people are solving this today.

Collapse
 
itskondrat profile image
Mykola Kondratiuk •

yeah monthly totals flatten everything. 4x retry at run level looks like a spike, same thing aggregated is just noise in the weekly chart.

Collapse
 
theuniverseson profile image
Andrii Krugliak •

The "how much vs what caused it" split is the real lesson. A total tells you the bleeding stopped but never where the wound is, and per-feature attribution is the only thing that turns a scary graph into an actual fix. Was the cause a retry loop or a prompt that quietly ballooned?

Collapse
 
arpitstack profile image
Arpit Gupta •

Neither actually. That's what made it so hard to catch.

No retry loop, no prompt bloat. The prompt was perfectly sized. The model was responding correctly. Every individual call looked healthy.

The bug was structural: the export trigger got accidentally wired into the autosave hook. So GPT-4o was being called every 30 seconds per active session, silently, with zero indication anything was wrong.

That's the worst kind, the wound isn't in any single call, it's in the frequency and frequency only becomes visible when you're rolling up cost by feature over time, not inspecting individual requests.

The "what caused it" question turned out to be a "why is this feature running 2,000 times a day" question. Attribution got us to the feature. The frequency pattern in the daily rollup pointed at the bug.

Collapse
 
theuniverseson profile image
Andrii Krugliak •

That's the scarier version, because every call passing its own check is why per-request monitoring never flags it. The signal only lived at the feature-frequency layer, which nothing was watching. It's why I stopped trusting per-step success in agents and started checking the outcome against something the agent didn't report itself.

Collapse
 
jasmine_park_dev profile image
Jasmine Park •

The $0-bug-that-bills-$1800 is the most relatable LLM-cost story there is. Ours was a retry loop that looked free per-call and was not. What turned these from recurring surprises into caught-in-an-hour was per-feature cost attribution plus a budget alert: tag every call by which feature triggered it, roll up cost-per-feature daily, page when one blows its budget. A single total bill hides exactly this kind of bug because the spike is averaged across everything. Was yours visible in the per-request logs, or only in the monthly total? The gap between those two is where these hide.

Collapse
 
arpitstack profile image
Arpit Gupta • • Edited

Only in the monthly total, honestly.

Per-request logs showed normal latency, normal response codes. Nothing suspicious. The cost was just... averaging out silently.

Took us building feature-level rollups to finally see it. Once we tagged calls by feature and watched daily cost trends, the spike was obvious within 3 days.

That's what CostReveal came out of, as we got tired of finding these in the bill
instead of before it.

How are you handling attribution currently? Manual tagging or something more automated?

Collapse
 
jkson_12ga profile image
Jackson •

Quick question on the SDK setup.

When you tag by tenantId and userId both, does CostReveal let you pivot the dashboard by either dimension independently?
Like can you see all spend for a specific tenant across all features, or only feature first then tenant breakdown?

Collapse
 
arpitstack profile image
Arpit Gupta •

To clarify on the tenant pivot: the dashboard breaks down by Feature, Service, and User independently.

So you can go By User and see every feature that user touched and what it cost, or go By Feature and see which users are driving spend on that specific feature.

The Unit Economics section then rolls this up into cost per user which is what exposed the pricing gap in the post.

No separate tenant grouping out of the box but user level attribution gets you there.

Collapse
 
kugarne profile image
Garne •

The autosave hook quietly calling GPT-4o every 30 seconds is such a nasty one, every single call looks completely healthy on its own. Once you had the feature level rollup running, was daily granularity fast enough to catch something like that, or did you end up needing closer to real time?