DEV Community

Cover image for You Can't Test Money Controls With a Free Model
Debashish Ghosal
Debashish Ghosal

Posted on AI-assisted

You Can't Test Money Controls With a Free Model

Every budget gate in my v0.2.0 field test was green. Every spend-velocity guard passed. Every tenant cap held. Every ROI flag was clean.

All of it was green-on-vacuum. The local model was priced at $0, so no run could ever breach. My "budget enforcement" was a unit test about arithmetic — it proved the math, not the gate.

When I sat down to test the cost system on HivePlane v0.2.0 — four budgets, spend-velocity guards, tenant spend caps, cost-per-completed-task, ROI flags, an attestation-bound result cache — I discovered that a zero-cost default model makes every money gate structurally untestable. This is the article I wish I'd written when I first mentioned the $0 problem in passing.

What a $0 model hides

Control What the green test "proved" What was actually true
Per-run budget Spend never exceeded the cap Spend was always $0.00 — the cap was unreachable
Spend-velocity guard No anomalous burn paused a run There was no burn — the guard had nothing to detect
Tenant spend cap No tenant was ever refused The cap was set against a denominator of zero
Cost-per-completed-task CPCT reported cleanly completed=False on every event — CPCT was always 0
ROI / expensive_low_value No agent flagged expensive No agent cost anything to flag

A zero-cost default model makes every budget, velocity, and ROI gate structurally untestable. You can write the test, it can pass, and it proves nothing about the seam where money is actually enforced.

The priced profile

The fix wasn't a code change. It was a test-design feature: an opt-in --priced profile that prices the local omlx model and drops the zero-cost prefix. Suddenly the budget gate fires at the right seam — admission, before the expensive work — instead of after the spend, where it's too late.

tip: If your budget enforcement test runs against a model that costs nothing, your test is about arithmetic, not enforcement. Price the model in your test profile, or you'll ship a budget gate that has never refused a real run.

Then the real bugs surfaced:

  • Production usage events were always completed=False, so cost-per-completed-task was always zero.
  • rollover() wiped mid-period spend at the period boundary, so a tenant that spent 80% of its cap lost the accounting at rollover.
  • Tenant caps only enforced for DAY periods — a weekly cap was decorative.
  • The local model's zero-cost prefix meant no run could ever breach, so the "tenant cap refused admission" gate had never once fired.

Each was a seam. Each was invisible until the model cost money.

The cache that outlived the certification

The deeper v0.2.0 money idea — and the one I'm most proud of — is the attestation-bound result cache. A cache hit requires a valid attestation. Re-certification invalidates the entry. The cache key includes the task input, the workload, the manifest version, the bundle hash, the config, and the model tier.

Why this matters: a cache that survives a behavior change is a silent regression you pay to repeat. The agent got worse; you re-certified; the old cached result is still served; you keep charging for an outcome the new artifact can no longer produce. Binding the cache to the attestation means a re-cert automatically evicts the stale result, and the savings you report are only ever savings against current certified behavior.

Field-test scenario S22 proves both halves: a cache hit reuses a result and shows the savings; a re-cert invalidates it. Before that binding, the cache (#547) had no attestation validation at all — a cached result outlived the certification that justified it.

What I learned

The product was fine; the test was lying. Price the model in your test profile — it's a test-design decision, not a code change — or your budget gate has never refused a real run. Enforce at admission, not after the spend. And bind your cache to your certification: a cache without attestation validation is a regression with a discount.

The v0.2.0 field test exercises showback (cost by tenant → team → workload with CPCT) and cache hit + invalidation as scenarios S21 and S22. The priced profile is what made them real. Full evidence in the field-test report; the cost/ROI subsystem is documented in the cost service design.

References

If your model costs $0 in CI, your budget gate has never refused a real run. What's the priced path in your test profile — and when did you last run it?

Top comments (0)