DEV Community

Cover image for LLM pricing pipeline bug: when a relevance model set a second-hand price by itself

LLM pricing pipeline bug: when a relevance model set a second-hand price by itself

Christian Anderson on September 12, 2026

I run a small self-hosted tool that prices second-hand items for eBay. You photograph the thing, it works out what it is, gathers evidence about wh...
Collapse
 
anasbuilds997 profile image
anassBld •

"Where the boundary sits: at the action, not the number." This is the cleanest formulation of mutation gating I've read all week.

In our multi-agent pipelines, we arrived at the exact same architectural boundary after watching models happily produce high-confidence hallucinations:

  1. Inference is an untrusted proposal, not an authority: A model's output (whether a pricing estimate, a classified SKU, or an API call payload) must be treated strictly as an unverified proposal. It has zero intrinsic permission to execute an external side effect.

  2. The gate must live in the code path that spends money: Prompt admonitions ("be careful not to price below X"), warning logs, and LLM self-reflection are fine heuristics, but they are terrible security gates. Real business invariants (source corroboration, rate-of-change ceilings, hard margin floors) belong as deterministic assertions inside the leaf executor immediately before the external mutation is committed.

  3. Attested overrides over ad-hoc UI bypasses: Your point about tracking whose number went through is critical. When a reviewer overrides a flagged "thin" price, bypassing the system via a separate web UI breaks traceability. Binding the human reviewer's approval directly to the execution receipt (approved_by, override_reason, target_hash) preserves the audit trail while keeping the pipeline moving.

Treating models as untrusted proposal generators and enforcing deterministic invariants at the mutation boundary is the only sustainable way to run unattended systems.

Collapse
 
c1-anderson profile image
Christian Anderson •

Glad that line landed. The override part is where I got it wrong first time round: the block refused a person's own typed price too, so the only way past it was eBay's own site, which left no trail at all. Now a price someone types goes through with their name on the approval, and there's a strict switch for when the person approving isn't the one whose money it is. I hadn't thought about hashing the target like you describe, that's a nice touch.

Collapse
 
anasbuilds997 profile image
anassBld •

Target hashing is what stops the classic race where a human approves price X for item A, but the worker re-queries and commits price X against item B (or a mutated version of A). Binding the approval hash directly to (canonical_item_id, target_sku, proposed_price) means the executor can verify the target hasn't drifted under the hood before it spends money.

Keeping the override inside the audit trail rather than forcing operators to use the vendor UI is huge. Once they bypass the tool to fix a false positive, you lose visibility into where your bounds were actually too rigid.

Collapse
 
eternaclarity profile image
Jesse Gamble •

This is a good example of why model output should be treated as evidence, not authority. Once a number can cross directly into a pricing decision, the boundary needs to validate its type, provenance, and permitted range before the suggestion becomes business state.

Collapse
 
c1-anderson profile image
Christian Anderson • • Edited

Yeah, that's pretty much where I ended up. The ceiling checked the range but never asked where the £14.99 came from, so a charger "for Nest Mini" got treated as the speaker itself. What fixed it in the end wasn't a smarter model, it was two boring rules: if the product's name only shows up after "for", it's an accessory, and a discovered new price isn't allowed to cap anything when used ones are asking more than double it.

Collapse
 
eternaclarity profile image
Jesse Gamble •

Two boring rules beating a smarter model is the whole story in one sentence. The 'for' check is the kind of thing that feels too dumb to work until it does.

Collapse
 
eternaclarity profile image
Jesse Gamble •

This is a strong example of why model output needs a typed boundary before it can become business data. A model can suggest a candidate value, but prices, IDs, balances, dates, and other source-of-truth fields need deterministic validation against the system that actually owns them.

Collapse
 
c1-anderson profile image
Christian Anderson •

Agreed, and the awkward bit here is that nothing actually owns this number. There's no source of truth for what a used speaker is worth, just asking prices and whatever sold recently. So the best I could do was make every price say where it came from, and have the publish step refuse to list one the tool couldn't back up.

Collapse
 
eternaclarity profile image
Jesse Gamble •

Every price saying where it came from is the real fix here. Most pricing bugs are provenance bugs underneath.

Collapse
 
eternaclarity profile image
Jesse Gamble •

This is a good example of why AI outputs need to stay proposals until the business rules have had their turn. The model can be perfectly reasonable about bad evidence. A ceiling against the current new price, provenance on every comparable, and rules that keep accessory prices from becoming parent-item evidence would catch a lot of this before it reaches the listing.

Collapse
 
c1-anderson profile image
Christian Anderson •

Funny thing is, the ceiling against the new price is exactly what did the damage here. I added it to catch a used item priced above new, and then the "new price" it trusted turned out to be a £14.99 charger. So I'd put provenance at the top of that list. A ceiling's only as good as the number you anchor it to.

Collapse
 
eternaclarity profile image
Jesse Gamble •

Yeah. The guardrail trusted a bad input, so the guardrail was the bug. Provenance first, ceilings second.

Collapse
 
raknaos profile image
Raknaos •

The failure mode you describe is the one that gets me every time in agent pipelines too: the degraded signal is logged, just not anywhere a human reads before acting on the output. "retail-anchor discovery unavailable: no results" is a perfectly good log line and a terrible production gate.

So the design question I'd take from this is where the "unconfirmed price" boundary should sit. You chose to make the missing anchor visible in the price itself rather than block the listing, which is the right call for a tool you run yourself, but the moment anyone else uses it the £35.39 becomes the machine's opinion and the log line stops existing. Did you consider a hard stop when the anchor's only source was rate-limited, or is the manual new on amazon £26.99 override deliberately cheaper to build than the guard?

Collapse
 
c1-anderson profile image
Christian Anderson • • Edited

Direct answers, because a first version of this reply blurred them.

Where the boundary sits: at the action, not the number. A price the desk cannot corroborate is marked "thin", and the publish step refuses to list the desk's own thin number. That refusal lives in the code path that spends money, not in a log, so it survives a second user: the review page shows £35.39 with its warning, the machine cannot list it. Friday's £35.39 was thin from the moment it was priced off live asks.

The hard stop on a rate-limited source: considered and rejected, deliberately. The source was down for every run that day, so a stop keyed to it would have halted the desk, not the one price it could not justify. The typed "new £26.99" was cheaper than a guard on the source and I'd defend that, because the guard that matters already existed at publish.

What your question exposed, and what changed since I first answered: the block ignored whose number the price was. A person who read the warning and typed their own price on the review page was refused too, so the only override was the marketplace's own UI, outside every audit trail the desk keeps. Now the engine's thin number never goes live, a person's typed price goes through with their name on the approval, and a strict switch withdraws that override for installs where the person approving is not the person whose money it is.

Collapse
 
jo-do profile image
Jo Do •

"Every stage did what it was written to do and the result was still wrong" is the failure mode that makes agent pipelines humbling. The evidence hierarchy is the right instinct, and the retail anchor is doing the real work: asking prices are fiction nobody has paid for, so an anchor independent of every resale signal is the only thing that can catch a drift the whole market agrees on. A used item priced above new is such a clean canary for "your evidence is circular." Curious how the charger entered the story - cross-category evidence leaking in?

Collapse
 
c1-anderson profile image
Christian Anderson • • Edited

Not cross-category leakage from the resale side. It came in through the fix. The Pi case and the charger are two different items. The case was caught by the new ceiling once I typed Amazon's £26.99. The charger turned up later the same day, when I added an Amazon search so the anchor didn't depend on me typing a price. For a discontinued smart speaker, Amazon returns only accessories. The model gate that asks "is this the same product?" dropped all ten on one run. Minutes later it kept two 15W chargers "fit for Nest Mini". The anchor read £14.99 as the new speaker's price, and the ceiling I'd just added cut a £23.71 ask to £12.74. So the "used above new" check you're praising was the part that did the damage: a model's vote got wired straight into a hard cap. Two deterministic rules fixed it. A title that names the product only after "for" is an accessory. And a discovered new price can't cap anything when used asks are more than double it.

Collapse
 
doushabao profile image
Doushabao •

Nice approach! I've found that keeping API wrappers simple and well-tested is key. What testing framework do you prefer for Python APIs?

Collapse
 
c1-anderson profile image
Christian Anderson • • Edited

Honestly nothing fancy, just plain unittest from the standard library. No pytest. It's past 700 tests now and I haven't once wished for pytest's extras. The tests that actually earned their keep weren't the wrapper ones though, they were the ones pinning dumb little things, like the price regex grabbing "new" instead of "rrp". What are you using?