DEV Community

Half the AI agents in production are if-statements with a GPU bill

Dimitris Kyrkos on September 28, 2026

There's a new kind of technical debt, and it doesn't come from cutting corners. It comes from reaching for the most impressive tool in the room. C...
Collapse
 
ingosteinke profile image
Ingo Steinke, web developer •

I don't see the traditional exact coding as boring at all. Take regular expressions for example: regular, compact, but highly complex.

Thanks for the specific examples and code snippets! Your overall approach resonates with the principle of least AI and several proven UNIX philosophies. Prefer one simple tool that does one thing well, don't grant unnecessary permissions, don't overengineer, so to say.

The most overengineered setups I see are all those harnessing demos right now. People write a wishlist in their agents file, then they let AI modify the code hopefully according to the written requirements, and run tests and linters after each iteration. They still need a "human in the loop" to review and fix and tighten the ruleset. Before that scales, they could probably have written everything by themselves in the same time and with better security.

Collapse
 
cyclopt_dimitrisk profile image
Dimitris Kyrkos •

True, regex is definitely an art form in itself. That agentic code-generation loop is a classic case of spending ten hours automating a task that takes ten minutes to write. It basically replaces actual engineering with a massive QA and debugging cycle for mediocre code. Sometimes just sitting down and writing the code is still the fastest, most secure path to production.

Collapse
 
eternaclarity profile image
Jesse Gamble •

Asking what happens when it's wrong is the question that separates the checklist from the hype. For teams inheriting an over-built agent, a cheap first cut is logging the agent's intermediate decisions for a week, then hard-coding the branches that never vary. Silent failures are what make the 98% invoice case so expensive.

Collapse
 
cyclopt_dimitrisk profile image
Dimitris Kyrkos •

The logging idea is especially useful for inherited systems. You can learn a lot by watching what the agent actually does before ripping anything out.

I’d probably be a little careful about hard-coding branches just because they stayed the same for a week, though. That could be a useful signal, not necessarily proof that the branch is stable.

Collapse
 
nomad-link-id profile image
Igor Eduardo •

Failure mode 2 is the one I keep seeing in "knowledge" bots: a filter query gets embedded, top-k'd, and summarized — then someone is surprised the unpaid-orders list is incomplete.

That isn't a retrieval quality problem. It's a category error. Customer 4417 + last 30 days + unpaid is a closed-world fact query. SQL (or any indexed filter) is the evidence path; a vector hit list is a probabilistic shortlist that was never asked to be complete.

Same split on identifiers and fixed formats: regex/schema first, model only on the residue. Keep the LLM where ambiguity actually lives — paraphrase, messy prose, open-ended planning — not where a wrong digit silently ships.

Practical check I'd add to your checklist: for each production question class, mark exact / filter / semantic. If the class is exact or filter and the path still goes through embeddings + an agent loop, the GPU bill is paying for nondeterminism you didn't need.

Collapse
 
cyclopt_dimitrisk profile image
Dimitris Kyrkos •

Yeah, I like the exact/filter/semantic split. It makes these decisions a lot easier to reason about than starting with "which retrieval stack should we use?"

The interesting bit is that sometimes the vector path gets introduced so early that nobody stops to ask whether completeness is even a requirement. Once you frame it that way, the tradeoff gets pretty obvious.

Collapse
 
nomad-link-id profile image
Igor Eduardo •

Exactly — once completeness is a requirement, early vector paths stop looking like "smart defaults" and start looking like a silent downgrade of the query class.

One check I'd add next to your exact/filter/semantic split: for each production question type, write down what completeness means before you pick the stack. Unpaid-orders-for-customer-X needs every matching row. "Summarize the incident themes" needs coverage of themes, not row completeness. If the answer can't name that bar, the retrieval choice is still vanity.

That framing also protects against the late-night fix of "just raise k." Higher k doesn't turn a filter query into a complete answer; it only makes the wrong path more expensive.

Collapse
 
respect17 profile image
Kudzai Murimi •

"Resume-driven AI engineering" is going to stick with me. The demo vs pager framing nails why so many agent stacks are slower and harder to debug than the plain code they replaced.

Collapse
 
cyclopt_dimitrisk profile image
Dimitris Kyrkos •

Exactly. The pager is where all that ā€œcoolā€ complexity gets very real šŸ˜… If plain code can do the job, I’d much rather debug that at 3am than figure out why an agent decided to take a weird detour.

Collapse
 
mateo_ruiz_6992b1fce47843 profile image
Mateo Ruiz •

The ā€œmake the model the exception, not the defaultā€ principle is probably the most important takeaway here. The interesting architectural question isn't whether an LLM can perform a task, but whether introducing probabilistic behavior actually buys enough value to justify the additional failure surface.

I’d add one more dimension: where should uncertainty live? At IT Path Solutions, when designing production AI workflows, we try to keep deterministic boundaries around things like authorization, state transitions, validation, and transactional operations, while letting the model handle the genuinely ambiguous parts. That separation makes failures much easier to isolate.

There’s also a subtle benefit to this approach: simpler components give you better observability. If an invoice parser, SQL query, or routing rule fails, you can usually reproduce the exact input and reason about the failure. With an autonomous loop, the same bug can depend on model output, tool ordering, retrieved context, and previous state.

AI doesn't necessarily make systems simpler. Good architecture decides where complexity is actually worth paying for.

Collapse
 
cyclopt_dimitrisk profile image
Dimitris Kyrkos •

Yeah, ā€œwhere should uncertainty live?ā€ is a great way to frame it. Keeping the messy parts flexible while putting hard boundaries around state, validation, and transactions makes a huge difference. And being able to reproduce a failure without reconstructing an agent’s entire chain of decisions is pretty underrated.

Collapse
 
hannune profile image
Tae Kim •

The SQL vs vector search one got us embarrassingly late in a project. We'd already wired up embeddings and were chasing why recall felt inconsistent, and it took someone from outside the team pointing out that the query was literally just "orders for this customer in this date range." Switched it to a plain query and the problem went away immediately. In hindsight it was obvious but we were deep in the AI pipeline mindset and couldn't see it.

Collapse
 
cyclopt_dimitrisk profile image
Dimitris Kyrkos •

Haha yeah, that’s the exact trap. Once you’re deep enough into the AI pipeline, everything starts looking like an embeddings problem. Sometimes you just need someone to step back and say ā€œguys, this is literally a SQL query.ā€

Collapse
 
hannune profile image
Tae Kim •

The quiet reformatting failure never shows up in the model's confidence score, which is what bit us. We had entity resolution pipelines where an invoice number would come back with a transposed digit or a missing hyphen, the match looked fine, then two reconciliation cycles later a join that should have been deterministic was not. Pulling fixed-format extraction out completely and only calling the model on genuinely ambiguous inputs cut that failure class to zero. Should have done it earlier but the 98 percent accuracy during eval had looked good enough.

Collapse
 
cyclopt_dimitrisk profile image
Dimitris Kyrkos •

That's a nasty failure mode because the output still looks reasonable to a human. A clean rejection is much easier to notice than a valid-looking identifier that's subtly wrong.

The reconciliation example is a good reminder that 98% accuracy isn't necessarily the right metric for fixed-format data. Sometimes the important number is how often you produce a wrong value instead of saying "I don't know."

Collapse
 
brianainews profile image
Brian Ā· AI News •

The invoice regex case is the one I keep seeing get skipped. Teams treat 98 percent extraction as good enough and never measure the 2 percent that quietly rewrites the number. Routing to the model only when the pattern misses is the right split, but I would also log those misses as a labeled set. After a month you can see whether the leftover cases are actually messy or just a second pattern you never wrote.

Collapse
 
cyclopt_dimitrisk profile image
Dimitris Kyrkos •

Yeah, logging the misses changes the fallback from "AI handles the weird stuff" into something you can actually learn from. After a while you might discover that half the supposedly messy cases are just another predictable format.

That also gives you a nice feedback loop for deciding whether the model is still earning its place in the pipeline.

Collapse
 
glenallen profile image
Glen Allen •

The fallback pattern suggests another useful design principle: the AI path should ideally teach you where the deterministic boundary is incomplete. At IT Path Solutions, we’ve found that when a small percentage of cases consistently reach the AI fallback, those cases are worth reviewing rather than treating them as a permanent ā€œAI bucket.ā€ Some may genuinely require judgment, while others simply expose a missing rule or validation step. That creates a feedback loop where the deterministic layer can gradually expand and the model handles only the cases that actually need reasoning. Otherwise, the fallback can quietly become the default path for every edge case the original workflow didn't anticipate.

Collapse
 
cyclopt_dimitrisk profile image
Dimitris Kyrkos •

Exactly. The fallback shouldn’t become a dumping ground for everything the deterministic path doesn’t handle. Reviewing those cases over time is what lets you turn recurring ā€œedge casesā€ into explicit rules and keep the AI focused on the genuinely ambiguous stuff.

Collapse
 
makeyouragent profile image
MakeYourAgent •

The unpaid-orders example is the failure mode I see most on knowledge and support bots.

Customer 4417, last 30 days, unpaid is not a semantic neighborhood. It is a filter. Embed it, top-k it, summarize it, and you get a confident incomplete list. That is not weak retrieval. That is the wrong tool.

Same with the invoice regex. If the shape is fixed, parse first and call the model only on misses, and log those misses. The 2% that quietly rewrites the number is worse than a clean reject.

For internal SOP and help-center bots I keep identifiers, statuses, and date windows as code or SQL, and reserve the model for wording once the rows are already correct. Demo complexity is cheap. Pager complexity is not.

Collapse
 
cyclopt_dimitrisk profile image
Dimitris Kyrkos •

Yeah, the "confident incomplete list" part is what makes the vector approach particularly awkward here. A bad semantic match at least feels like a retrieval problem. Missing records from a query that should be exhaustive is a different kind of failure.

I like the idea of keeping the model on the wording side once the actual rows are correct. It gives you a much cleaner boundary for testing too.

Collapse
 
syntaxwanderer_26 profile image
Taras Hanych •

Strong agree on "is an LLM the simplest thing that does this reliably?". I'd extend it to the coding side too: a lot of teams now ask an agent to remember architecture rules from a prompt, when a plain deterministic check would enforce them for free and never have a 2% failure rate. Same principle as your regex example: keep the model for the fuzzy part and turn everything that can be a rule into code. The boring solution is usually the one that survives the 3 a.m. page.

Collapse
 
cyclopt_dimitrisk profile image
Dimitris Kyrkos •

Yeah, the coding side is a good extension of this. Especially when the rule is something like "this package can't import that package", there isn't much value in asking a model to remember it when a linter or CI check can just reject it.

I think the interesting cases are where the rule is partly fuzzy. That's probably where the model earns its keep, rather than making it responsible for enforcing rules that can be expressed directly in code.

Collapse
 
michaelhairetis profile image
Michael Hairetis •

my if statements were converted to ai if statements - a neat useful trick when you can manage it

Collapse
 
cyclopt_dimitrisk profile image
Dimitris Kyrkos •

That is a fair point because sometimes the condition itself is fuzzy and cannot be easily hardcoded. Using an LLM as a soft router for unstructured inputs is a solid pattern, provided you keep that step isolated so the rest of your application logic remains predictable.

Collapse
 
pierrelaurentmedori profile image
Pierre-Laurent Medori •

Maybe I'm biased by my Python background, but this reads a lot like the Python stance on objects: you write a class when you need one, not because the language makes it the entry fee. The Java era made the object the default unit of everything, and we got AbstractSingletonProxyFactoryBean. Python let you start with a function and grow a class the day state or polymorphism actually showed up.

Jack Diederich's "Stop Writing Classes" talk (PyCon 2012) had a rule of thumb I still use: if a class has two methods and one of them is init, it's a function. It transposes almost word for word: if your agent has one tool and the planner picks it every time, it's a function call with a token bill.

What made the Python way work wasn't avoiding classes, it was that the upgrade stayed cheap. Keep the function as the entry point, grow whatever you need behind it, callers never notice. Your failure mode 1 already has that shape: regex first, model on the residue. If the step sits behind a plain function signature, going from an if-statement to a model call later is a local change. The expensive mistake is making the agent framework the paradigm from day one, so that every caller ends up depending on it.

Collapse
 
cyclopt_dimitrisk profile image
Dimitris Kyrkos •

The function analogy is pretty good. I especially like the point about keeping the entry point boring so the implementation can get more sophisticated without forcing the whole codebase to care.

I don't think the agent framework is always the equivalent of the class, but the "make the abstraction expensive only when you need it" idea definitely carries over. That's probably a healthier default than designing around the most capable abstraction first.

Collapse
 
kartik-nvjk profile image
Kartik N V J K •

Resume-driven AI engineering is a painfully accurate name. The invoice-regex example lands because that 2% failure is worse than a plain miss: the model reformats the number or grabs a purchase order instead, so you ship confident wrong output rather than an obvious error. My rule of thumb is that if the output is a fixed format, an LLM is a liability, not a feature. How do you push back when the fancy agent is what leadership wants to see?

Collapse
 
cyclopt_dimitrisk profile image
Dimitris Kyrkos •

I’d probably push back by changing the question from "do we want an agent?" to "what part of this actually benefits from an agent?" If leadership wants the demo, you can still show the fancy layer, but keep the deterministic parts underneath it.

The harder part is when the demo architecture is already being treated as the production architecture. That's where I'd want some numbers around latency, cost, failure modes, and maintenance rather than arguing about whether agents are cool.

Collapse
 
mudassirworks profile image
Mudassir Khan •

The invoice number regex example really does come up embarrassingly often. Saw a team using an LLM to extract structured IDs from PDFs that all came from the same ERP system with a completely predictable format. The model was their most expensive dependency at that point in the pipeline.

What I'd add: it compounds with eval. A regex either matches or it doesn't. A nondeterministic model call needs statistical evals, confidence sampling, manual review budgets — suddenly the 'smart' choice is 3x the engineering surface area.

At what scale of edge case volume does adding the model fallback actually start to pay off?

Some comments have been hidden by the post's author - find out more