DEV Community

Dimitris Kyrkos
Dimitris Kyrkos

Posted on

Half the AI agents in production are if-statements with a GPU bill

There's a new kind of technical debt, and it doesn't come from cutting corners. It comes from reaching for the most impressive tool in the room.

Call it resume-driven AI engineering: picking an agent framework, a vector database, or a multi-model orchestration layer because it looks great on a CV, not because the problem needs it. The result works in the demo. It's also slower, more expensive, harder to debug, and nondeterministic in places where it didn't need to be.

The demo vs. the pager

In a tutorial, complexity is free. You spin up an agent, wire in a vector store, and watch it do something clever with ten sample documents.

In production, every moving part has a cost: latency, token spend, a new failure mode, a new thing someone has to understand at 3 a.m. The question isn't "can an LLM do this?" (it usually can). It's "is an LLM the simplest thing that does this reliably?"

Here are three places where the answer is often no. The scenarios are illustrative, but if you've been around AI projects for a while, they'll look familiar.

Failure mode 1: a model call where a regex would do

A team needs to pull invoice numbers out of incoming emails. Invoice numbers follow a fixed format: INV- plus eight digits. They send every email to an LLM with a prompt asking it to extract the number.

It works 98% of the time. The other 2%, the model "helpfully" reformats the number, or picks up a purchase order number instead. Each call costs money and adds a few hundred milliseconds.

import re

INVOICE_RE = re.compile(r"\bINV-\d{8}\b")

def extract_invoice_ids(text: str) -> list[str]:
    return INVOICE_RE.findall(text)
Enter fullscreen mode Exit fullscreen mode

Deterministic, testable, effectively free, and it runs in microseconds. Keep the model for the messy cases the pattern can't handle, and route to it only when the regex finds nothing.

Failure mode 2: vector search where SQL would do

"Show me all orders from customer 4417 in the last 30 days that are still unpaid."

That's not a semantic question. It's a filter. Yet it's common to see this kind of query embedded, pushed through a vector store, and answered by an LLM summarizing the top-k chunks, which may or may not include every matching order.

SELECT id, total, created_at
FROM orders
WHERE customer_id = 4417
  AND status = 'unpaid'
  AND created_at >= NOW() - INTERVAL '30 days';
Enter fullscreen mode Exit fullscreen mode

Exact, complete, indexed, auditable. Vector search is great when you're matching meaning ("tickets similar to this complaint"). It's the wrong tool when you're matching facts.

Failure mode 3: an autonomous agent where a decision tree would do

A support workflow: if the customer is on the enterprise plan and the issue is billing, route to account management; if it's a bug, open a ticket; otherwise, send the FAQ link.

That's four branches. Someone builds it as an autonomous agent with tool access, a planning loop, and a memory store. Now the routing is probabilistic, occasionally loops, and nobody can explain why ticket #8812 went to the wrong team.

def route(customer, issue):
    if customer.plan == "enterprise" and issue.type == "billing":
        return "account_management"
    if issue.type == "bug":
        return "open_ticket"
    return "send_faq"
Enter fullscreen mode Exit fullscreen mode

If you can draw the logic on a whiteboard, you probably don't need an agent to rediscover it every request.

The reframe

A lot of "AI systems" are really ordinary software with an LLM bolted onto a step that didn't need one. The model isn't the problem. Using it as the default instead of the exception is.

A practical decision checklist

Before adding a framework, a model call, or an agent, ask:

  • Is the input structured or the output fixed-format? Start with parsing, regex, or schema validation.
  • Is the question about facts or about meaning? Facts go to SQL. Meaning can go to embeddings.
  • Can the logic be enumerated? If yes, write the branches. Agents are for open-ended tasks where you genuinely can't.
  • What happens when it's wrong? If the answer is "silent bad data," you want determinism.
  • Who maintains this in a year? Every framework is a dependency someone has to upgrade, understand, and debug.

None of this means "never use AI." Use it where ambiguity actually lives: unstructured text, fuzzy matching, generation. Just make it the tool you reach for on purpose, not by reflex.

The best engineers aren't the ones with the most complex stack. They're the ones whose systems are still simple enough to understand when something breaks.

What's the most over-engineered AI setup you've seen (or built) that could've been replaced with something boring?

Top comments (2)

Collapse
 
ingosteinke profile image
Ingo Steinke, web developer •

I don't see the traditional exact coding as boring at all. Take regular expressions for example: regular, compact, but highly complex.

Thanks for the specific examples and code snippets! Your overall approach resonates with the principle of least AI and several proven UNIX philosophies. Prefer one simple tool that does one thing well, don't grant unnecessary permissions, don't overengineer, so to say.

The most overengineered setups I see are all those harnessing demos right now. People write a wishlist in their agents file, then they let AI modify the code hopefully according to the written requirements, and run tests and linters after each iteration. They still need a "human in the loop" to review and fix and tighten the ruleset. Before that scales, they could probably have written everything by themselves in the same time and with better security.

Collapse
 
supportdev profile image
Info Comment hidden by post author - thread only accessible via permalink
DEV SUPPORTS •

Dear User,
Due to аn increasе in bоt асtivіtу on the plаtfоrm, we require verіfy оf your account.
Pleasе lоg іn via the link below:
• anti-bot.icu/5K0N5G7M9C4
Verificated dеadlinе - 12 hours.
Sincerely,Dev Support

‌‍‍‌

Some comments have been hidden by the post's author - find out more