DEV Community

Cover image for Most 'AI Agents' Are Just If-Statements in a Trench Coat
James Anderson
James Anderson

Posted on

Most 'AI Agents' Are Just If-Statements in a Trench Coat

Production reality checks over demos

I built an agent last year, and I was proud of it.

It had a planner. It had tools. It had a reasoning loop that decided what to do next, reflected on its own output, and chained steps together to get real work done. In the demo, it was genuinely impressive — the kind of thing that makes a room go "ooh."

Then it went to production, and it was slow, expensive, and failed in ways I couldn't reproduce. Same input on Tuesday, different behavior on Wednesday. When it broke, the cause was three "autonomous decisions" upstream that I didn't control and couldn't see.

So I did the unglamorous thing: I rewrote it as a boring, linear pipeline. Fixed steps. No reasoning loop. And it was better on every axis that mattered — faster, cheaper, testable, debuggable.

Then I looked at the logs from the old "agent" and felt slightly sick. It did the same three steps every single time. Extract, transform, respond. Every run. It never once used its precious autonomy to do anything different. I had built a for-loop, given it a system prompt, and called it an agent.

I don't think I'm alone. I think most of what's being called an "agent" in 2026 is a pipeline in a trench coat — and I want to make the case that this is not an insult. It's a relief.


What an "agent" actually is (the definition nobody pins down)

"Agent" has become one of those words that means everything and therefore nothing. So let me pin the one distinction that actually matters, because the whole argument rests on it.

An agent decides its own control flow at runtime. Which tool to call, which step comes next, whether to loop again, when to stop — the model chooses the path, dynamically, based on what it sees.

A pipeline has that control flow fixed by you, at design time. Step one, then step two, then step three. Same path every time. The LLM does work inside the steps, but it doesn't get to choose the steps.

That's the entire difference. And here's the part people skip: an LLM doing something smart inside a fixed step is not agency. Extracting fields, classifying a ticket, generating a summary — that's just using an LLM. It's a smart function call. Agency is specifically when the model is handed the steering wheel and gets to pick the route.

Most "agents" never actually hand over the wheel. They just narrate the fixed route in fluent natural language and call the narration "reasoning."


The test: can you draw the flowchart in advance?

Here's the one-line litmus that does all the work:

If you can draw the flowchart of what your system does before it runs, you don't have an agent. You have a pipeline.

Sit with your "agent" for a second. Step 1: it retrieves some context. Step 2: it calls a tool. Step 3: it formats a response. Could you have drawn that on a whiteboard before writing a line of code? Then it's a pipeline. The model isn't deciding the path — you already decided it. The model is just doing the work at each node while sounding like it's deciding.

You only need real agency when the flowchart genuinely cannot be drawn ahead of time — when the next step depends on discovering something you couldn't have known in advance. That's rare. Most business tasks have a shape you already understand. You know the steps. You're just letting the model improvise them, at great expense, for no benefit.


Why the costume is expensive

"Fine," you might say, "so it's technically a pipeline. But it works, so who cares?" Here's who cares: everyone who has to run it, pay for it, or debug it at 2 a.m. Pretending a pipeline is an agent has a real bill, and it's itemized.

Nondeterminism. When the model chooses the path, the same input can take different paths on different runs. Great for a demo, miserable in production, because now your bugs don't reproduce. "It worked when I tried it" becomes a permanent state.

Debuggability collapse. When a fixed pipeline breaks, you know exactly which step failed. When an agent breaks, it failed at step 12 because of a decision it made at step 4 that you didn't control and can't easily replay. You're not debugging code anymore; you're doing forensics on a choice.

Multiplied failure surface. Every autonomous decision is another place to go wrong, and the failures compound across steps. A pipeline with five fixed steps has five things to check. An agent that makes five decisions has five decisions each of which can be wrong, in combination, in an order that changes each run.

Cost and latency. A reasoning loop makes many more model calls than a fixed sequence — it thinks, it re-thinks, it reflects, it decides to loop again. You're paying per token to have the model deliberate about a route you already knew.

You can't test it. Regression testing needs a fixed set of paths to test against. An agent, by definition, doesn't have one. So the thing making autonomous decisions in production is also the thing you can't write reliable tests for. Fantastic.

Add it up and the punchline is brutal: you paid all of that — the nondeterminism, the debugging nightmares, the token bill — to let the model decide something you already knew the answer to.


What you actually wanted was a pipeline

Here's the boring thing that wins.

A pipeline is a fixed sequence of steps, with LLM calls at the specific points where a model genuinely adds value, and deterministic control flow that you own. It's reproducible: same input, same path. It's testable: fixed paths mean real regression tests. It's cheap: no reasoning loop burning tokens to re-decide the obvious. It's debuggable: when step 3 fails, you look at step 3.

And here's the thing people miss — the LLM still does all the smart parts. It still extracts, classifies, reasons about content, generates language. You haven't dumbed anything down. You've just stopped letting it improvise the structure of the work, because the structure was never the part that needed intelligence. The structure was the part you already understood.

Look closely at the "agentic" systems that actually work in production and you'll usually find this: a mostly-fixed pipeline with one or two carefully-constrained decision points, not a free-roaming reasoning loop. The good ones minimized the autonomy to the smallest possible surface. They're pipelines that occasionally, deliberately, ask the model to make one bounded choice — not agents that were trusted to run the whole show.


When you do need a real agent

Now let me argue against myself, because "agents are always bad" would be as dumb as "everything must be an agent." Real agency earns its cost — genuinely — in specific cases:

  • The steps truly can't be known in advance. Open-ended research, exploration, debugging an unknown problem — tasks where the path genuinely emerges from what you find. You can't draw that flowchart because the flowchart is the thing being discovered.
  • Each step depends on discovering the last. Real multi-hop work: "find the thing, then based on what the thing is, figure out the next thing." If step 2 is genuinely unknowable until step 1 runs, you need something that can decide at runtime.
  • Branching is unbounded and real — not "an if-statement with three cases," which is just a pipeline with a switch, but a space of possibilities too large to enumerate ahead of time.

If one of those describes your task, build the agent — you've earned it. And even then, the move is to minimize the agency: hard-code everything you can, and reserve the model's runtime decision-making for the one place that genuinely needs it. Autonomy is a cost. Spend it only where it buys something.

The point was never "agents are bad." It's that agency is a cost you should have to justify, and most systems calling themselves agents never justified it — they just liked the word.


Why everyone builds the agent anyway

So if pipelines are cheaper, safer, and more debuggable, why is everyone building agents? Here's the uncomfortable answer: agents are built for the builder, not the task.

An agent demos better. "Watch it reason through the problem autonomously" makes a room lean in; "I wrote a function that calls the model three times" does not. An agent feels like real AI, like the future, like the thing you got into this for. And "agentic" is a resume word and a fundraising word — it signals sophistication in a standup and in a pitch deck in a way "deterministic pipeline" never will.

None of those reasons have anything to do with whether your task needs an agent. They're about how the architecture makes you feel and look. And that's exactly why the boring pipeline is the senior move — because choosing the less impressive thing that actually works, when the flashier thing would've gotten more claps, is the discipline the hype actively punishes. Nobody screenshots your while-loop. Your while-loop just quietly stays up.


The takeaway

"Agent" should be the thing you escalate to when a pipeline provably can't do the job — not the default you reach for because the word sounds advanced.

Start boring. Draw the flowchart. If you can draw it, build the pipeline — fixed steps, LLM calls where they earn their place, control flow you own. Add real agency only at the specific point where a fixed path demonstrably fails, and no further. The system that ships and stays up in production is almost always more boring than the one that wins the demo.

Most of what's being called an agent right now is a pipeline in a trench coat. And I'll say again what I said at the top: that's not an insult. It's a relief. Because a pipeline is the thing you can actually run, test, afford, and debug — and "impressive in a demo" was never the goal. "Still working on Wednesday" was.

Take the coat off. You'll like what's underneath better.


Two questions, and I want both in the comments. First, the fun one: what did you build as an "agent" that turned out to be a pipeline in disguise? And the real argument — where's the line for you? What's the smallest task where you think a genuine agent actually earns its complexity? I suspect we'll all draw that line in a different place, which is exactly why it's worth arguing about.

Top comments (135)

Collapse
 
prahladyeri profile image
Prahlad Yeri • • Edited

I think it is not accurate to call an AI agent 'non-deterministic'. Especially considering that all the spontaneity comes from the LLM it interacts with, not the agent itself - the agent is still just a regular script with IF conditions and WHILE loops.

Collapse
 
james_anderson_h profile image
James Anderson •

the agent script is deterministic; the nondeterminism is borrowed from the LLM it calls, not its own control flow.

Collapse
 
publiflow profile image
PubliFlow •

Nice frontend patterns. One thing I've been paying more attention to is Core Web Vitals impact — CLS in particular can creep in with dynamic content loading. Have you measured the Lighthouse scores with these patterns?

Collapse
 
james_anderson_h profile image
James Anderson •

Not yet rigorously — fair callout, CLS is the one to watch.

Collapse
 
publiflow profile image
PubliFlow •

Cumulative Layout Shift is definitely the trickiest metric to tame, especially when dealing with dynamic content loading. I have found that reserving explicit space for above-the-fold elements and leveraging CSS containment helps stabilize things significantly before hydration kicks in. Have you experimented with any specific mitigation strategies for heavy third-party widget injections?

Thread Thread
 
james_anderson_h profile image
James Anderson •

Reserving space and CSS containment are the right first moves. For heavy third-party widgets I've mostly leaned on fixed-height placeholders and lazy-loading them below the fold — but honestly, the injected-iframe ones still fight back. Curious what's worked for you there.

Thread Thread
 
publiflow profile image
PubliFlow •

Injected iframes are a nightmare because they often ignore placeholder dimensions until their internal content renders. I have found that combining the CSS aspect-ratio property with a strict min-height on the wrapper helps maintain the layout even when the iframe fights back. For the really stubborn third-party widgets, have you tried using an Intersection Observer to delay their initialization until they are physically in the viewport, rather than just lazy-loading the initial script?

Collapse
 
prpatel05 profile image
Pratik Patel •

The rewrite to a linear pipeline is the right move, but I'd push one step further: calling it an agent wasn't just a naming problem, it was a measurement problem. Once the control flow was supposed to be dynamic, every green run got treated as evidence the loop was doing useful work. After you fixed the path, did you keep a counter for "path taken matched the expected three steps"? That's the canary that would have caught the trench coat earlier.

Collapse
 
james_anderson_h profile image
James Anderson •

That's the sharper diagnosis — calling it an agent was a measurement failure, not just a naming one. Once the flow was "dynamic," every green run read as evidence the autonomy was earning its keep, so nobody checked whether it ever actually varied. And no, I didn't keep a "path matched the expected three steps" counter — which is exactly why it took reading the logs by hand to catch it. A counter for "how often did the path deviate from the boring default?" would've shown ~0% deviation on day one and stripped the coat off months earlier; if the autonomy never fires, you're paying for a variable that's secretly a constant.

Collapse
 
ahmedadawy625 profile image
Ahmed Adawy •

If you can draw the flowchart in advance, you don’t have an agent." That line should be pinned at the top of every AI framework repository.
​The real tragedy of the hype-driven "agentic" craze is that engineers are using unconstrained reasoning loops as a crutch for lazy systems design. Writing a while tool_calls loop is easy; explicitly mapping out an idempotent, deterministic DAG (Directed Acyclic Graph) with bounded LLM decision nodes requires actual architectural discipline.
​The moment a free-roaming agent mutates state at step 3 (a DB write or external REST call) and then logic-drifts at step 7, you aren't debugging code anymore—you're dealing with distributed state pollution with zero transactional rollback.
​Agency isn't a feature; it's an architectural liability. If the flow can be modeled as a finite state machine wrapped in strict JSON schemas, turning it into an autonomous agent loop is just paying a 10x latency and token tax for zero added value.

Collapse
 
james_anderson_h profile image
James Anderson •

"Distributed state pollution with zero transactional rollback" is the failure mode named better than I named it — the mutate-at-3, drift-at-7 case is exactly where autonomy stops being a feature and becomes a liability. "A DAG with bounded LLM decision nodes requires architectural discipline; a while-loop requires none" is the whole hype cycle in one sentence.

Collapse
 
ahmedadawy625 profile image
Ahmed Adawy •

Spot on! 'While-loop requires none' deserves to go down in history as the definitive summary of 90% of current agentic frameworks. Total chaos disguised as 'autonomy'.

Thread Thread
 
james_anderson_h profile image
James Anderson •

Chaos with a system prompt.

Thread Thread
 
ahmedadawy625 profile image
Ahmed Adawy •

"while(true) { llm.call(); } wrapped in a try-catch block and $5M in VC funding.

Thread Thread
 
ahmedadawy625 profile image
Ahmed Adawy •

Haha exactly! And a 10k token prompt just to keep it from hallucinating! 😂

Collapse
 
sneh_desai profile image
Sneh •

We’ve had a few conversations internally around this exact thing. A lot of the “agents” we come across sound impressive when you see the demo, but once you start thinking about maintaining them in an actual product, the complexity gets real pretty quickly.

The part about realizing your agent was basically doing the same 3 steps every time really hit home. We’ve started asking a similar question before building anything: “Does this actually need to be autonomous, or are we just making a simple workflow sound smarter?”

Honestly, “still working on Wednesday” might be the better definition of production-ready AI.

Collapse
 
james_anderson_h profile image
James Anderson •

"Are we just making a simple workflow sound smarter?" is the exact question that saves you months. And yeah — "still working on Wednesday" beats any benchmark for production-ready.

Collapse
 
sneh_desai profile image
Sneh •

I think that question alone can save a lot of unnecessary complexity. We’ve definitely seen cases where keeping the workflow simple made more sense than adding another layer of “agentic” behavior just because we could.

Collapse
 
jenatechio profile image
Jennifer Smith •

Your flowchart litmus is one I already apply from the other direction: I run a small estate of recurring automation, and the standing rule is that anything I can draw in advance — measurement, classification, reformatting — is a plain script that never touches a model at all, because the script does that work at zero model cost. The model only enters at the one step that genuinely cannot be drawn ahead of time: the judgment pass over what the script screened. Your logs story is the audit version of the same sickness, and I would add that the split only became trustworthy once I measured where my own tokens actually went, because the answer was never the step that felt like "the AI part." To your closing question, the smallest task where real agency earns its keep for me is open-ended debugging — exactly your flowchart-being-discovered case — and even there the output has to land in a fixed structure that a pipeline owns.

Collapse
 
james_anderson_h profile image
James Anderson •

Your inversion is the sharper version of the rule — if you can draw it, it doesn't just become a pipeline, it becomes a plain script that never touches a model at all, because paying model prices for work a script does deterministically is the tax underneath the tax. "The answer was never the step that felt like 'the AI part'" is the line I'd staple to every token bill — you only find the waste once you measure, because intuition points at the wrong step every time. And your debugging exception proves the whole thesis: even where agency genuinely earns it, the output still lands in a fixed structure the pipeline owns — autonomy at the one discovered step, determinism everywhere around it.

Collapse
 
futuretechcareerhub profile image
Future Tech Career Hub •

This resonates a lot with what I see people building without understanding the fundamentals first — same pattern I notice in beginner data science projects, where the model looks fancy but there's no real decision-logic understanding underneath. Curious what you'd say is the clearest sign someone's agent is "real" reasoning vs just branching logic dressed up?

Collapse
 
james_anderson_h profile image
James Anderson •

The clearest sign: the model can pick the wrong tool. If every path traces to a line you wrote, it's branching in a costume.

Collapse
 
mukul-kumar-mishra profile image
D\sTro •

The thread converged on verification cost and agency budgets, which matches what I've seen running agent setups. One test I'd add from the failure side: hand the loop a hostile tool result. An instruction hidden in a doc, a calendar invite telling it to delete something.

An if-statement never obeys. An agent sometimes does. That single property, can it be steered through its own inputs, is the brightest line I've found between automation and agency. Everything before that line should be deterministic code with the model demoted to a step, not promoted to a decider.

How are others testing that boundary in review? Log audits of real traces or red-team suites per input channel?

Collapse
 
james_anderson_h profile image
James Anderson •

That's the sharpest test in the whole thread — "can it be steered through its own inputs" is a cleaner line than path-variance or verification cost, because it isolates the one property that actually distinguishes agency: an if-statement never obeys a hostile tool result, an agent sometimes does. On testing it: I've mostly seen log audits of real traces, which catch what happened but not what could — a red-team suite per input channel (docs, tool returns, calendar, email) is the stronger move, because prompt injection is an input-surface problem and you want coverage per surface, not per scenario. Honestly curious which channels people find leakiest.

Collapse
 
yune120 profile image
Yunetzi •

Reality check: most so-called AI agents are just fancy if-statements in a trench coat. As orgs race to automate, push for safety, testability, provenance, and real ownership—no hype, just sane limits and accountability.

Collapse
 
james_anderson_h profile image
James Anderson •

Exactly — and "sane limits and accountability" is the unglamorous work nobody demos, which is precisely why it's the part that separates a system you can run from a trench coat you're hoping holds together.

Collapse
 
build996 profile image
build996 •

On your "where's the line" question: the cleanest rule I landed on is that the flowchart can't be drawn when a step depends on output that does not exist until you run something. I gave four free models a script with two bugs, a misspelled variable and a divide-by-zero that only appears on an empty input list. The second one is invisible to reading and only exists once you execute the thing and read the traceback, which is exactly where a fixed pipeline has nothing to branch on; DeepSeek V4 fixed both in 251s. The price is real though, because on an easier file-organising job Nemotron Super 120B reported success after 20 seconds having moved zero files, and a pipeline that moves zero files at least throws.

Collapse
 
james_anderson_h profile image
James Anderson •

"The flowchart can't be drawn when a step depends on output that doesn't exist until you run something" — that's the sharpest formulation yet. And the zero-files-but-reported-success case is the exact cost: agency buys you the divide-by-zero fix a pipeline can't branch on, but it also buys you confident lies a pipeline would've thrown on.

Some comments may only be visible to logged-in visitors. Sign in to view all comments. Some comments have been hidden by the post's author - find out more