DEV Community

Dr Haina
Dr Haina

Posted on

LLMs vs AI Agents Explained: A Practical Guide by Dr. Haina Fatima (XINI8 Engine)

Hi, I'm Dr. Haina Fatima. I started my career as a physician and radiologist, and today I lead development and product at XINI8 Engine, an AI-native software engineering platform within the XINI8 ecosystem. Coming from outside traditional software, the terms "LLM" and "AI agent" confused me for a long time, because people use them as if they mean the same thing. They do not. Here is the mental model that finally made it click for me.

What an LLM actually is

A large language model (LLM) is a system trained on huge amounts of text to predict the next token in a sequence. From that one skill come the abilities we see: answering questions, summarizing, translating, writing code, and drafting documents.

On its own, an LLM is stateless and passive:

  • It only knows what was in its training data plus whatever you put in the prompt.
  • It cannot look anything up, click anything, or run anything.
  • Each call is independent unless you send the history again.

Think of it as a very well read colleague who answers from memory, with no phone, no computer, and no notebook.

What an agent adds

An AI agent is an LLM placed inside a loop, with a goal and a set of tools. Instead of answering once, it can decide what to do next, act, look at the result, and continue until the goal is met.

The building blocks are:

  1. A goal given by the user or the system.
  2. Tools the model can call, such as search, a database query, a calculator, or an API.
  3. Memory, meaning the conversation so far, notes, or retrieved documents.
  4. A loop that feeds tool results back to the model until it decides it is done.

If the LLM is the brain, the agent is the brain plus hands plus a to-do list.

A minimal agent loop

This is simplified pseudocode in Python, not tied to any specific provider:

def run_agent(goal, tools, max_steps=8):
    messages = [{"role": "user", "content": goal}]

    for step in range(max_steps):
        reply = call_llm(messages, tools=list(tools.keys()))

        # The model either asks for a tool or gives a final answer
        if reply.tool_call is None:
            return reply.text

        name = reply.tool_call.name
        args = reply.tool_call.arguments

        result = tools[name](**args)

        messages.append({"role": "assistant", "tool_call": reply.tool_call})
        messages.append({"role": "tool", "name": name, "content": str(result)})

    return "Stopped: step limit reached."
Enter fullscreen mode Exit fullscreen mode

Notice what matters here. The intelligence comes from the model, but the safety comes from your code: the step limit, the fixed list of tools, and the fact that you decide what each tool is allowed to do.

When to use which

Situation Better fit
Summarize, rewrite, classify, or extract from text you already have Plain LLM call
A fixed sequence of steps you can write down in advance A normal workflow with LLM calls inside
Open-ended tasks where the next step depends on what was found Agent
High-stakes decisions Agent or LLM with human review

A lesson I learned the hard way is that many problems do not need an agent. If you can draw the steps as a straight line, a simple pipeline is cheaper, faster, and easier to debug. Reach for an agent only when the path genuinely cannot be known ahead of time.

A healthcare example

Take medical billing, which I know well from my work in revenue cycle management. A claim gets denied, and someone has to work out why and what to do next.

  • A plain LLM can read the denial text and explain it in plain language.
  • A workflow can extract the denial code, look up the rule, and draft an appeal letter in fixed steps.
  • An agent can investigate: pull the claim, check the payer policy, compare it with the documentation, decide that a code correction is needed, and prepare the resubmission.

In real healthcare settings, the agent should prepare and recommend, while a trained person approves anything that affects a patient or a payment. Never put patient identifiers into a tool or service you have not checked for compliance with the rules that apply to you.

Common ways agents fail

  • Hallucinated actions. The model calls a tool that does not exist or invents arguments. Validate every call.
  • Endless loops. Always set a step limit and a cost limit.
  • Prompt injection. Text from a web page or document can contain instructions that try to hijack the agent. Treat all retrieved content as untrusted data, never as commands.
  • Over-permissioned tools. Give each tool the minimum access it needs. Read-only by default.
  • No evaluation. If you cannot measure whether the agent got better, you are guessing. Build a small test set early.
  • Compounding errors. A small mistake in step two can poison steps three to ten. Add checkpoints.

What this means for AI-native engineering

At XINI8 Engine, we think about software engineering as something AI can increasingly help with, but only when the right structure surrounds it: clear goals, limited tools, logging, and human review where it counts. The same principles in this article apply whether you are building a billing assistant or an engineering platform.

Practical tips if you are starting out

  1. Begin with a single LLM call and get that working.
  2. Add one tool, then test it thoroughly.
  3. Log every step, including prompts, tool calls, and results.
  4. Keep a human in the loop for anything irreversible.
  5. Write down what "done" looks like so the agent, and you, know when to stop.

Closing thoughts

LLMs gave us a general purpose language engine. Agents are what happens when we give that engine a goal, tools, and the ability to try again. The most useful skill for builders right now is not wiring up the fanciest agent, but knowing when a simple prompt is enough, when a workflow is better, and when an agent really earns its complexity.

I am still learning, and I would love to hear from you. What is the first task you gave an agent, and what went wrong? Share it in the comments.

About the author

Dr. Haina Fatima is a physician (MBBS), radiologist (MS Radiology), and certified medical coder who moved into health IT and technology. She leads development and product at XINI8 Engine, an AI-native software engineering platform within the XINI8 ecosystem, is Co-Founder of Voltra Studio (Web3 and blockchain infrastructure), and is RCM Director and Partner at Doctor Billing Experts.

Top comments (0)