DEV Community

Cover image for Five agent episodes, and the build decision to take from each
Conor Bronsdon
Conor Bronsdon

Posted on Originally published at chainofthought.show

Five agent episodes, and the build decision to take from each

An agent loop can look fine in a notebook and fall apart across sessions, teams, and production data. The Chain of Thought collections group episodes for people shipping agents as working systems. Below is one build decision per episode from that agents set, tied only to what those guests argue on the show.

Why isn't my framework the moat anymore?

Jerry Liu helped turn LlamaIndex into one of the most installed AI frameworks of the last three years, then argued on The AI Framework Era Is Over: Why Context Is the Moat that the framework era is over. His point for builders: as agent loops and harnesses get good enough, the scaffolding stops being where you win. Context quality is what compounds.

Practical move: budget engineering time for the context layer (what gets retrieved, parsed, and kept accurate) before you spend it swapping orchestration libraries. If your use case is document heavy, Jerry treats very high parse accuracy as the bar for legal, insurance, and financial work (he cites 95%+ as the real target in those verticals). If the agent reads bad inputs, I would not expect prompt polish to rescue it.

What should I build so the agent remembers the next session?

Richmond Alake, Director of AI Developer Experience at Oracle, treats memory as the layer that decides whether an agent works across sessions, and says most teams under-build it. On Agent Memory: The Last Battleground in the AI Stack, he separates memory engineering from prompt engineering and context engineering as its own discipline.

Two failure modes he calls out often: the wrong mental model, and deleting data instead of forgetting it. Production memory should decay and deprioritize (he points to relevance, recency, and importance style scoring from the Generative Agents line of work) instead of hard-deleting them. He maps working, episodic, semantic, and procedural memory to how you segment what goes in the window versus what lives in stores. Files are fine for prototypes; he argues databases win in production.

When do many agents beat one, and what breaks?

Mikiko Chandrasekhar of MongoDB frames multi-agent systems as a reliability problem first. On Mastering Multi-Agent Systems, she says agents offer serious capability, but mastering reliability is what decides whether agentic systems work in production. Her recommendation is to treat agents as software products: observability, evaluations, and guardrails are part of the product for non-deterministic systems.

Add agents only when the task needs split roles or parallel work, because coordination adds failure modes. Plan for chat logs, vector search, and debugging trails before the second agent joins.

What layers am I stacking when I call something an agent?

João Moura, founder and CEO of CrewAI, uses The Emerging AI Agent Stack to map what enterprises assemble when they move past curiosity: orchestration, provisioning, authentication, and measurement for large numbers of agents in production. Knowledge-work automation at scale needs measurability, reliability, and trust in his framing; a demo that completes once is the start.

The episode talks about provisioning, authentication, and measurement for hundreds or thousands of agents. Before you add agents, list what must exist at that scale: identity, access, metrics, and interoperability between components. If you cannot name how you will measure success and catch regressions, you are still in experiment mode. João and his co-panelists also stress that non-deterministic multi-agent setups multiply failure modes, so testing has to go beyond what worked for traditional services.

How do I tell an agent from a fancy workflow?

The earliest anatomy episode in this set is still the clearest line in the sand. On Got Agents? Agentic Workflows & Architecture, João Moura draws the line at agency: a fixed if-this-then-that flow is not an agent. Brian Raymond and Bob van Luijt sit in the same conversation on data management, deployment, and generative feedback loops, but that definition is the decision rule.

Use this before you rename pipelines:

# Sketch: agency check before you call it an "agent"
def is_agentic(step_plan, can_replan: bool, tool_choice: str) -> bool:
    if step_plan == "fixed_dag" and not can_replan:
        return False  # workflow with an LLM step, not an agent
    if tool_choice == "model_picks_per_turn":
        return True
    return can_replan
Enter fullscreen mode Exit fullscreen mode

If every branch is predetermined, you likely need better workflow tooling and evals more than an autonomous label on the README.

Checklist

  • Invest in context and document accuracy where Jerry Liu says the moat moved; treat framework churn as secondary.
  • Name memory engineering explicitly: prefer controlled forgetting over deletes, and plan stores past file-based prototypes (Richmond Alake).
  • Add multiple agents only with product-grade observability and evals; treat reliability as the gate (Mikiko Chandrasekhar).
  • Stack provisioning, auth, and measurement before you scale agent count (João Moura on the agent stack).
  • Refuse the agent label when the system is a fixed if-this-then-that flow (João Moura, EP 2).

More on this topic, with the related episodes, is on Chain of Thought. It draws on this episode.

Subscribe to the Chain of Thought newsletter for new episodes and write-ups like this one.

Drafted with AI assistance from the episode transcripts.

Top comments (0)