DEV Community

Cover image for I stopped letting one agent do everything and my tool calls got way less cursed
Lars Winstand
Lars Winstand

Posted on Originally published at standardcompute.com

I stopped letting one agent do everything and my tool calls got way less cursed

I realized I was asking one poor model to do four jobs badly.

  1. Plan the task
  2. Pick the right tool
  3. Recover from failures
  4. Write a clean final answer like nothing exploded

That setup looks great in a demo.

In production, it’s how you get a Slack bot that says “done” after failing on step two.

I’ve seen the same failure pattern in n8n, LangChain, OpenAI function-calling loops, and custom Python agents:

  • one giant GPT node with access to everything
  • retries handled by vibes
  • broken API responses treated like “reasoning problems”
  • too much context dumped into one loop

The fix was not a better prompt.

The fix was splitting orchestration from generation.

Stop asking one model to be the planner, dispatcher, retry loop, validator, and summarizer.

Give it supervision.

The problem: one agent is doing workflow engine work

A lot of agent setups are really orchestration engines wearing chatbot costumes.

We tell the model:

  • think through the task
  • choose tools
  • inspect outputs
  • detect failures
  • retry if needed
  • decide when it’s done
  • summarize the result

That is not one job.

That is a control plane.

And once you treat it like one, a lot of weird agent behavior suddenly makes sense.

What actually makes tool-calling agents flaky?

Most tool-use failures are not because GPT-5, Claude Opus 4.6, or Grok 4.20 are “bad at reasoning.”

They’re usually workflow bugs.

Common ones:

  • Too many tools exposed to one agent
  • No explicit retry policy when an API fails
  • Retrieval is unbounded, so the model gets flooded with junk
  • State from one task leaks into another
  • The same model is doing reasoning, validation, exception handling, and cleanup

That last one is the killer.

When the agent breaks, people usually respond by making the prompt longer.

More instructions.
More examples.
More all-caps warnings.

Sometimes that helps.
Usually it just creates a more polite failure.

The better question is:

Which parts of this workflow should not be left to model improvisation?

That question leads straight to the supervisor pattern.

The pattern: supervisor + narrow workers

This is the version I keep coming back to:

  1. A supervisor plans the work
  2. Specialist workers handle narrow tasks
  3. The runtime handles retries, validation, tracing, and state transitions
  4. A final step summarizes the result

The supervisor is not there to be clever.

It’s there to keep the system honest.

It decides:

  • which worker gets the task
  • which tools are allowed
  • whether a failure should retry or stop
  • what counts as success

That one change removes a shocking amount of chaos.

A minimal example with the OpenAI Agents SDK

The OpenAI Agents SDK already nudges you in this direction.

Minimal shape:

from agents import Agent, Runner

planner = Agent(
    name="Planner",
    instructions="Break work into steps and delegate to the right worker."
)

result = await Runner.run(planner, "Find the latest API pricing changes and summarize them")
print(result.final_output)
Enter fullscreen mode Exit fullscreen mode

That snippet is tiny, but the important part is architectural:

  • Agent defines behavior
  • Runner owns the loop

That is already better than raw “model, keep calling tools until you feel done.”

If you add handoffs, validation, and tracing, you get a much safer setup without building everything from scratch.

Why narrow workers beat one giant generalist

Because tool choice gets worse as the tool list grows.

If you give one agent 12 tools, it will eventually get creative in a bad way.

A web research worker should do web research.
A file retrieval worker should do file retrieval.
A reporting worker should format output.

Not every worker needs access to every tool.

That’s how you reduce blast radius.

Example: bounded retrieval matters

One of the easiest reliability wins is limiting retrieval.

Don’t do this:

  • fetch everything
  • dump it all into context
  • hope the model sorts it out

Do this instead:

from agents import Agent, FileSearchTool, WebSearchTool

agent = Agent(
    name="ResearchWorker",
    tools=[
        WebSearchTool(
            search_content_types=["image", "text"],
            image_settings={"max_results": 3, "caption": True}
        ),
        FileSearchTool(
            max_num_results=3,
            vector_store_ids=["VECTOR_STORE_ID"],
        ),
    ],
)
Enter fullscreen mode Exit fullscreen mode

Even here, I’d usually split this into two workers.

  • one worker for web
  • one worker for internal files
  • supervisor decides which one to call

The point is not elegance.
The point is fewer weird failures.

What breaks real agents is boring operational stuff

Most agent demos don’t die on reasoning.

They die on:

  • HTTP timeouts
  • 429s from Notion
  • malformed Jira payloads
  • partial failures halfway through a long run
  • state that gets corrupted after a retry

This is why long-running workflow tooling matters.

If your automation touches Gmail, HubSpot, PostgreSQL, Slack, and an internal API, the main problem is often not “smarter prompting.”

It’s resumability.
Checkpointing.
Explicit failure handling.

That’s why tools like LangGraph are interesting.

Not because they make prompts prettier.
Because they treat agents like workflows.

Install is simple:

pip install -U langgraph
Enter fullscreen mode Exit fullscreen mode

The bigger shift is mental.

You have to admit your agent is not just a chat loop.

It is a workflow with failure modes.

State is not one thing

This part gets ignored a lot.

People say “context” like it’s one big blob.

It isn’t.

There’s session state, tool state, conversation history, sandbox state, intermediate outputs, retry metadata, and final outputs.

When teams blur all of that together, agents get haunted fast.

You start seeing things like:

  • today’s invoice parser using yesterday’s customer record
  • a failed code execution polluting the next task
  • retries reusing stale tool outputs

A supervisor layer helps because it can control state transitions explicitly.

Something like:

  • open a fresh job session
  • pass only required state to each worker
  • checkpoint outputs after each step
  • isolate execution environments
  • clean up resources at the end

That is much easier to debug than one giant conversational loop with mystery memory.

Which stack should you use?

My blunt version:

Option Best use case
OpenAI Responses API Short-lived workflows where you want direct control over tool dispatch, loop logic, and state
OpenAI Agents SDK Multi-agent workflows where you want handoffs, guardrails, sessions, tracing, and a managed runner
LangGraph Long-running, stateful automations where durable execution and resume-after-failure matter most

And the tradeoff view:

Option What you’re really signing up for
OpenAI Responses API More plumbing, more control, less hidden behavior
OpenAI Agents SDK Good default if you want supervisor-style structure without building every runtime piece yourself
LangGraph More workflow engineering up front, much better operational behavior for complex automations

If you’re already building in Python, the OpenAI Agents SDK is a pretty reasonable place to start.

If your automations run long and touch lots of systems, I’d look hard at LangGraph.

If the workflow is short and deterministic, the Responses API may be enough.

When this is overkill

Not every automation needs a mini control plane.

If your agent:

  • uses one or two deterministic tools
  • runs for a short time
  • has simple success/failure logic

then a single loop with strict schemas may be fine.

I would not build a tiny bureaucracy for a support-ticket draft assistant that reads one knowledge base and writes one reply.

The point is not “always build a supervisor.”

The point is:

stop pretending one model call should be the supervisor.

My rule now

If an agent:

  • has more than a couple tools
  • touches external APIs
  • runs long enough that restart is painful
  • needs retries that aren’t trivial

then I assume it needs supervision.

Not because the models are weak.

GPT-5, Claude Opus 4.6, and Grok 4.20 can all be excellent workers.

But production reliability usually comes from workflow design, not model heroics.

The smartest model in the world will still have a bad day if you make it do planning, routing, validation, retries, cleanup, and summarization inside one fuzzy loop.

Make the model good at the thing it should do.

Keep orchestration in the runtime.

That one change made my tool calls a lot less cursed.

Practical checklist

If you want to clean up a flaky tool-calling agent this week, I’d start here:

[ ] Split planner/supervisor from worker tasks
[ ] Reduce tool access per worker
[ ] Add explicit retry rules outside the prompt
[ ] Bound retrieval results (ex: max_num_results=3)
[ ] Validate tool outputs with code, not instructions alone
[ ] Separate session state from task state
[ ] Add tracing so failures are inspectable
[ ] Decide what should retry vs escalate vs fail fast
Enter fullscreen mode Exit fullscreen mode

That gets you much further than another 300 tokens of prompt engineering.

One more thing if you run agents all day

Once you start breaking agents into supervisors, workers, retries, and long-running loops, usage goes up fast.

That’s when per-token pricing starts messing with architecture decisions.

Teams end up optimizing for cost anxiety instead of reliability.
They avoid retries.
They shrink context too aggressively.
They hesitate to let automations run continuously.

That’s exactly why tools like Standard Compute are interesting for agent-heavy workflows.

It gives you an OpenAI-compatible API with flat monthly pricing, so you can run automations, supervisors, and worker loops without staring at token burn all day.

If you’re building in n8n, Make, Zapier, OpenClaw, or your own Python stack, predictable pricing changes the way you design systems.

You can afford to build the safer architecture instead of the cheapest-looking demo.

And for agents, that trade is usually worth it.

Top comments (0)