DEV Community

Cover image for Your Agent Is Paying an LLM to Run a For-Loop: Why Most Agent Harnesses Route Every Tool Call Through Inference
mech.app
mech.app

Posted on Originally published at mech.app

Your Agent Is Paying an LLM to Run a For-Loop: Why Most Agent Harnesses Route Every Tool Call Through Inference

Most agent frameworks make the LLM the bus between every tool call. Ask an agent to read every Markdown file and extract headings, and the default execution path looks like this: one inference to find the files, another to read file A, another to read file B, another to filter, another to aggregate. None of those hops requires reasoning. They are a deterministic sequence you could write in ten lines of Python. Yet you pay for every one of them in tokens, context window space, and wall-clock time.

The fix is small and almost everyone building production agents has converged on it: give the model a tool that runs a script, and move the loop out of the model and into the script. This is not a new idea. It is worth writing down anyway, because the why decides how you build it, and the problems it creates decide whether it survives contact with production.

The Loop Everybody Writes First

The default agent harness treats the LLM as the control plane for every step. Here's what happens when you ask it to read Markdown files and extract headings:

  1. LLM inference: "Find the Markdown files"
  2. Tool call: find_files("*.md")
  3. Result: A large list of file paths, back into context
  4. LLM inference: "Read file A"
  5. Tool call: read_file("A.md")
  6. Result: File contents, back into context
  7. LLM inference: "Read file B"
  8. Tool call: read_file("B.md")
  9. Result: File contents, back into context
  10. LLM inference: "Filter and aggregate"
  11. Final answer

Look at what is actually happening. The only genuinely non-deterministic step is deciding that the task should be "read the files, pull the headings." Everything after that (which files, hold this result, go read the next, keep only the headings, count them) is bookkeeping. The model gets invoked for the bookkeeping anyway, once per hop, and each hop drags its payload back through the context window.

The determinism is the tell: nothing in the middle of that chain needs a model. It is a for-loop wearing a chat interface.

Why Frameworks Ship This Way

Most agent frameworks default to LLM-as-bus for three reasons:

Simplicity: Every tool call goes through the same path. The orchestration loop is uniform. You don't need special cases for "this tool can chain" or "this tool needs a script runtime."

Flexibility: If the model decides halfway through that it needs to change direction, it can. The LLM sees every intermediate result and can re-plan.

Debuggability: Every step is a discrete inference. You can log it, replay it, and inspect the tool call and response.

These are real benefits. The problem is that they scale linearly with the number of steps, and most of those steps don't need the flexibility. You are paying for optionality you don't use.

The Cost Structure

Here's what you actually pay when the LLM is the bus:

Cost Dimension Per-Hop Impact Example (10-file loop)
Input tokens Full context + tool result each time 10 × (context + result size)
Output tokens Tool call JSON + reasoning 10 × ~50-200 tokens
Latency Inference + network round-trip 10 × ~500ms-2s
Context window Accumulates all intermediate results Can hit limits on large datasets
Rate limits One request per step 10 requests for a deterministic loop

If you are reading 100 files, you are making 100+ inference calls. If each file is 2KB and your context is 10KB, you are processing 1.2MB of input tokens for a task that could run in a single 50-line Python script.

The math gets worse when you add retries, error handling, or nested loops. Every conditional branch becomes another inference.

The Script Tool Pattern

The fix is to give the model a tool that runs code. Instead of:

LLM -> find_files -> LLM -> read_file(A) -> LLM -> read_file(B) -> ...
Enter fullscreen mode Exit fullscreen mode

You get:

LLM -> run_script("""
files = find_files("*.md")
headings = []
for f in files:
    content = read_file(f)
    headings.extend(extract_headings(content))
return headings
""") -> LLM
Enter fullscreen mode Exit fullscreen mode

The model writes the loop once. The script runtime executes it. The result comes back in a single hop.

This is not hypothetical. Most production agent harnesses (LangChain's PythonREPL, AutoGPT's execute_python_code, Semantic Kernel's code interpreter) ship a version of this tool. The implementation details vary, but the shape is the same: the LLM writes code, the harness runs it in a sandbox, the result goes back into context.

What You Gain

Token efficiency: One inference to write the script, one to process the result. No intermediate hops.

Latency: One round-trip instead of N. If the script takes 100ms to run, you save (N-1) × inference_latency.

Context window: Intermediate results stay in the script runtime. Only the final output goes back to the model.

Rate limits: One request instead of N. Matters when you are hitting per-minute quotas.

What You Lose

Observability: The loop is now opaque. You don't see each file read or each heading extracted unless you instrument the script runtime.

Error handling: If the script fails halfway through, the model doesn't see which step broke. You need to surface stack traces and partial results.

Re-planning: The model can't change direction mid-loop. If it realizes halfway through that it needs different data, it has to start over.

Security boundary: You are now running arbitrary code. Sandboxing becomes critical. You need to decide what the script can access (filesystem, network, environment variables) and enforce it.

Hybrid Control Flow

The smart move is not to route everything through the script tool. It is to route deterministic sequences through the script tool and keep the model in the loop for decisions.

Here's a decision tree:

  • Use the LLM directly when:

    • The next step depends on reasoning about the previous result
    • You need to re-plan based on intermediate data
    • The task is a single tool call or a short chain (2-3 steps)
  • Use the script tool when:

    • The task is a loop over a known set
    • The logic is deterministic (filter, map, reduce)
    • You are hitting token or latency budgets
  • Use a hybrid when:

    • The loop needs checkpoints (run 10 files, report back, decide whether to continue)
    • You need observability into intermediate steps
    • The script might fail and you want the model to handle retries

Example hybrid flow:

# LLM decides the plan
plan = llm.invoke("Read all Markdown files and extract headings")

# Script tool executes the deterministic part
script = """
files = find_files("*.md")
results = []
for f in files[:10]:  # Process in batches
    content = read_file(f)
    results.append({"file": f, "headings": extract_headings(content)})
return results
"""
batch_1 = run_script(script)

# LLM decides whether to continue
decision = llm.invoke(f"Processed 10 files. Results: {batch_1}. Continue?")
Enter fullscreen mode Exit fullscreen mode

This keeps the model in the loop for decisions but moves the execution out of the inference path.

Instrumentation

If you route through a script tool, you need to instrument it. Otherwise, you lose visibility into what the agent is actually doing.

Minimum instrumentation:

  • Log the script source before execution
  • Log the return value and any exceptions
  • Track execution time
  • Count tool calls made inside the script

Better instrumentation:

  • Trace each function call inside the script (file reads, API calls, database queries)
  • Surface partial results if the script fails mid-execution
  • Emit structured logs that your observability stack can parse
  • Track token costs per script invocation (input tokens for the script source, output tokens for the result)

Best instrumentation:

  • Sandbox telemetry: CPU, memory, network I/O
  • Replay capability: save the script and inputs so you can re-run it outside the agent
  • Diff tracking: if the model writes a similar script twice, show what changed

Security Boundaries

Running arbitrary code is a privilege escalation. The model can now do anything the script runtime allows. You need to decide what that is.

Filesystem access: Read-only? Write to a temp directory? Full access?

Network access: Localhost only? Specific domains? Unrestricted?

Environment variables: Can the script read secrets? Which ones?

Execution limits: Timeout? Memory cap? CPU quota?

Most harnesses use a sandboxed Python interpreter (RestrictedPython, PyPy sandbox, or a container with resource limits). The sandbox enforces the policy. The model doesn't get a vote.

If you are running in a multi-tenant environment, you need process isolation. One agent's script should not be able to read another agent's data or exhaust shared resources.

When the Script Tool Fails

The script tool introduces a new failure mode: the model writes code that doesn't run.

Syntax errors: The model writes invalid Python. The sandbox rejects it. You need to surface the error and let the model retry.

Runtime errors: The script runs but throws an exception (file not found, division by zero, API timeout). You need to decide whether to return the exception to the model or fail the entire task.

Partial results: The script processes 50 files and crashes on file 51. Do you return the first 50 results, or do you treat it as a total failure?

Infinite loops: The model writes a loop that never terminates. You need a timeout and a way to signal that the script was killed.

The best practice is to return structured errors:

{
  "status": "error",
  "error_type": "runtime",
  "message": "FileNotFoundError: config.yaml",
  "partial_results": [],
  "traceback": "Traceback (most recent call last):\n  File \"script.py\", line 5, in <module>\n    config = read_file('config.yaml')\nFileNotFoundError: config.yaml"
}
Enter fullscreen mode Exit fullscreen mode

This gives the model enough information to fix the script and retry.

Cost Comparison

Here's a concrete example. Task: read 100 Markdown files, extract headings, count them.

LLM-as-bus:

  • 1 inference to plan
  • 1 inference to find files
  • 100 inferences to read files (one per file)
  • 1 inference to aggregate
  • Total: 103 inferences
  • Estimated input tokens: 103 × 10KB context = 1.03MB
  • Estimated output tokens: 103 × 100 tokens = 10,300 tokens
  • Estimated latency: 103 × 1s = 103 seconds
  • Estimated cost (GPT-4): ~$30

Script tool:

  • 1 inference to plan
  • 1 inference to write the script
  • 1 script execution (100 file reads, local)
  • 1 inference to process the result
  • Total: 3 inferences
  • Estimated input tokens: 3 × 10KB context + script source = 30KB + 2KB
  • Estimated output tokens: 3 × 100 tokens = 300 tokens
  • Estimated latency: 3 × 1s + 0.5s script execution = 3.5 seconds
  • Estimated cost (GPT-4): ~$1

The script tool is 30× cheaper and 30× faster. The gap widens as the loop size increases.

Technical Verdict

Use the script tool when:

  • You are processing more than 5-10 items in a deterministic loop
  • Token costs or latency are a constraint
  • The logic is simple enough that the model can write it in one shot
  • You have a sandboxed runtime and can instrument it
  • The task is embarrassingly parallel (reading files, calling APIs, filtering data)

Avoid the script tool when:

  • The task requires re-planning based on intermediate results
  • You need full observability into every step for compliance or debugging
  • The model is unreliable at writing correct code for your domain
  • You don't have a secure sandbox or can't enforce resource limits
  • The loop is short (fewer than 5 iterations) and the overhead of script generation exceeds the savings

Hybrid approach when:

  • You need checkpoints (process in batches, report back, decide whether to continue)
  • The loop might fail and you want the model to handle retries with context
  • You are debugging and need to see intermediate steps without sacrificing production performance
  • The task mixes deterministic execution with decision points (process files until you find a match, then stop and ask what to do next)

The script tool is not a replacement for the LLM. It is a way to keep the model in the loop only when reasoning is actually needed. Most production agents need both: the LLM for planning and decision-making, the script runtime for execution. The trick is knowing when to route through which path, and that decision is driven by cost, latency, and observability requirements.

Source Links

Top comments (1)

Collapse
 
autenai profile image
Auten •

The same for-loop shows up even more painfully with agents that drive a GUI: every click is an inference plus a screenshot or accessibility dump back into context, so a 12-step form fill can cost more than the reasoning that chose to fill it.

One thing your "When the Script Tool Fails" section points at that gets sharper on screens: the deterministic sequence drifts. A button moves or gets renamed and the recorded steps break, not because of a syntax error but because the world changed. What has worked for us is your hybrid shape with the checkpoint at the failure point: replay the recorded steps with no model call, and only hand control back to the model at the exact step that no longer matches, then save the repaired step so the next replay is model-free again.

Your "replay capability" bullet under best instrumentation ends up being the whole product in that setting. (Disclosure: I'm on the Auten team, we build this for Claude Code / Cursor driving a real screen.)

Some comments may only be visible to logged-in visitors. Sign in to view all comments.