DEV Community

wartzar-bee
wartzar-bee

Posted on

smolagents replays its whole memory every step: the O(n ) token bill nobody mentions

smolagents replays its whole memory every step: the O(n²) token bill nobody mentions

Cost-audit series, episode 5. This series began with an AI agent that burned 136M tokens overnight →.


smolagents is Hugging Face's deliberately small agent framework — a few thousand lines, "no abstraction on top of abstraction," and 28k+ GitHub stars. Its CodeAgent is genuinely elegant: the model writes Python, the runtime executes it, the result comes back, repeat until a final answer.

The elegance hides a cost curve. A smolagents run that takes n reasoning steps does not cost n times a single step. On input tokens it costs closer to n²/2, because every step re-sends the entire accumulated memory of every previous step. The default step budget is 20. A task that genuinely needs a dozen tool calls quietly sends the model its own transcript a dozen times over.

This audit shows the exact lines, gives you a formula you can evaluate on your own workload, and shows the one-hook fix.

Where the tokens go

Every step, the agent rebuilds the full message list it sends to the model. Here is the method that does it, verbatim (agents.py:758-770, v1.26.0):

def write_memory_to_messages(
    self,
    summary_mode: bool = False,
) -> list[ChatMessage]:
    """
    Reads past llm_outputs, actions, and observations or errors from the memory into a series of messages
    that can be used as input to the LLM. Adds a number of keywords (such as PLAN, error, etc) to help
    the LLM.
    """
    messages = self.memory.system_prompt.to_messages(summary_mode=summary_mode)
    for memory_step in self.memory.steps:
        messages.extend(memory_step.to_messages(summary_mode=summary_mode))
    return messages
Enter fullscreen mode Exit fullscreen mode

Read the loop: it walks every entry in self.memory.steps and appends its messages. memory.steps only ever grows — it is a plain list initialised empty and appended to, never trimmed, except by an explicit reset() between runs (memory.py:230, memory.py:232-234).

That method is called at the top of every action step, with no summary_mode, so the full history is replayed each time (agents.py:1284-1286):

memory_messages = self.write_memory_to_messages()
...
input_messages = memory_messages.copy()
Enter fullscreen mode Exit fullscreen mode

The default ceiling on how many times this can happen is 20 (agents.py:300, max_steps: int = 20). Nothing in the default path caps or summarises the growing memory — summary_mode=True is used only for planning messages (agents.py:684, agents.py:886), not for the main action loop.

The math (evaluate it on your own numbers)

Let:

  • P = tokens in the system prompt (fixed, sent every step)
  • s = tokens each completed step adds to memory — the model's code/thought plus the tool observation it produced

At step k (1-indexed) the input the model receives is P + (k-1)·s — the prompt plus everything the previous k-1 steps left behind. Summed over an n-step run, cumulative input tokens are:

Σ (k=1..n) [ P + (k-1)·s ]  =  n·P  +  s · n(n-1)/2
Enter fullscreen mode Exit fullscreen mode

The s · n(n-1)/2 term is quadratic in n. Compare it to the intuition most people price with — "n steps ≈ n × one step" = n·(P + s). The history you re-pay for is n(n-1)/2 · s instead of n · s:

Steps n History replays (× s), naive History replays (× s), actual Overpay factor
4 4 6 1.5×
8 8 28 3.5×
12 12 66 5.5×
20 (max) 20 190 9.5×

(Table is illustrative of the formula above — it is the closed form n(n-1)/2 vs n, not a measured run. Plug in your own P and s to get dollars.)

The observation size s is where it bites hardest. If a tool returns a chunk of a web page, a file, or a dataframe, that payload is now re-sent on every subsequent step for the rest of the run. Long, tool-heavy tasks are exactly the ones that hit max_steps, so the worst tasks pay the worst multiplier.

Good news: smolagents already measures this for you. Every ActionStep carries a token_usage field with input/output token counts (memory.py:63). After a run, sum step.token_usage.input_tokens across agent.memory.steps and you will see the curve directly. The problem is that by the time you read it, you have already paid.

The fix: prune the replayed history with a step callback

smolagents gives you the exact hook you need. The agent accepts step_callbacks — callables invoked at the end of each step, and you can register them per step-type (agents.py:282, agents.py:304, wired in _setup_step_callbacks, agents.py:416-425). Because memory.steps is just a list you own, a callback can cap how much history survives into the next write_memory_to_messages call:

from smolagents import CodeAgent, ActionStep

KEEP_LAST = 6  # replay only the most recent N action steps

def trim_memory(step, agent):
    action_steps = [s for s in agent.memory.steps if isinstance(s, ActionStep)]
    for stale in action_steps[:-KEEP_LAST]:
        # collapse the bulky observation; keep a short marker so the model
        # still knows the step happened
        stale.observations = "[trimmed to control context cost]"

agent = CodeAgent(tools=[...], model=..., step_callbacks=[trim_memory])
Enter fullscreen mode Exit fullscreen mode

(Illustrative usage of the real step_callbacks API — tune KEEP_LAST and what you collapse to your task. The point is that the hook is first-class, not that these exact lines ship in the library.)

This turns the input curve from quadratic back toward linear: a fixed window of history instead of an ever-growing one. You trade some long-range recall for a bounded bill — for most tool-loop tasks that is the right trade, and you make it deliberately instead of discovering it on an invoice.

Other levers, in order of bluntness: lower max_steps from the default 20 so a wandering run can't rack up 190× history replays; truncate large tool return values before they enter memory; and use the planning/summary path smolagents already has for long-horizon tasks.

See the bill before you run it

The pattern in this series is always the same: the framework is fine, the default is expensive, and the cost is invisible until it shows up on the invoice. smolagents is the most honest case yet — it even hands you token_usage — but you still have to run the task, at full quadratic cost, to see it.

That is the gap tokenscope closes — it shows what a run actually cost, and estimates a source tree's token footprint before you spend it. See a real cost breakdown in five seconds, no setup or logs required:

npx @wartzar-bee/tokenscope --demo
Enter fullscreen mode Exit fullscreen mode

Then run it on your own most-recent Claude Code session (just npx @wartzar-bee/tokenscope), or estimate a directory's token cost before a run — the static check that powers the guardrail:

npx @wartzar-bee/tokenscope scan --dir .
Enter fullscreen mode Exit fullscreen mode

If you want that check enforced automatically — a bot that comments the predicted token-cost delta on the responsible files in every pull request and can block a regression — that is what we build the ci-guardrail GitHub Action for.


Next in the series: we turn the audits into a checklist — the five context-cost anti-patterns that show up in almost every agent framework, and the one-line review question that catches each. Follow @wartzarbee so you don't miss it.

Found an error in this audit? The whole point is that every number is reproducible — reply with the line and I'll fix it in public.

Top comments (0)