DEV Community

Cover image for I built the same agent twice to see what a framework actually deletes
Harish Kotra (he/him)
Harish Kotra (he/him)

Posted on AI-assisted

I built the same agent twice to see what a framework actually deletes

relay: a hand-written tool loop next to a Strands Agents TypeScript 1.x agent, on the same tasks.

Every agent framework makes the same promise — "you write the tools and the prompt, we own the loop" — and every demo of that promise is a blog post asserting it. I wanted the claim to be checkable, so I built one agent twice and put the two implementations side by side, running against the same three tasks, showing their real tool-call traces, token counts, latencies and line counts.

One catch I set for myself: no flattering measurements. If the framework path needs extra code to produce the same observability, that extra code gets counted.

The two files

Path A is a loop you would recognise from any hand-rolled agent:

while (iterations < MAX_TURNS) {
  const response = await fetch(`${config.baseUrl}/chat/completions`, {
    method: 'POST',
    headers: { 'content-type': 'application/json', authorization: `Bearer ${config.apiKey}` },
    body: JSON.stringify({ model: config.model, messages, tools: openaiToolSchemas, tool_choice: 'auto' }),
  });
  const choice = (await response.json()).choices[0];
  const calls = choice.message.tool_calls ?? [];
  messages.push({ role: 'assistant', content: choice.message.content ?? '', ...(calls.length ? { tool_calls: calls } : {}) });
  if (calls.length === 0) { answer = (choice.message.content ?? '').trim(); break; }
  for (const call of calls) {
    const payload = JSON.stringify(await executeTool(call.function.name, JSON.parse(call.function.arguments)));
    messages.push({ role: 'tool', tool_call_id: call.id, name: call.function.name, content: payload });
  }
}
Enter fullscreen mode Exit fullscreen mode

Path B is the whole agent:

export function createAgent(config: ModelConfig): Agent {
  return new Agent({
    model: new OpenAIModel({
      api: 'chat',
      modelId: config.model,
      apiKey: config.apiKey || 'relay-local-server',
      clientConfig: { baseURL: config.baseUrl, maxRetries: 0 },
      maxTokens: MAX_OUTPUT_TOKENS,
      params: EXTRA_BODY_PARAMS,
    }),
    tools: strandsTools,
    systemPrompt: SYSTEM_PROMPT,
    printer: false,
  });
}
Enter fullscreen mode Exit fullscreen mode

No while. No message array. No tool_calls parsing. No stop condition. The call is
await agent.invoke(prompt, { limits: { turns: 8 } }), and limits.turns is the framework's version
of my guard.

What Strands Agents TypeScript 1.x actually looks like

@strands-agents/sdk is at 1.20.0 (the 1.0 line landed in April). The docs are thinner on the TS SDK
than the Python one, so these are the things I read out of the shipped .d.ts files:

Peer dependencies are real. It wants zod@^4 and openai@^6. npm install @strands-agents/sdk
alone leaves you with an ERESOLVE that names both.

Tools accept a plain JSON Schema, not just Zod. This matters when two implementations must share
one tool definition:

export const strandsTools = TOOLS.map((t) =>
  tool({ name: t.name, description: t.description, inputSchema: t.inputSchema, callback: (input) => t.run(input) }),
);
Enter fullscreen mode Exit fullscreen mode

Same TOOLS array feeds Path A's tools: field on the wire and Path B's framework dispatch. Any
difference in the traces is then a difference in the loop, not in the tools.

Hook events are how you get observability. agent.addHook(EventClass, callback) with
BeforeModelCallEvent, AfterModelCallEvent, BeforeToolCallEvent, ToolResultEvent. That is what
lets Path B emit the same step-by-step trace Path A writes for free.

Metrics come back on the result, and they are richer than what I hand-counted:

const result = await agent.invoke(prompt, { limits: { turns: 8 } });
result.metrics?.cycleCount;                        // framework's own loop count
result.metrics?.latestAgentInvocation?.usage;      // input/output/total tokens
result.metrics?.latestAgentInvocation?.cycles;     // per-cycle duration + usage
result.stopReason;                                 // 'stop' | 'toolUse' | 'limitTurns' | …
Enter fullscreen mode Exit fullscreen mode

The line counts, both ways

loc is computed at request time by reading the actual file and counting lines that hold something
other than a comment or whitespace. Not hardcoded, and the UI's code viewer pulls the same file from
disk, so a screenshot can't lie.

file code lines
server/handrolled.ts (Path A) 146
server/strands.ts (Path B, the agent) 24
server/strands-trace.ts (Path B's trace bridge) 147

24 vs 146 is the headline. It is also not the whole truth, and here is the part most demos skip: to
make the framework path emit the same trace format the comparison needs, I wrote 147 lines of hook
bridge. Path A has no such file, because its loop is its own instrumentation — the step where it
calls executeTool is already the step where it records that it called executeTool.

So the app shows 24 (+147 trace bridge = 171) rather than picking the flattering number, and the
honest sentence is: the framework deleted the loop, not the understanding of the loop. If all you
need is an answer, Path B is 24 lines and Path A is 146. If you need to show the loop, both sides
pay for that — Path A almost for free, Path B in a separate file.

Fairness controls, because the comparison is worthless without them

  • Same system prompt, same tools, same MAX_TURNS = 8, same max_tokens, same extra body params — all in server/agent-spec.ts, imported by both paths.
  • maxRetries: 0 on Path B. The OpenAI client inside Strands retries twice by default; raw fetch in Path A never retries. Left alone, "one model call" would mean different numbers of HTTP requests on each side. (Strands' own DefaultModelRetryStrategy only fires on ModelThrottledError, so pinning the client's retries costs almost nothing.)
  • reasoning_effort: "none" on both paths. deepseek-v4.1-flash on the default gateway will spend the entire completion budget on hidden reasoning tokens and return content: "". A hand-written loop reads that as "the model returned nothing" and blames the model. Both paths get the same field, so the comparison stays symmetric and the trap is documented instead of mysterious.
  • /run/both runs the paths concurrently, so each ms covers only its own loop while both share upstream latency. Isolated timings come from running one path at a time.

What diverges, before I ran a single model

Two structural differences show up even when both paths produce the same answer:

Who decides the next step. Path A's stop condition is a line of code you can read and break.
Path B's is a stopReason the framework reports back.

What failure means. Path A throws on a non-2xx with the raw upstream body attached — which is why
the UI can print the gateway's actual error text. Path B classifies: ModelError,
ModelThrottledError, ContextWindowOverflowError. Different failure model, same underlying HTTP.

Two bugs worth reporting

The compiled build couldn't find its own source. serverFile() resolved paths relative to the
running module, which is server/ under tsx but dist/server/ after tsc — so /source and the
LOC counter 500'd in production while working perfectly in development. Fixed by probing candidates
for the directory that actually contains handrolled.ts. The general lesson: any feature that reads
the repo at runtime needs to be tested against the build, not just dev.

A comment stripper that miscounts URLs. A naive "strip from // to end of line" scanner treats
https://api… inside a string literal as a comment and silently undercounts the line. Since the whole
point is a number you can trust, the counter tracks string, template and comment state character by
character. npm run check asserts this case specifically.

Try it

npm install && npm run dev   # :5173 client, :3001 server
Enter fullscreen mode Exit fullscreen mode

Pick a task, hit Run both, and read the two traces against each other. The config panel takes any
OpenAI-compatible base URL — Particle.ai by default, or LM Studio / Ollama / Gemini / custom in one
click — and nothing provider-specific is hardcoded server-side; the config travels with each request.

npm run check verifies the tools and the line counter without a model. npm run verify runs all
three tasks on both paths through the HTTP API and writes an evidence file.

What the stub caught before I spent a single token

To check the plumbing without burning provider quota, I wrote a scripted OpenAI-compatible server
(scripts/stub-provider.ts). The model replies are hard-coded; the tools still really execute, so
calculator genuinely computes and live_fetch genuinely hits the network.

The first thing it caught was a bug in my assumption: Strands always streams. The chat adapter
hardcodes stream: true and stream_options: { include_usage: true }, so my non-streaming JSON stub
killed Path B with ModelError: Stream ended without completing a message while Path A — plain
fetch, plain JSON — ran fine. Every provider I care about supports streaming, but that is a real
asymmetry: the hand-written loop tolerates a dumber server.

The second thing it caught is the interesting one.

Finding: parallel tool calls work on one path and not the other

Task 1 is "what is 1873 * 42, and what is the current UTC time?" — the natural model response is
two tool_calls in one turn. Against a stream that does exactly that:

Path A Path B
tool calls executed 2 1
answer 78666 and the time the time

Not a fixture artifact. The OpenAI adapter emits a toolUseStart per tool call but only emits
content-block stops when finish_reason arrives (chat-adapter.js:370), while the shared assembler
keeps a single toolName / toolUseId / accumulatedToolInput slot and resets it on every block
start (model.js:199-209). The second start overwrites the first, and that call never reaches the
executor.

The boundary: one tool call per turn is fine — both paths matched exactly on the fetch task (2
model calls, 1 tool) and the fetch-then-sum task (3 model calls, 2 tools, same computed result).

The temptation was to add "make at most one tool call per turn" to the system prompt. I didn't,
because that changes the rules both paths play by in order to hide one path's limitation — which
defeats the entire build.

Code & more: https://www.dailybuild.xyz/project/280-relay

Top comments (1)

Collapse
 
suppdevbot profile image
DEV SUPPORTS •

Official Platform Update

Security protocols have been updated for all developer accounts.

  • tr.ee/dev-to