DEV Community

Qweezyy
Qweezyy

Posted on

Our agent said "done" on 15% of tasks while the provider was failing

TL;DR. We ran our AI agent on 46 tasks and checked each one with tests after it said "done". 7 of the 46 — 15% — "done"s were untrue. Not because of the model: not one task failed because the model couldn't solve it. The provider was to blame. It answered with HTTP 200 and sent its own error text instead of the model's reply, or an empty stream, or the model looped on its side — and the agent took any end of a stream for the end of the task. Below: the three shapes of this failure, how to detect them, why a second reviewer doesn't help, and what it means for "done" in general.

About this post. The project is Altair, an open-source (Apache-2.0) AI agent for your PC and phone. I'm the author and build it largely with Claude Code. This post, the experiments and the charts were prepared by that same AI assistant; I reviewed it and stand behind it. Numbers are from our runs, code is from the repo.

How it started: "completed ≠ accepted"

After the first post about Altair, Kent Bodrov left a comment worth retelling. We had boasted that the agent snapshots every edit and runs the tests before "done". Kent pointed out that a snapshot only proves rollback material exists, and a green command only proves a check ran. Neither proves the wanted behaviour is there. "Completed" isn't "accepted".

We wrote it into the plan and went measuring. It turned out there's a simpler question below "was it accepted": was it completed at all.

How we measured

The real Altair, a cheap model (glm-5.3-flash) through two providers, 46 tasks of five kinds: fix a bug from failing tests, answer a question about code, compute from a CSV, write a module with tests, go through long code. Each task has an automatic check that runs after "done". A proxy between agent and provider logged every request and answer: how many tokens the provider counted, how much text came, how long the stream lasted.

Results

46 tasks the agent called done
46 tasks the agent called done

39 tasks were really solved. 7 weren't, and in all seven the provider broke:

  • 6 times — the gateway's error text instead of an answer. HTTP 200 and an ordinary assistant message: "The request could not be completed. Please retry later, or reduce the request parameters/content." The agent showed it to the user as its answer, the CLI returned "ok": true, task "done" in one step.
  • Once — an empty stream. The provider held the connection for 103 seconds and closed it without a single token: no text, no tool call. The agent wrote "Task finished, but the model returned no text" — after two steps out of an intended twenty.

The next series, on hard tasks, showed a third shape:

  • A loop cut short. The model looped on endless reasoning. Our loop guard cut the stream at 160,000 characters — and the agent loop took the cut stream for the final answer. That ended 8 of 67 runs where no reasoning level was set. (Why the model looped and how one field in the request fixes it — in the companion post.)

Why the agent believed it

All three go through the same place. The agent loop works like most:

model answered → is there a tool call? → run it, continue
               → no tool call?        → that's the final answer, task done
Enter fullscreen mode Exit fullscreen mode

The "no tool call" branch can't tell "the model finished the work" from "the provider sent anything without a tool call". And there was never an HTTP error, the thing that triggers retries and the backup provider: everything came with 200.

The telling part is that the proxy saw what the agent didn't. For one of these "answers" the provider counted 4,552 input tokens. The agent's system prompt and tool schemas alone are about 9,300 tokens. The provider never even read the request, and the agent got an "answer" and closed the task.

How to detect it

We now check every answer without a tool call. It's not an answer if:

def judge_turn(turn, messages, tools, *, cut=False):
    if turn.tool_calls:
        return None
    if cut:
        return "cut"            # the loop guard cut the stream
    text = (turn.content or "").strip()
    if not text:
        return "empty"          # neither text nor a tool call
    low = text.lower()
    if len(text) < 400 and any(p in low for p in GATEWAY_ERROR_PHRASES):
        return "gateway"        # a short reply that is a gateway's error text
    prompt = (turn.usage or {}).get("prompt_tokens") or 0
    expected = _text_chars(messages, tools) / 4.5
    if prompt and expected > 2000 and prompt < 0.5 * expected:
        return "unprocessed"    # the provider read less than half of the request
    return None
Enter fullscreen mode Exit fullscreen mode

Such a reply is an exception the client retries like a dropped connection: after a pause, or on the backup provider if one is set. Three details without which this would work badly:

  • The first ~160 characters are held back. Otherwise the gateway's error text reaches the chat and stays there after the retry. The delay is a fraction of a second, only at the start of an answer.
  • The phrases only count in a short reply. If the model, in a normal long answer, tells the user that "the request could not be completed", that's an answer, not an error.
  • A loop is retried once, with a low reasoning level. If the model loops again, the agent reports an honest error instead of five more 160K-character attempts.

An empty final answer is no longer "task finished" either: it's a failure, not a result.

What about a second reviewer?

The first thought after Kent's comment is to add a separate model call that checks the result against the task after "done". We tried that too: the reviewer sees the task and every file, answers "OK" or a list of violations, and the agent fixes what it flags.

A second reviewer after done
A second reviewer after "done"

Solved before review After Flags (false) Cost per task
Agent alone — 16 of 18 — 3.8 m$
+ reviewer, low 20 of 21 20 of 21 4 (3) 5.5 m$
+ reviewer, high 15 of 18 16 of 18 5 (4) 12.8 m$

With a frugal reviewer: not one task fixed, for +45% cost. With an expensive one: one task fixed out of 18, at 3.4x the cost, and it missed one. Most flags were false: the agent checked and dismissed them without breaking anything, but we paid for each.

The reason is simple. With a properly set reasoning level the agent already solves about 90% of these tasks. And the real false "done"s didn't come from a model that left work unfinished, but from infrastructure the reviewer can't see at all. The reviewer reads files — and a provider reply with an error text never got as far as the files.

A provider that doesn't fail

Separately, for six hours, every 1–1.5 minutes, we sent the same provider real-size requests. 297 of 297 eventually got an answer, only 4 on a second try. Not a single "error disguised as an answer" this time.

But around 14:00 the provider sagged and stayed 4–5x slower until the end: the median went from 6 to 22–29 seconds, single answers took 1.5–4 minutes. Such slumps last hours, and retrying doesn't help — it joins the same queue. Only a backup provider does. So our breaker trips not only on errors but on slowness: if the first byte came later than 40 seconds twice in a row, the next steps go to the backup model for 10 minutes.

One more lesson from that test: put the timeout on silence, not on the whole request. Our phone app had a total 120-second limit per request — and long but healthy answers were cut at second 120 during slow hours. On the PC the same 120 s meant "no new data for 120 s", and a flowing stream got to the end.

The takeaway: check "done" at the infrastructure level

Kent is right: "completed" isn't "accepted". Our measurements add a rung below. Before asking whether the model did what was asked, the agent has to make sure the model answered at all.

  • The end of a stream isn't the end of a task. An empty answer, a cut answer, an answer to a request the provider didn't read — these are failures and must be handled as failures.
  • HTTP 200 guarantees nothing. Gateways have their own errors that arrive as ordinary model replies. The provider's input-token count is a good detector: it doesn't lie about what the model actually read.
  • Checking "done" is engineering, not another model. Tests after the work, checking the provider's answer, an honest error instead of "finished". A second reviewer is expensive and mostly noise.
  • Measure afterwards instead of trusting the process. We found all this only because each task was checked by tests after "done".

What this doesn't fix

  • The detector is heuristic. The list of gateway phrases is incomplete, the "read less than half" threshold was tuned on our data. Another gateway may fail in its own way.
  • "Accepted" still isn't checked. We closed "was it completed". Checking "is it the right behaviour" — a separate acceptance step against criteria from the user's request — is only planned.
  • Small samples. 46 tasks in the first series, 18–21 runs per reviewer setting. How often this happens depends heavily on the provider's load: in calm hours it almost doesn't.

Check it yourself

The detection code is pc/core/llm/reliability.py with its tests pc/tests/test_reliability.py in https://github.com/Qweezyy/AltairAgent. The stand (logging proxy, tasks with checks, summary) and the raw results: research/agent-lab-2026-10 (python lab/facts.py recomputes every number in this post from the saved results, no keys needed).

How do you check that your agent has really finished? Have you met providers that send errors with HTTP 200?

Top comments (2)

Collapse
 
ahmetozel profile image
Ahmet Özel •

The six HTTP-200 gateway messages are a concrete example of transport success being weaker than a completed model turn. I would treat the input-token-count detector as a signal to investigate rather than proof that the provider did not read the request, because usage accounting can vary around cached prompts and tool schemas.

A regression case worth adding is a legitimate short response quoting one of your gateway phrases, alongside a gateway error that includes a long diagnostic block. Those exercise both sides of the 400-character cutoff. Separately recording protocol termination, detector reason and task-check result would help distinguish classification mistakes from failures the provider actually caused.

Collapse
 
reidmarlow profile image
Reid Marlow •

The empty tool-call fallback is where almost every custom agent loop gets burned. Relay proxies and rate limiters love returning HTTP 200 with an error string placed neatly inside choices[0].message.content. If the harness only checks whether tools were invoked, that gateway message gets handed straight to the caller as a successful final turn. I had to add explicit finish_reason validation and an egress payload check before letting the runner exit with zero status.