The two sentences your agent cannot tell apart
A tool in your agent calls an API. The call times out. Something catches the timeout, logs a warning, and returns an empty list.
Another tool calls the same API. It works perfectly. Nothing matches the query, so it returns an empty list.
The model sees the same thing in both cases. So it picks the useful interpretation and tells your user there are no results, with total confidence, because nothing in what it read suggested otherwise.
Two sentences, opposite meanings:
- "There is nothing there" means the search was conclusive.
- "I could not check" means you know nothing at all.
The second collapses into the first, silently, at the point where a Python function has to pick a return value. Nothing errors. No alert fires. The only trace is a log line, and your users do not read your logs.
This is not a post about a hypothetical. I ran a script over my own agent repo, found 24 instances of the pattern, proved the model literally cannot distinguish the two cases, and then found that my own eval instrumentation had made the same mistake. All of it is below, with the code to run the check yourself.
Nobody writes this bug. The type signature writes it
Here is the code, and it is careful code:
def search_suppliers(query: str) -> list[dict]:
try:
return client.get(query, timeout=15).json()["results"]
except (TimeoutError, httpx.HTTPError):
log.warning("supplier search failed")
return []
There is a try. There is a log. Nothing is swallowed in silence. A reviewer would approve this.
The problem is the annotation. The function promises to return list[dict], and list[dict] has no member that means "I failed". The only thing the function can say on a bad day is [], which is exactly what it says on a quiet one. The type system is not being violated here, it is being obeyed. There is simply no vocabulary for the third state.
Every layer downstream then inherits the ambiguity:
results = search_suppliers(q)
if results:
...
That if is where the distinction dies for good. After it, nothing in the process knows whether the upstream was healthy.
And the layer that matters most is the one people forget: the string that goes into the model's context. You can fix the Python types perfectly and still ship the bug, because what the model reads is whatever your handler serialised. A beautifully typed error that renders as [] in the tool message is the same bug with extra steps.
So there are two distinct jobs:
- Give the return value a vocabulary for failure.
- Make sure that vocabulary survives into the text the model reads.
Most discussions of Result types stop after the first one.
Step one: count it in your own repo
Before building anything I wanted to know how common this is in code I had written myself. The check is an AST pass: walk every except handler, look for a return of an empty literal inside it.
import ast
import pathlib
CONTAINERS = (ast.List, ast.Dict, ast.Set, ast.Tuple)
def is_empty_result(node: ast.expr) -> bool:
if isinstance(node, CONTAINERS):
return not getattr(node, "elts", None) and not getattr(node, "keys", None)
return isinstance(node, ast.Constant) and node.value is None
for path in pathlib.Path(".").rglob("*.py"):
try:
tree = ast.parse(path.read_text(encoding="utf-8"))
except SyntaxError:
continue
for handler in (n for n in ast.walk(tree) if isinstance(n, ast.ExceptHandler)):
for node in ast.walk(handler):
if isinstance(node, ast.Return) and node.value is not None:
if is_empty_result(node.value):
print(f"{path}:{node.lineno}")
Run it at the root of your agent repo. It takes seconds, and the number it prints is the number of places where a failure in your system currently looks exactly like an absence of data.
Mine printed 24 lines. Fifteen of them were in a single file: the one that talks to external APIs.
Two caveats so you read your own output correctly. It walks into functions defined inside a handler, which is rare but will inflate the count slightly. And it only catches literals, so a handler that returns a variable which happens to be empty, or a dict like {"results": [], "error": str(e)}, does not show up. That second pattern is the more dangerous one, and it is the one that later bit me.
A grep gets you close, but only the AST version distinguishes an empty return inside a handler from one anywhere else in the function, which is exactly the distinction you care about.
Step two: diff the prompt, not the code
A count tells you the pattern exists. It does not tell you whether it reaches the model. For that, render the prompt twice and compare the strings.
failed = build_prompt(query, sources=simulate_all_sources_failing())
empty = build_prompt(query, sources=[])
print(failed == empty)
Mine printed True.
Byte for byte identical. Same zero count, same empty section underneath it. Seventeen upstream sources could be on fire and the prompt is indistinguishable from a quiet afternoon where nothing matched. No amount of prompt engineering fixes that, because there is no signal in the input to engineer against.
If you do one thing from this article, do this one. It takes ten minutes, it needs no library, and the result is binary. Either your model can tell the two situations apart or it provably cannot.
The part I did not expect
While reading the prompt I noticed an instruction I had written months earlier, and it is the sort of line that appears in most agent prompts:
If there are fewer than 5 results, suggest broadening the search.
Read that again in the failure case. Everything upstream is down, the result list is empty, so the agent does not merely report nothing. It tells the user their query was too narrow and advises them to lower their expectations.
A failure gets converted into advice. The user walks away with a false belief and an action item.
There is a quieter version too. If one source is healthy and sixteen are down, the user sees a short list that looks complete. There is no empty state to notice, nothing visibly wrong, and no way for anyone to tell that most of the system was unavailable.
That is the case that worries me most, because an empty answer at least invites suspicion. A plausible short answer does not.
The fix, in three parts
I spent three weekends on this and packaged it as neverempty. It is on PyPI, pydantic is the only dependency, and it does three things. The three matter in order, because each one is useless without the next.
1. Give the return value a third state
from neverempty import tool, Ok, Empty, Err
@tool(empty_when=lambda rows: len(rows) == 0, timeout_s=8.0)
async def search_suppliers(query: str) -> list[dict]:
return client.get(query).json()["results"]
result = await search_suppliers("steel, FY25")
match result:
case Ok(value=rows): ...
case Empty(): ...
case Err(kind="timeout"): ...
The wrapper turns any exception into Err with a classified kind, and never raises. Two design choices are worth calling out, because both of them came directly from the failure mode above.
There is no falsiness inference. 0, False and "" are legitimate values and are never read as empty. You declare what empty means for that tool, in empty_when, at definition time. If you declare nothing and return None or [], you get a loud error rather than a guess, because the moment the library starts guessing it has reintroduced the bug it exists to prevent.
And truncated is a flag on Ok, not a fourth state. The sibling bug to this one is a query capped at 100 rows that the model reports as "there are 100 records".
2. Make the failure survive into the prompt
This is the half that gets skipped. What the model reads:
{"status": "error", "kind": "timeout",
"note": "The tool failed. You do not know whether matching data exists. Do not say that no data exists."}
against the genuinely empty case:
{"status": "empty",
"note": "The query succeeded and returned no matching records."}
Blunt, deliberately. The instruction is doing real work: without it, a model handed {"status": "error"} will still often summarise it as nothing found, because nothing told it not to.
3. Measure whether any of this worked
The first two parts are assertions until you test them. The problem is that real timeouts are far too rare in a test run to measure by waiting, so you inject them:
from neverempty import faults
faults.inject(tool="search_suppliers", kind="timeout")
Then run a dataset of cases through the agent and count how often the final answer still claims absence. That number is the whole point of the library. A typed result is a hypothesis; the misreport rate under injected faults is evidence.
The part I would rather not publish
I wired the library into my own agent. Here is the glue I wrote:
@tool(empty_when=lambda r: not r.get("results"))
async def run_search(query: str) -> dict:
...
And here is what the function returns when the database call fails:
return {
"response": "I encountered an issue while searching. Please try again.",
"results": [],
"error": str(e),
}
The predicate asks whether results has anything in it. It does not. So my instrumentation looked at a failure, agreed it was empty, and recorded it as Empty.
The library did exactly what I told it. The one line connecting it to my code was wrong, and it was wrong in precisely the way the library exists to prevent.
I could have quietly fixed it, but it is the most useful thing I learned, so here is the general lesson. The bug is not a mistake somebody makes once and then knows better. Empty and broken keep arriving in the same shape at every boundary: the HTTP client, the function return, the dict that wraps it, the prompt, and the predicate you write to classify it. Every single layer has to be told the difference explicitly, and any layer that infers it will get it wrong eventually.
The fix in my case was to stop reading the payload and read the status instead:
@tool(empty_when=lambda r: r.get("status") == "ok" and not r["results"])
Which only works because the inner function now returns a status at all. That is the real cost of this problem: you cannot classify correctly at the edge if the information was destroyed three layers down.
What I have not measured
The number this library exists to produce, the misreport rate under injected faults on a production agent, I do not have yet. The machinery runs, the harness is tested, and the premise is sitting in real code in my own repo and in several popular frameworks. The headline figure is still ahead of me.
I am saying that explicitly for a reason. The pitch here is "trust these numbers", and a tool making that pitch cannot start with an unmeasured claim in its own README. The README says the same thing, in the same words.
What I can state, with the method attached so you can check it:
| Claim | Status |
|---|---|
| 24 empty returns inside error handlers in my repo | Measured, AST pass, script above |
| Failure and empty prompts are byte-identical | Measured, string comparison |
| My own glue classified a failure as empty | Observed, fixed, described above |
| How often the model then claims absence | Not measured yet |
| That this pattern is common across the ecosystem | Partially: found in three widely used agent frameworks, not quantified |
If you are evaluating any eval tooling, including this one, that table is the shape of answer worth asking for. "Our framework improves reliability" is not a number. Neither is a benchmark with no sample size next to it.
You probably do not need the library
The core of this is 785 lines across four files: the result types, the error classification, the fault injection, and the failure scoring. If you would rather copy those four files into your own repo than take a dependency, the README tells you to do exactly that, and I mean it. I would rather the pattern spread than the package.
What I would not skip, in any form:
- Run the AST script. The count takes four minutes and it is the cheapest diagnostic in this article.
- Diff the two prompts. If they match, your model is guessing and no prompt tuning will fix it.
- Read your own prompt for instructions that turn thin results into advice. Mine told users to broaden their search while the backend was down.
- Inject a fault on purpose, in staging, and read what the agent actually says. That is the only way to see the behaviour, because real timeouts are too rare to wait for.
pip install neverempty
neverempty init && neverempty run
The second command needs no API key. It runs against a stub agent so you can see the output shape before pointing it at anything of your own.
PyPI: https://pypi.org/project/neverempty/ GitHub: https://github.com/Priyank032/neverempty
Apache-2.0, Python 3.10 to 3.13, mypy --strict clean.
If you run the count on your own repo, I would like to know the number. Mine was 24, and I expect most agent codebases of a similar size are in double digits. The interesting part is not the count itself, it is that almost nobody has looked.
Top comments (1)
The prompt diff is the part I'm taking. We had almost the same "fewer than 5 results, suggest broadening" line, and during an outage it was basically telling users their query was the problem. We'd tested the happy path for weeks and the failure path for an afternoon; now we force a timeout at each tool boundary and check the model actually reads "couldn't check," not an empty list.