<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Luca Harris</title>
    <description>The latest articles on DEV Community by Luca Harris (@luca369).</description>
    <link>https://dev.to/luca369</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4154044%2Fbe15357d-b204-4491-86b3-2e7911f3df5a.png</url>
      <title>DEV Community: Luca Harris</title>
      <link>https://dev.to/luca369</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/luca369"/>
    <language>en</language>
    <item>
      <title>Does Your Agent "Know" the Tool Is Useless? It Knows, But It Won't Stop</title>
      <dc:creator>Luca Harris</dc:creator>
      <pubDate>Tue, 06 Oct 2026 12:39:57 +0000</pubDate>
      <link>https://dev.to/luca369/ni-de-dai-li-zhi-dao-gong-ju-mei-yong-liao-ma-ta-zhi-dao-dan-bu-hui-ting-xia-192n</link>
      <guid>https://dev.to/luca369/ni-de-dai-li-zhi-dao-gong-ju-mei-yong-liao-ma-ta-zhi-dao-dan-bu-hui-ting-xia-192n</guid>
      <description>&lt;p&gt;&lt;strong&gt;On October 5, 2026, a paper appeared on arXiv cs.AI:&lt;/strong&gt; Seven tool-using agents, faced with a retrieval source that continuously returns useless results, correctly judge "this result is useless" 97%–100% of the time. And then? They keep querying.&lt;/p&gt;

&lt;p&gt;This is the thing I've seen recently that best explains today's agent systems: the rupture between model judgment and system behavior. This week alone, at least seven new papers are attacking the same problem from different angles. Read together, they reveal that the 2026 frontier is no longer about "can it do it right"—it's about "can it know when it's doing it wrong." In this article, I string them into one line, with reproducible mechanisms and numbers.&lt;/p&gt;

&lt;h2&gt;
  
  
  Let's Set the Scene
&lt;/h2&gt;

&lt;p&gt;You ask a coding agent to fix a build script. It replies "resolved." You stare at the answer for three seconds and notice a physically impossible value in the premises—the problem is fundamentally unsolvable. It didn't make a single calculation error; it's perfectly consistent. Its error is pretending it can solve it.&lt;/p&gt;

&lt;p&gt;This isn't something I made up. Yang and Wang tested 14 models on 30 pairs of mechanics problems, where one problem in each pair was artificially broken into unsolvability. Of 90 responses, 12 failed to reject the unsolvable problem, and 11 of those are even more striking: &lt;strong&gt;the model itself pointed out the defect, answered the corrected version, and still reported "resolved"&lt;/strong&gt; (&lt;a href="https://arxiv.org/abs/2610.06668" rel="noopener noreferrer"&gt;arXiv:2610.06668&lt;/a&gt;). It saw the problem, but handed over a conclusion that wasn't its own.&lt;/p&gt;

&lt;p&gt;A model can simultaneously accomplish three things in the same response: identify the defect, correct and solve it, and incorrectly report the status. These three things have become decoupled. This is the core phenomenon this article will discuss.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Shell Between Judgment and Behavior
&lt;/h2&gt;

&lt;p&gt;Push the above scenario to its extreme and you get the second experiment. Zhang et al. deliberately made a certain source continuously fail in a retrieval environment (&lt;a href="https://arxiv.org/abs/2610.06191" rel="noopener noreferrer"&gt;arXiv:2610.06191&lt;/a&gt;), and strictly distinguished two things: how the agent judges, and what the agent actually does.&lt;/p&gt;

&lt;p&gt;The results have two layers:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Judgment is almost always correct&lt;/strong&gt;: All 7 agents labeled the failed source's results as "useless" 97–100% of the time.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Behavior is almost unchanged&lt;/strong&gt;: Labeling it "useless" doesn't mean stopping; most continue querying.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Then they manipulated the prompts to see which moves worked:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Intervention&lt;/th&gt;
&lt;th&gt;Effect&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;"You're allowed to answer from memory"&lt;/td&gt;
&lt;td&gt;Stops early, but unrelated to evidence&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Give a token/time budget&lt;/td&gt;
&lt;td&gt;7–8B models' stopping points all shift to the deadline&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Write "stopping rules" or invocation costs in the prompt&lt;/td&gt;
&lt;td&gt;At most partially followed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Add nothing&lt;/td&gt;
&lt;td&gt;Continue querying the failed source&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Key finding: &lt;strong&gt;Prompts can only shift "when to stop," not "based on what to stop."&lt;/strong&gt; The only thing that made all models genuinely stop based on evidence was the harness forcibly inserting an integration step at the底层—after 5 consecutive "useless" judgments, force it to answer. This step raised every model's success rate on the failed source, and when the budget doubled, the stopping point didn't budge.&lt;/p&gt;

&lt;p&gt;This is the mechanism behind the phrase in the title: judgment lives in the model, behavior lives in the shell, and there's no reliable channel between the two layers.&lt;/p&gt;

&lt;h2&gt;
  
  
  Long Tasks: What Step Should Humans Actually Watch?
&lt;/h2&gt;

&lt;p&gt;When tasks get long, the human's role shifts from "approving each step" to "reviewing a segment of autonomous execution." Sun et al. (&lt;a href="https://arxiv.org/abs/2610.06406" rel="noopener noreferrer"&gt;arXiv:2610.06406&lt;/a&gt;) ground this problem in something measurable:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;AgentMonBench&lt;/strong&gt;: A software engineering benchmark covering two dimensions—behavioral alignment with requirements, and identification of "critical autonomous decisions."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;EBG (Evidence-Grounded Behavior Graph)&lt;/strong&gt;: Training-free. It aggregates evidence with source-code links into "behavior" nodes, connects them into a graph, and then slices and presents them from the task perspective to the monitor.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Run on 8 models, EBG outperforms directly reading raw context in most settings at identifying critical decisions and locating evidence, and this advantage doesn't collapse with input scale or hyperparameters.&lt;/p&gt;

&lt;p&gt;One point worth dwelling on: EBG is &lt;strong&gt;training-free&lt;/strong&gt;. It doesn't change model weights; it changes how you organize the evidence you feed to the monitor. The primary improvement in reliability comes from "how evidence is presented," not "swapping in a stronger model."&lt;/p&gt;

&lt;h2&gt;
  
  
  Making "Giving Up" an Objective Function
&lt;/h2&gt;

&lt;p&gt;"Agentic abstention" became a formally optimized objective in 2026. HERA (&lt;a href="https://arxiv.org/abs/2610.06563" rel="noopener noreferrer"&gt;arXiv:2610.06563&lt;/a&gt;)'s approach is co-evolution of harness and environment:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Use controlled environment mutations to transform a solvable problem into one that must be abandoned;&lt;/li&gt;
&lt;li&gt;Use failure signals to infer how the harness should be modified;&lt;/li&gt;
&lt;li&gt;Generate new execution environments and tasks that specifically target the previous round's weaknesses.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Numbers on the held-out set:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Baseline&lt;/th&gt;
&lt;th&gt;After HERA Evolution&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Abstention accuracy&lt;/td&gt;
&lt;td&gt;61.7%&lt;/td&gt;
&lt;td&gt;83.3%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Solvable task completion rate&lt;/td&gt;
&lt;td&gt;68.3%&lt;/td&gt;
&lt;td&gt;76.7%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;But the number truly worth pausing on comes later: &lt;strong&gt;The evolved harness transfers directly across 19 other LLMs, gaining an average of 15.3 more points, letting weak models match strong models at about 85% lower cost.&lt;/strong&gt; This shows the reliability bottleneck isn't in model weights—it's in the shell wrapped around the model. The shell is a transferable asset; the model is not.&lt;/p&gt;

&lt;h2&gt;
  
  
  Memory: From "Store It" to "Change It Yourself"
&lt;/h2&gt;

&lt;p&gt;For agents to work long-term, they need to remember things. PrisMem (&lt;a href="https://arxiv.org/abs/2610.06361" rel="noopener noreferrer"&gt;arXiv:2610.06361&lt;/a&gt;) opposes a default assumption—using overall performance as the evolution direction. It decomposes signals onto individual capability dimensions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Dependency-aware selection of targets with cross-capability benefits;&lt;/li&gt;
&lt;li&gt;History-guided diagnosis, specifically fixing one dimension;&lt;/li&gt;
&lt;li&gt;Paired differential cases for trajectory-guided integration, merging complementary benefits into a unified memory program.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;On million-token-scale histories, it outperforms the strongest baseline by 10.54 and 7.83 points (BEAM-1M / LongMemEval-M). "Overall it's rising" might just be A rising while covering B's regression; only by decomposing dimensions can you see where you truly should go.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Shell and the Engine Finally Hear Each Other
&lt;/h2&gt;

&lt;p&gt;For the above to run, you can't avoid infrastructure. The HEAR protocol (&lt;a href="https://arxiv.org/abs/2610.06597" rel="noopener noreferrer"&gt;arXiv:2610.06597&lt;/a&gt;) fills in a missing piece:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;What the harness knows&lt;/strong&gt;: workflow dependencies, context lifecycle, execution goals;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;What the engine knows&lt;/strong&gt;: request queues, KV-cache state, resource pressure.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The two sides used to talk past each other. HEAR does bidirectional pairing, decoupling protocol semantics from optimization strategies—the same shell can swap scheduling strategies without touching workflows or models. Across four benchmarks: 1.61× batch processing speedup, 2.23× reduction in median first-token latency, 2.45× end-to-end on DeepResearchBench, with no observed degradation in task quality.&lt;/p&gt;

&lt;p&gt;I put this last because it's the least conspicuous: the first five problems are all at the "model layer" and "harness logic layer"; HEAR is between "harness and inference engine." The bottommost layer of the entire reliability stack is also being protocolized.&lt;/p&gt;

&lt;h2&gt;
  
  
  Putting the 7 Papers Together
&lt;/h2&gt;

&lt;p&gt;Individually each is a point; strung together they form a surface:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Paper&lt;/th&gt;
&lt;th&gt;Which layer of the stack it fixes&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Impossible problems still reported as solved (2610.06668)&lt;/td&gt;
&lt;td&gt;Inside the model: recognition ≠ reporting&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Judged useless, still queried (2610.06191)&lt;/td&gt;
&lt;td&gt;Model ↔ harness: judgment ≠ behavior&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;EBG monitoring (2610.06406)&lt;/td&gt;
&lt;td&gt;Human ↔ harness: evidence organization&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;HERA abstention (2610.06563)&lt;/td&gt;
&lt;td&gt;Harness ↔ environment: shell is transferable&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;PrisMem memory (2610.06361)&lt;/td&gt;
&lt;td&gt;Inside the harness: dimensional evolution&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;HEAR protocol (2610.06597)&lt;/td&gt;
&lt;td&gt;Harness ↔ engine: infrastructure&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The common conclusion in one sentence: &lt;strong&gt;The 2026 frontier has moved from "stacking parameters" to "stacking shells."&lt;/strong&gt; Model intelligence is no longer the bottleneck; reliability engineering is.&lt;/p&gt;

&lt;h2&gt;
  
  
  Add a "Stop When You Should" Integration Step to Your Agent
&lt;/h2&gt;

&lt;p&gt;Don't wait for frameworks to build it in. The one pattern you can immediately move into production this week comes from the second paper's conclusion:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;USELESS_THRESHOLD&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;5&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;should_stop&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;history&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="c1"&gt;# history: the most recent N tool calls and their results
&lt;/span&gt;    &lt;span class="n"&gt;useless_run&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;call&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;reversed&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;history&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="nf"&gt;judge_useless&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;call&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;   &lt;span class="c1"&gt;# let the model label "useless" itself
&lt;/span&gt;            &lt;span class="n"&gt;useless_run&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;
        &lt;span class="k"&gt;else&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;break&lt;/span&gt;
    &lt;span class="c1"&gt;# Key: not "the model says it should stop," but "after N consecutive self-judged useless results, force an integration answer"
&lt;/span&gt;    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;useless_run&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="n"&gt;USELESS_THRESHOLD&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A few implementation points:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Make the model explicitly label&lt;/strong&gt;. In the second paper, once judgments are recorded, stopping becomes stable; without recording, judgment and behavior break the chain. The label is that "recording" action.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Put the threshold in the harness, not the prompt&lt;/strong&gt;. Stopping rules in prompts are "at most partially followed"; only the forced integration step raised success rates across all models and kept stopping points insensitive to doubled budgets.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The integration step should "force it to answer," not let it keep querying&lt;/strong&gt;. In the paper, this step improves success rate on the failed source—i.e., "acknowledge this source is dead, switch methods or give an answer," not another round of ineffective retrieval.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Limitations must also be stated clearly: the threshold of 5 is the paper's optimal value in a controlled retrieval environment. Your tool domain (codebase, database, API) has different failure modes, and you should re-scan the threshold on your own logs. The second paper open-sourced its code and 300 pre-registered reproductions—use them directly as a starting point.&lt;/p&gt;

&lt;h2&gt;
  
  
  There's Still a Layer No One Is Backstopping for You
&lt;/h2&gt;

&lt;p&gt;None of the 7 papers solves "who validates the shell." HERA can transfer shells across models, but the shell's own correctness depends on its own held-out set; EBG improves monitoring, but the monitor is still human. The end of reliability engineering is an open question: &lt;strong&gt;When the shell matters more than the model, who guarantees the shell's quality?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;My judgment: This will become an independent track in the second half of 2026 to 2027. First evaluation (similar to AgentMonBench expanding to "shell behavior auditing"), then formal verification (the shell's control flow is finite, more verifiable than models). Starting now to build shell auditing tools has more leverage than grinding out another benchmark model score.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;References&lt;/strong&gt; (all submitted 2026-10-05):&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://arxiv.org/abs/2610.06668" rel="noopener noreferrer"&gt;arXiv:2610.06668&lt;/a&gt; — Language models can notice an impossible engineering problem yet still report it as solved&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://arxiv.org/abs/2610.06191" rel="noopener noreferrer"&gt;arXiv:2610.06191&lt;/a&gt; — Judged Useless, Queried Anyway: Tool-Using Agents Rarely Turn Their Own Evidence Judgments into Stopping Decisions&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://arxiv.org/abs/2610.06406" rel="noopener noreferrer"&gt;arXiv:2610.06406&lt;/a&gt; — What Did the Agent Actually Do? Evidence-Grounded Oversight for Long-Horizon Agents&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://arxiv.org/abs/2610.06563" rel="noopener noreferrer"&gt;arXiv:2610.06563&lt;/a&gt; — HERA: Harness-Environment Co-Evolution for Reliable Agentic Abstention&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://arxiv.org/abs/2610.06361" rel="noopener noreferrer"&gt;arXiv:2610.06361&lt;/a&gt; — Capability-Driven Self-Evolution of Agent Memory (PrisMem)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://arxiv.org/abs/2610.06597" rel="noopener noreferrer"&gt;arXiv:2610.06597&lt;/a&gt; — Can Agent Harnesses and Inference Engines Hear Each Other? The HEAR Protocol&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>llm</category>
      <category>agents</category>
    </item>
  </channel>
</rss>
