<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Raja Rajak</title>
    <description>The latest articles on DEV Community by Raja Rajak (@raja_rajak).</description>
    <link>https://dev.to/raja_rajak</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4174746%2F64c02006-2ba5-4cef-9952-76bcdd619c79.jpg</url>
      <title>DEV Community: Raja Rajak</title>
      <link>https://dev.to/raja_rajak</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/raja_rajak"/>
    <language>en</language>
    <item>
      <title>Beyond Leaderboard Illusions: Benchmarking Multi-Turn Agentic Feedback Loops in Autonomous Software Engineering with TFD-Bench</title>
      <dc:creator>Raja Rajak</dc:creator>
      <pubDate>Sat, 10 Oct 2026 06:58:48 +0000</pubDate>
      <link>https://dev.to/raja_rajak/beyond-leaderboard-illusions-benchmarking-multi-turn-agentic-feedback-loops-in-autonomous-software-40fj</link>
      <guid>https://dev.to/raja_rajak/beyond-leaderboard-illusions-benchmarking-multi-turn-agentic-feedback-loops-in-autonomous-software-40fj</guid>
      <description>&lt;h1&gt;
  
  
  Beyond Leaderboard Illusions: Benchmarking Multi-Turn Agentic Feedback Loops in Autonomous Software Engineering with TFD-Bench
&lt;/h1&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Submission for the Kaggle Benchmarking Challenge on DEV&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
&lt;em&gt;Author:&lt;/em&gt; Raja Rajak (&lt;a href="https://dev.to/rajrajak99"&gt;@rajrajak99&lt;/a&gt;)&lt;br&gt;&lt;br&gt;
&lt;em&gt;Kaggle Benchmark Dataset:&lt;/em&gt; &lt;a href="https://www.kaggle.com/datasets/rajrajak99/gemma4-tfd-agentic-trajectories" rel="noopener noreferrer"&gt;Gemma 4 TFD Agentic Trajectories&lt;/a&gt;&lt;br&gt;&lt;br&gt;
&lt;em&gt;Kaggle Evaluation Notebook:&lt;/em&gt; &lt;a href="https://www.kaggle.com/code/rajrajak99/tfd-bench-multi-turn-agentic-swe-benchmark-eval" rel="noopener noreferrer"&gt;TFD-Bench on Kaggle&lt;/a&gt;&lt;br&gt;&lt;br&gt;
&lt;em&gt;GitHub Repository:&lt;/em&gt; &lt;a href="https://github.com/rajrajak99/gemma4-tfd-agent" rel="noopener noreferrer"&gt;github.com/rajrajak99/gemma4-tfd-agent&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  ⚡ Executive Summary &amp;amp; TL;DR
&lt;/h2&gt;

&lt;p&gt;Standard AI coding benchmarks—like HumanEval, MBPP, and static LeetCode-style puzzles—suffer from a fatal blindspot: &lt;strong&gt;they test one-shot code synthesis in isolation, completely ignoring how software engineering actually happens in the real world&lt;/strong&gt;. Real software development is not a single prompt-response transaction; it is &lt;strong&gt;stateful, iterative, hypothesis-driven, and governed by test feedback&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;When deployed inside live repositories, state-of-the-art LLMs suffer from what we call the &lt;strong&gt;"Illusion of One-Shot Competence"&lt;/strong&gt;: they generate elegant, syntactically clean code patches that &lt;strong&gt;fail 64.2% of real-world regression test suites&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;To measure what actually matters, we designed &lt;strong&gt;TFD-Bench (Test-Feedback-Driven Benchmark)&lt;/strong&gt;: an evaluation suite of &lt;strong&gt;50 multi-turn debugging tasks across 10 Python error classes&lt;/strong&gt; adapted from SWE-bench Lite. We benchmarked 5 frontier model configurations across &lt;strong&gt;Pass@1 Resolve Rate&lt;/strong&gt;, &lt;strong&gt;AST Tool Calling Validity&lt;/strong&gt;, &lt;strong&gt;Context Token Consumption&lt;/strong&gt;, and &lt;strong&gt;Loop Stalling Frequency&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Our key finding? &lt;strong&gt;Closed-loop test execution feedback acts as a massive reasoning accelerator&lt;/strong&gt;: enforcing autonomous test reproduction before code generation jumps issue resolution from &lt;strong&gt;29.5% to 45.6%&lt;/strong&gt; while &lt;strong&gt;slashing context token waste by 49.5%&lt;/strong&gt;.&lt;/p&gt;




&lt;h2&gt;
  
  
  🎯 What task(s) did you run?
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. The Core Problem with Static Coding Benchmarks
&lt;/h3&gt;

&lt;p&gt;Current LLM evaluation often relies on single-turn benchmarks where an LLM is given a docstring and generates a self-contained function. This fails to evaluate four indispensable developer capabilities:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Codebase Exploration:&lt;/strong&gt; Can the model navigate multi-file repositories using symbol search and selective reading without exceeding context limits?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Defect Reproduction:&lt;/strong&gt; Can the model synthesize an isolated, minimal test script that independently confirms the bug (&lt;code&gt;exit_code != 0&lt;/code&gt;) before touching existing code?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Surgical Patching:&lt;/strong&gt; Can the model apply targeted line replacements rather than rewriting entire files?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Automated Verification:&lt;/strong&gt; Can the model execute its reproducer, parse terminal tracebacks, and self-correct if the test fails?&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  2. The TFD-Bench Protocol
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;TFD-Bench&lt;/strong&gt; tests models across a strict, 5-phase closed-loop cycle:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Phase 1: Discovery:&lt;/strong&gt; Codebase localization via &lt;code&gt;search_code&lt;/code&gt; and &lt;code&gt;view_file&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Phase 2: Test-First:&lt;/strong&gt; Minimal standalone reproduction script synthesis (&lt;code&gt;create_reproducer&lt;/code&gt;) asserting pre-fix failure (&lt;code&gt;exit_code != 0&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Phase 3: Reasoning &amp;amp; Repair:&lt;/strong&gt; Surgical patch generation targeting minimal line changes (&lt;code&gt;edit_file_replace&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Phase 4: Verification:&lt;/strong&gt; Sandbox execution re-running the reproducer to assert resolution (&lt;code&gt;exit_code == 0&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Phase 5: Self-Correction Loop:&lt;/strong&gt; If verification fails, traceback parsing and automated retry (max 3 cycles).&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  3. Benchmark Task Taxonomy (50 Real-World Repository Tasks)
&lt;/h3&gt;

&lt;p&gt;The dataset contains 50 gold-standard, repository-grade tasks evenly distributed across 10 Python bug categories (5 curated tasks per class):&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Bug Category&lt;/th&gt;
&lt;th&gt;Real-World Scenario&lt;/th&gt;
&lt;th&gt;Primary Failure Mode in LLMs&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;ZeroDivisionError&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Metrics / Normalization calculations when denominator is 0&lt;/td&gt;
&lt;td&gt;Missed edge-case guard clauses&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;IndexError&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Stream parsing / Token chunking on empty sequences&lt;/td&gt;
&lt;td&gt;Out-of-bounds slicing assumptions&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;KeyError&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Nested JSON schema validation &amp;amp; configuration parsing&lt;/td&gt;
&lt;td&gt;Unhandled missing dictionary keys&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;TypeError&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Dynamic type conversions &amp;amp; nullable payload inputs&lt;/td&gt;
&lt;td&gt;String/None concatenation bugs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;AttributeError&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Method calls on optional/uninitialized object instances&lt;/td&gt;
&lt;td&gt;NoneType member access&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;FileNotFoundError&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Relative path resolution across different execution roots&lt;/td&gt;
&lt;td&gt;Working directory path drift&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;ValueError&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Parsing non-standard string formats / timestamps&lt;/td&gt;
&lt;td&gt;Unchecked input validation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;RecursionError&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Circular tree traversals &amp;amp; recursive graph resolvers&lt;/td&gt;
&lt;td&gt;Missing base termination conditions&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;BoundaryCondition&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Off-by-one pagination &amp;amp; sliding window edge cases&lt;/td&gt;
&lt;td&gt;Inclusive vs exclusive range mistakes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;ResourceLeak&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Unclosed socket/file descriptors in exception blocks&lt;/td&gt;
&lt;td&gt;Missing context managers (&lt;code&gt;with&lt;/code&gt; statements)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h2&gt;
  
  
  🤖 Which models did you run it against?
&lt;/h2&gt;

&lt;p&gt;To rigorously assess open vs. closed models and the impact of specialized fine-tuning, we evaluated 5 representative models across different scales and paradigms:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;TFD-Agent &lt;a href="https://dev.toOur%20System"&gt;Google Gemma 4 31B + QLoRA&lt;/a&gt;:&lt;/strong&gt;
Google's Gemma 4 31B fine-tuned with 4-bit QLoRA targeting all linear projection layers (&lt;code&gt;q, k, v, o, gate, up, down_proj&lt;/code&gt;) specifically conditioned on the Test-Feedback-Driven protocol.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Google Gemma 4 31B (Standard ReAct):&lt;/strong&gt;
The identical base weights of Gemma 4 31B deployed in a vanilla ReAct prompt loop without TFD conditioning.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Meta Llama 3.1 70B Instruct (SWE-agent Framework):&lt;/strong&gt;
Frontier open-weights generalist model evaluated via the established Princeton SWE-agent bash interface.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Qwen 2.5 Coder 32B Instruct (ReAct):&lt;/strong&gt;
A state-of-the-art open code-specialized model known for high HumanEval scores.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;OpenAI GPT-4o mini (ReAct Baseline):&lt;/strong&gt;
Proprietary frontier lightweight model representing commercial API agent backends.&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  💡 What are the main insights?
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. The Benchmark Leaderboard
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model Architecture&lt;/th&gt;
&lt;th&gt;Parameter Scale&lt;/th&gt;
&lt;th&gt;Pass@1 Resolve Rate (%)&lt;/th&gt;
&lt;th&gt;Tool Syntax Validity (%)&lt;/th&gt;
&lt;th&gt;Mean Tokens / Solved Task&lt;/th&gt;
&lt;th&gt;Loop Stalling Rate (%)&lt;/th&gt;
&lt;th&gt;Avg Turns to Fix&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;TFD-Agent &lt;a href="https://dev.toOurs"&gt;Gemma 4 31B&lt;/a&gt;&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;31B&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;45.6%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;96.9%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;19,400&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;2.1%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;7.2&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Llama 3.1 70B (SWE-agent)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;70B&lt;/td&gt;
&lt;td&gt;42.4%&lt;/td&gt;
&lt;td&gt;90.2%&lt;/td&gt;
&lt;td&gt;32,100&lt;/td&gt;
&lt;td&gt;8.5%&lt;/td&gt;
&lt;td&gt;9.8&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Qwen 2.5 Coder 32B (ReAct)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;32B&lt;/td&gt;
&lt;td&gt;41.0%&lt;/td&gt;
&lt;td&gt;89.8%&lt;/td&gt;
&lt;td&gt;29,800&lt;/td&gt;
&lt;td&gt;9.6%&lt;/td&gt;
&lt;td&gt;10.5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;GPT-4o mini (ReAct)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Frontier API&lt;/td&gt;
&lt;td&gt;38.2%&lt;/td&gt;
&lt;td&gt;87.6%&lt;/td&gt;
&lt;td&gt;28,400&lt;/td&gt;
&lt;td&gt;11.4%&lt;/td&gt;
&lt;td&gt;11.2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Gemma 4 31B (Vanilla ReAct)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;31B&lt;/td&gt;
&lt;td&gt;29.5%&lt;/td&gt;
&lt;td&gt;81.5%&lt;/td&gt;
&lt;td&gt;38,400&lt;/td&gt;
&lt;td&gt;18.2%&lt;/td&gt;
&lt;td&gt;14.8&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h3&gt;
  
  
  2. Discovery #1: The "Illusion of One-Shot Competence" (Silent Regressions)
&lt;/h3&gt;

&lt;p&gt;When models operate without mandatory test reproduction, they exhibit an alarming failure mode:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;In 64.2% of failed attempts&lt;/strong&gt;, the vanilla models generated a patch that looked completely correct to human reviewers at a glance, but either failed to handle zero-length inputs, introduced a secondary exception, or broke downstream calling functions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;By forcing the model to write a reproducer FIRST (&lt;code&gt;assert func(...) == ...&lt;/code&gt;) and confirm it fails with &lt;code&gt;exit_code != 0&lt;/code&gt;&lt;/strong&gt;, the agent grounds its reasoning in execution reality. This single procedural constraint boosted solve rates by &lt;strong&gt;+16.1 percentage points&lt;/strong&gt; across the board.&lt;/li&gt;
&lt;/ul&gt;




&lt;h3&gt;
  
  
  3. Discovery #2: AST Tool Syntax Collapse Under Turn Drift
&lt;/h3&gt;

&lt;p&gt;One of the most surprising findings was how tool invocation quality degrades over multi-turn conversations:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;In turns 1 through 4, models achieve ~92% valid JSON tool calls.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Past turn 7, vanilla models suffer a 3.8x spike in tool syntax errors&lt;/strong&gt; (malformed JSON strings, trailing commas, hallucinated tool parameters like &lt;code&gt;path&lt;/code&gt; instead of &lt;code&gt;file_path&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;Once an error occurs, vanilla models often enter &lt;strong&gt;"hallucination cascades"&lt;/strong&gt;, repeating invalid tool arguments indefinitely.&lt;/li&gt;
&lt;li&gt;Fine-tuning with turn-level supervision (TFD-Agent) kept tool validity above &lt;strong&gt;96.9%&lt;/strong&gt;, eliminating catastrophic loop stalling (dropping from 18.2% to 2.1%).&lt;/li&gt;
&lt;/ul&gt;




&lt;h3&gt;
  
  
  4. Discovery #3: Test Feedback as a Reasoning &amp;amp; Token Compression Engine
&lt;/h3&gt;

&lt;p&gt;Many developers assume adding automated test runs increases inference costs. &lt;strong&gt;Our benchmark proved the exact opposite:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Standard ReAct Loop:&lt;/strong&gt; 38,400 tokens per solved task (model wanders aimlessly searching files and trying random variations).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;TFD-Agent Closed Loop:&lt;/strong&gt; &lt;strong&gt;19,400 tokens per solved task (a 49.5% context reduction)&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Why?&lt;/strong&gt; A concrete traceback with &lt;code&gt;exit_code: 1&lt;/code&gt; and line numbers acts as a &lt;strong&gt;semantic anchor&lt;/strong&gt;. The model doesn't need to guess where the fault is; the Python runtime tells it directly, pruning redundant exploration paths.&lt;/li&gt;
&lt;/ul&gt;




&lt;h3&gt;
  
  
  5. Discovery #4: Fine-Tuning Beats Scale in Agentic Loops
&lt;/h3&gt;

&lt;p&gt;Notice that &lt;strong&gt;TFD-Agent (31B)&lt;/strong&gt; outperformed &lt;strong&gt;Llama 3.1 70B&lt;/strong&gt; (45.6% vs 42.4%) despite having less than half the parameters:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;General-purpose 70B models have superior world knowledge, but they are not conditioned to adhere to the strict protocol of &lt;em&gt;Locate -&amp;gt; Reproduce -&amp;gt; Surgical Edit -&amp;gt; Verify&lt;/em&gt;.&lt;/li&gt;
&lt;li&gt;When a 31B model is aligned to treat terminal output as ground-truth feedback, it achieves state-of-the-art software engineering competence on consumer/edge workstations.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  🔍 Case Study: Anatomy of an Autonomous Fix
&lt;/h2&gt;

&lt;p&gt;Here is a real example from task &lt;code&gt;tfd_zerodivisionerror_1&lt;/code&gt; demonstrating the power of the loop:&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 1: Locating the Fault
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;
json
{"tool": "search_code", "arguments": {"query": "def calculate_precision_recall"}}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>devchallenge</category>
      <category>kagglechallenge</category>
    </item>
  </channel>
</rss>
