DEV Community

AI Tech Connect
AI Tech Connect

Posted on Originally published at aitechconnect.in

SWE-Serve Rejects a Third of Patches Other Tests Pass

Originally published on AI Tech Connect.

What you need to know The rejection rate is the finding. End-to-end serving tests reject roughly one third of the patches that pass other evaluations. That is a statement about how we measure coding agents, not about any one model. The tasks are real. 53 repository-grounded tasks derived from merged SGLang production pull requests — work that shipped, reviewed by people who maintain an inference server for a living. The ceiling exists. The best-performing model-effort configuration reached 75% mean pass@1, across 11 models tested. Agents are not hopeless at inference engineering; they are just badly measured by function-shaped tests. It is open. Reference implementation runs with Harbor; the code sits at github.com/NVIDIA/swe-serve, and the paper is arXiv 2609.26777. The benchmark, in one…


Read the full article on AI Tech Connect →

Top comments (0)