DEV Community

Cover image for NVIDIA measured it: coding agents scored 19% without a spec file, 100% with one
Piekwerk
Piekwerk

Posted on

NVIDIA measured it: coding agents scored 19% without a spec file, 100% with one

NVIDIA put a number on something most of us argue about anecdotally: what is a curated config file actually worth to a coding agent? On October 1 the DOCA team published agent skills on GitHub and paired the release with a 65-prompt evaluation. Same agents, same models. Without the skills, agents satisfied 19% of the graded checklist items. With the skills loaded, 100% across all 65 prompts. No fine-tuning, no new model, just a specification the agent could read at inference time.

What actually shipped

The bundle lives in the NVIDIA skills repository and is scoped per DOCA component: Flow, GPUNetIO, RDMA, PCC, and friends. Each skill is a directory built around a SKILL.md file carrying real function signatures, hardware capability requirements, build constraints, and known failure modes with mitigations. NVIDIA is explicit that these are not compressed documentation. They are machine-readable specifications the agent reasons against directly, and the contributing rules state the bundle is guidance-only, never generated code.

That last part matters. The skill does not do the work. It tells the agent what the real API surface looks like before the agent writes a line, which is exactly the job a good rules file does in any stack.

The failure taxonomy is the real story

The headline number is 19% to 100%, but the breakdown of what agents got wrong without skills is more useful, because every failure mode has a cousin in mainstream development:

  • Invented or misused APIs and flags: 59 of 65 prompts
  • Hardware capability never verified before writing code: 46 of 65
  • Wrong tool routing (right goal, wrong tool): 39 of 65
  • Skipped smoke tests: 34 of 65
  • Guessed version numbers: 30 of 65

Read that list with a webdev hat on. Invented flags become hallucinated config keys in a Next.js build. Unverified capabilities become imports from a library version you do not ship. Guessed versions become a lockfile that lies. I wrote about why agents ignore rules in this piece, and NVIDIA's taxonomy matches the failure classes I see in ordinary repos, just measured at scale on a niche SDK.

The demo: 189 lines versus 695

NVIDIA also ran a side-by-side demo, two agents building the same Go program that sends real RDMA traffic on a BlueField-3. Both succeeded. The with-skills agent needed 189 lines of handwritten code against 695 without, a 73% reduction, and issued 20 hardware commands instead of 37.

I find the line count more interesting than the checklist score. Less handwritten code means fewer correction cycles between task and working code. The without-skills agent was not failing exactly, it was rediscovering the API by trial and error, and every trial was a debugging session for the developer holding the wheel. That is the quiet cost of an under-specified agent in any domain.

Why inference-time knowledge scales

The architectural point NVIDIA makes is one worth stealing: the knowledge lives in a layer the agent reads at run time, so when the API changes, you update the skill, not the model. No retraining cycle, no waiting for the next foundation model to have memorized your framework's current quirks. The skills format follows the open agentskills.io specification, and NVIDIA's verified-skills post notes the same SKILL.md works across Claude Code, Codex, and Cursor. Portable instructions, scanned and signed before publication. That is the direction the whole config layer is moving: from personal CLAUDE.md notes toward versioned, verifiable artifacts.

The honest caveats

Vendor evaluation, vendor grading. NVIDIA wrote the 65 prompts, wrote the checklists, and ran the scoring, and the blog post itself says the 100% is checklist satisfaction, not flawless production software. DOCA is also close to a worst case for bare models: a fast-moving, hardware-bound SDK where training data is thin and stale. The effect size on a React codebase, where every model has seen a million Next.js apps, will be smaller. I would not quote 19-to-100 at your team lead as a promise. But the direction of the effect is not in doubt, and the failure modes it fixes are the ones that burn review time every week.

What transfers to your stack

You do not need BlueField hardware to use the finding. Take NVIDIA's top failure classes and write the opposite into your agent config:

## Versions (do not guess)
- Next.js 15.x App Router only; check package.json before importing.
- Never invent config keys: if unsure a key exists, read next.config.ts
  instead of assuming.

## Capability checks before code
- Verify a dependency exists in package.json AND the lockfile before
  importing it. If missing, stop and ask, do not add it silently.

## Build truth
- The build command is `pnpm build`. Do not run `npm run build`.
- After any config change, run the build once before reporting done.
Enter fullscreen mode Exit fullscreen mode

Four rules, each aimed at a measured failure class. That is usually worth more than a 2,000-line CLAUDE.md, and it fits inside the 300-line instruction budget with room to spare. NVIDIA ships an evals.json test set alongside every skill so the guidance gets graded, which is the same reason I run validators against config kits instead of trusting that they read well.

The takeaway

NVIDIA has enormous model resources and still chose to fix this problem with files, not training. When the people who can fine-tune anything say the cheapest reliable win is a verified spec the agent reads at run time, that is a strong signal for where your own effort should go: a small, current, version-pinned config beats a large, stale one. If you want a starting point, my AgentConfig Studio kits apply exactly this pattern for common stacks, or grab the free Next.js sample kit and compare it against your current rules file. Then measure your own 65 prompts. The gap might surprise you.

Top comments (0)