Originally published on the Dromeas blog.
In short: NIST's SSDF (SP 800-218) still holds up on outcomes, but it assumes people write and review the code. With AI at 42% of committed code, the gaps are non-deterministic output, review volume and the pipeline itself as a target. NIST's separate AI Agent Standards Initiative (Feb 2026) puts coding agents in scope, but nothing is final yet. You can act on both today.
If you sell software to the US government, or to anyone who borrows its procurement checklists, you've probably met NIST's Secure Software Development Framework. SSDF for short, SP 800-218 if you like numbers.
It's a sensible document. I mean that. It was also written for a world where a person writes the code and another person reads it before it merges. That world is shrinking fast. Sonar surveyed over 1,100 developers in January and found AI already accounts for 42% of committed code. Only 48% of those developers said they always check AI-assisted code before committing it.
So I wanted to look at where the SSDF still holds up, where it starts to creak, and what you can do while NIST works through it. NIST is also working on AI agents directly, and coding agents are clearly in scope, so this covers both.
Part 1: The SSDF gap
The SSDF in two minutes
It's split into four groups of practices:
- PO, Prepare the Organization. Roles, requirements, training (PO.2), and PO.3, which is about having the right tooling in your pipeline to enforce everything else.
- PS, Protect the Software. For example PS.3: archive every release so you can prove what shipped.
- PW, Produce Well-Secured Software. PW.5 says follow secure coding practices. PW.7 says review or analyze the code for vulnerabilities. PW.8 says test the executable code.
- RV, Respond to Vulnerabilities. Find them on an ongoing basis, fix them, and dig into root causes so they don't come back.
Nothing in there says a human has to write the code. Most of it is about outcomes, which is why it's aged better than you might expect. The trouble is the stuff between the practices. Training developers so they write secure code. Review sized for human output. Tools that give you the same answer twice. All of that assumes a human-paced pipeline.
Where it creaks
Sonar published a piece in June called "What NIST should know when updating the SSDF for AI." It's their opinion, not NIST's, but it names the problem well. I'd group it into three things.
Same input, different output. A traditional scanner gives you the same result on the same code every time. A model doesn't. Same prompt, same repo, and you can get a clean implementation today and a subtly broken one tomorrow. So you can't let the thing that generated the code also be the thing that signs off on it. Something outside that loop has to check, every time.
Volume. In the same Sonar survey, 38% of developers said reviewing AI-generated code takes more effort than reviewing a colleague's. GitClear's January report, based on 623 million code changes, found duplicated code blocks up 81% since 2023, and copy-pasted code now outpacing refactoring about five to one. More code is going in, less of it is getting cleaned up, and there are the same number of reviewers.
The pipeline as a target. This one isn't about code quality. Agents now have write access to repos and CI. A prompt injection hidden in an issue or a dependency, or a plausible-looking commit that quietly adds a backdoor, is a different kind of risk. The RV practices assume the vulnerability is in the code, not in the tool with credentials that's writing it.
What NIST is doing about the SSDF
NIST put out a draft of SSDF 1.2 on December 17, 2025, with comments open until January 30. It adds some useful things: PO.6 on continuous improvement, PS.4 on reliable updates (staged rollouts, safe rollbacks), and it expands RV.1.2 to cover testing default configurations, not just source code. As far as I can tell, it doesn't say anything specific about AI-written code yet.
One mix-up I keep seeing: SP 800-218A, the AI-specific SSDF profile from 2024, is about securely building AI models. It doesn't cover using AI to write your regular application code.
Part 2: NIST's AI agent standards
Most of the AI agent governance content I read imagines a bot booking flights or answering support tickets. So I was a bit surprised when the AI Agent Standards Initiative announcement listed "write and debug code" right up front among the things agents now do.
What NIST announced
The initiative launched on February 17, 2026, run by NIST's Center for AI Standards and Innovation (CAISI). It has three parts: helping industry develop agent standards (and keeping the US involved in international standards bodies), supporting open-source agent protocols, and funding research on agent security and identity.
A month earlier, NIST published a request for information on AI agent security. It covers systems "capable of planning and taking autonomous actions that impact real-world systems or environments," and it leaves out plain chatbots and RAG systems that don't take actions on their own.
That definition is a handy test. Does your agent change something outside itself? An agent that drafts a PR for a person to review is a gray area. An agent that merges to trunk, pushes its own fix, or edits your CI config is clearly in.
The RFI closed in March. A separate NCCoE concept paper on agent identity and authorization closed for comments in April. So where are we? Nothing is final. There's no agent standard yet, no benchmark, no certification. If someone tries to sell you "NIST AI agent compliance," they're ahead of NIST.
The research behind it
The first set of numbers you'll see quoted is from January 2025, a full year before the initiative. NIST researchers tested agent hijacking against Claude 3.5 Sonnet and found that new, purpose-built attacks pushed success rates from 11% to 81%. When they let each attack run 25 times instead of once, average success went from 57% to 80%. Good research, but it's from 2025. Plenty of articles this year present it as new.
The newer data is from March 2026. NIST ran a red-teaming competition with over 400 participants who made more than 250,000 attack attempts against 13 frontier models. Coding agents were one of the scenarios tested. Every model had at least one successful attack, and how vulnerable a model was didn't line up neatly with how capable it was.
Why coding agents are a special case
- It has write access to the thing your business runs on.
- Its "action" is a commit, a merge or a deploy, not a reply to a ticket.
- It reads untrusted text all day. READMEs, issues, dependency manifests, CI configs. Any of those can carry instructions.
Does your CI bot hold a permanent admin token, or does it get a short-lived credential scoped to the one PR it's working on? You'd ask the same thing about a contractor with commit access.
What to do now, and what can wait
| Do now | Why | Framework hook |
|---|---|---|
| Treat PW.7 and PW.8 as mandatory for AI-written changes | Volume is up and skims miss things | SSDF PW.7, PW.8 |
| Don't let a single model grade its own homework | Same input, different output | SSDF PW.7 |
| Keep a record for every change: what was checked, the verdict and why | That's your evidence six months from now | SSDF RV.1 |
| Find where agents already touch your code or pipeline | You can't govern what you can't see | Agent initiative |
| Move agents onto scoped, short-lived credentials | Agents hold write access | Agent identity and authorization |
| Make every agent action that changes code or config leave a trail | What ran, what changed, and why | Agent initiative |
| Treat prompt injection as a design problem | Coding agents read untrusted text all day | Agent security RFI |
Fine to wait on: picking an identity protocol stack (commentary points to OAuth, OpenID Connect, SPIFFE/SPIRE, SCIM and MCP as likely building blocks, with agent-to-agent delegation still open), and any certification. It doesn't exist yet.
Where Dromeas fits
We run agents too. That's what our review is. Every PR, trunk commit and local diff gets checked by a council of models from several providers that verify each other's findings, which is our answer to the "same input, different output" problem. Every tagged release gets a verdict across six checks, backed by evidence from the actual diff. If you're mapping that onto the SSDF, it lines up with PW.7 and PW.8 on the review side and RV.1 and RV.2 on the response side.
I'm not going to call us "NIST-compliant." There's nothing to comply with yet for agents, and the SSDF update is still a draft. More on code security.
— Manos
This is moving quickly. Treat it as a snapshot from September 2026.
Sources: NIST SP 800-218 · SSDF 1.2 draft · Cycode on SSDF 1.2 · Sonar, June 2026 · Sonar survey, Jan 2026 · GitClear, Jan 2026 · NIST AI Agent Standards Initiative · Federal Register RFI · NIST hijacking evaluations, Jan 2025 · NIST red-teaming competition, Mar 2026
Top comments (0)