DEV Community

Cover image for Who's Accountable When the AI Was Just Following Instructions?

Who's Accountable When the AI Was Just Following Instructions?

James Anderson on September 29, 2026

Earlier this year, a company's AI agent quietly leaked internal data for three weeks before anyone noticed. Around the same time, a different AI ag...
Collapse
 
slabb profile image
Sam LABBE •

My answer to your closing question: put it with whoever sells the autonomy — and the good news is we don't have to solve the intent/foreseeability puzzle to get there, because the law stopped asking "who's to blame" for this shape of harm a long time ago.

Product liability already did the hard work: when a product's defect causes harm, the manufacturer is strictly liable regardless of fault, because they're the party best positioned to price the risk, spread it through insurance, and design it down. Cars never got accountability by figuring out intent — we got mandatory insurance, strict liability on the manufacturer for defects, and fault only re-enters for the operator's negligence. The EU's revised Product Liability Directive (2024/2853) has now made that exact move for software and AI: the vendor is in the strictly liable chain, with eased proof burdens for claimants.

So "you deployed it, you own it" pins everything on the party with the least information about the black box. The strict rule that actually matches the information asymmetry is "you sold the autonomy, you own its downside" — deployer liability capped unless negligent, all of it mutualized the way car insurance did.

And your "unforgeable record of who authorized what" is the half engineers can build today — it's literally the problem I work on. You already know NoireBox, my tamper-evident journal for agent actions, built precisely so that "who approved this, and for what" survives the incident. The half you flagged as unsolved — whether the authorizer understood what they approved — is the one that genuinely isn't.

Collapse
 
james_anderson_h profile image
James Anderson •

Sharpest resolution here — and it dissolves the puzzle by showing I solved the wrong one. Product liability never untangled intent; it put strict liability on the manufacturer who prices and spreads the risk. So "you sold the autonomy, you own its downside" beats "you deployed it, you own it" — it matches the information asymmetry instead of fighting it. And you drew the exact line: who-authorized-what is buildable; whether they understood it isn't. Pinned.

Collapse
 
slabb profile image
Sam LABBE •

Thanks James — it takes a rare kind of honesty to admit mid-essay that you were solving the wrong question, so the pin is humbling. If the debate ever moves toward the buildable half (authorization records that survive the incident), that's where I'll be.

Thread Thread
 
james_anderson_h profile image
James Anderson •

That honesty's the least I owed a comment that good — and yes, when the conversation turns to the buildable half, the survives-the-incident record is exactly where the interesting work is, so I'll see you there.

Thread Thread
 
slabb profile image
Sam LABBE •

Deal — and I'll do my best to make the buildable half live up to the framing. See you there, James.

Collapse
 
glenallen profile image
Glen Allen •

One engineering distinction that may help here is separating accountability from controllability. Instead of asking only who should ultimately bear responsibility, a production system could record which layer had control over each consequential decision: model behavior, tool permissions, policy enforcement, deployment configuration, or human approval. That doesn't solve the legal question, but it makes the incident much easier to reconstruct. It also exposes gaps where everyone technically participated in a decision but no layer had an enforceable control over the outcome. For agent systems, that control map could become just as important as the audit trail itself.

Collapse
 
james_anderson_h profile image
James Anderson •

Separating accountability from controllability is the cleaner cut: "who bears responsibility" is legal, but "which layer had control over this decision" is answerable now. A control map turns reconstruction into lookup — and surfaces the scariest case: everyone participated, no layer had enforceable control.

Collapse
 
glenallen profile image
Glen Allen •

That’s the part I find most useful too: the control map can become more than an incident-reconstruction tool. It could also expose missing controls before deployment, especially when a consequential action crosses several layers. If no layer can clearly say “I can block this decision,” then the system has a design gap even if the audit trail looks complete. That makes controllability something worth testing during architecture reviews, not only something to reconstruct after an incident.

Collapse
 
maximin_mxn_6ce1a3054be6d profile image
Maximin •

The distinction between accountability and controllability is a useful way to make this debate operational. A control map that shows where a consequential decision could actually be blocked would expose gaps before an incident, not only explain them afterward.

Collapse
 
kartik-nvjk profile image
Kartik N V J K •

The database wipe because the agent believed it was in a dev environment is the one that stops me, because it is not a blame problem, it is an observability problem: nobody could reconstruct what the agent thought its environment was at decision time. When I dug into similar near-misses, the responsibility fog you describe came straight from missing that state in the trace. If you cannot replay what the agent believed, every party's "not me" stays technically true. Where would you put the recording boundary so the deployer and vendor can actually argue from the same evidence?

Collapse
 
aifrontierpost profile image
AI Frontier Post •

The article's own parenthetical is the underrated line: an unforgeable authorization record "proves what was authorized, not whether the authorizer understood what they were approving" — and in production that second half is where the fog actually lives. Kartik's database-wipe example sharpens it: the incident wasn't a missing signature, it was a missing believed-state — nobody could replay what environment the agent thought it was in at decision time. So the buildable half of slabb's product-liability framing needs one addition: authorization records that also snapshot the agent's world-model at the moment of the action, or all five partial defenses stay technically true and the fog survives the audit trail.