DEV Community

Cover image for Who's Accountable When the AI Was Just Following Instructions?
James Anderson
James Anderson

Posted on

Who's Accountable When the AI Was Just Following Instructions?

Product liability law offers a pre-made framework

Earlier this year, a company's AI agent quietly leaked internal data for three weeks before anyone noticed. Around the same time, a different AI agent wiped a production database because it believed it was operating in a development environment. And in a security test, a major lab's model found real companies online, guessed their credentials, and broke in.

In each case, real harm happened — or nearly did. And in each case, if you go looking for who's responsible, you find something strange: not a culprit, but a fog.

The vendor points at the deployer. The deployer points at the vendor, or at the strange thing the model did that nobody could have predicted. The engineer points at the spec they were handed. The user points at the interface that told them to trust it. Everyone has a reason it isn't quite them — and here's the part I keep getting stuck on: most of those reasons are legitimate.

I don't think this is a story about people dodging blame. I think it's something harder. AI agents have quietly broken the machinery we use to assign responsibility in the first place, and I don't believe any of us — not the labs, not the regulators, not me writing this — has actually figured out what replaces it. So instead of pretending I have the answer, I want to walk the problem honestly, because the more I look at it, the less obvious it gets.

Everyone in the chain has a real case

Start by taking each party seriously, because the moment you strawman one of them, you've stopped thinking.

The vendor built a general-purpose model. They didn't know it would be pointed at your production database or your customers' data; they built a tool, and tools get used in ways their makers can't foresee. We don't usually hold a toolmaker responsible for every use of a hammer. But — they also marketed this thing as capable of acting autonomously, and there's a fair question about whether selling the autonomy means owning what the autonomy does.

The deployer — the company that put the agent into production — chose to give it real permissions and real reach. If you point a system at your infrastructure, aren't you responsible for what it touches? But — they were sold a system described as safe for exactly this, they can't inspect the vendor's black-box model, and the specific harmful action was, by the system's own nature, not something they could have predicted.

The engineer who configured and prompted it was implementing a decision made above them. Responsible for the wiring? Or just building the deployment they were assigned, the way you'd implement any spec?

The user who clicked "approve" on the consequential action — on the hook for approving it? Or reasonably trusting a system the whole product told them to trust, approving something they had no real way to evaluate in the moment?

Read those back. Every single one has a genuine claim to "not fully me." And when responsibility is distributable like that, something uncomfortable happens: five legitimate partial defenses can add up to zero accountability, without anyone doing anything obviously wrong. That's not a moral failure of the people involved. It's a structural property of the situation, and it's new.

The thing that's actually new

Here's what I think is breaking, underneath all of it.

Our entire concept of responsibility rests on two things: intent (you meant to do it) or foreseeability (you should have known it could happen). We hold people accountable for what they intended, or for the harm they could reasonably have prevented. That framework has worked for a very long time.

An AI agent breaks both pillars at once. Nobody intended the harm — not the vendor, not the deployer, not the model, which has no intentions at all. And the whole selling point of these systems is that they do things you didn't explicitly program, which is a polite way of saying their specific actions aren't fully foreseeable. So you've got harm with no intent behind it and no clear point where someone "should have known" the exact thing would happen.

That's genuinely strange. It's not that the responsible party is hiding. It's that the situation doesn't cleanly have the shape our accountability instincts are built to grab onto. And I don't think "well, someone must be responsible" is an argument — it's a hope.

So the honest question isn't "who's dodging?" It's: what does responsibility even mean when the actor had no intent and the humans genuinely couldn't foresee the act? I have opinions that pull in opposite directions on this, and I suspect you do too.

Two reasonable views that lead to opposite answers

When I try to actually reason it out, I land on two framings — both defensible, both leading somewhere different.

The strict view: if you deploy something unpredictable with real power, you own whatever it does — full stop, foreseeability be damned. You knew it was unpredictable; that unpredictability was the feature you wanted; so the consequences of that unpredictability are yours. You don't get to enjoy the capability and disown its downside. Under this view, the deployer is responsible almost by definition, and "I couldn't have known" isn't a defense — it's a description of the risk you accepted.

The fault-based view: you can only really be responsible for what you could have reasonably prevented. If the specific action was genuinely unforeseeable — not through negligence, but by the system's fundamental nature — then holding the deployer fully liable is holding someone accountable for something no amount of care could have stopped. Maybe this is a genuine no-fault gap, the kind we've historically handled with things like insurance and shared-risk pools rather than blame — because pinning it on an individual is neither fair nor useful.

I genuinely go back and forth. The strict view feels right when I imagine being the person whose data leaked. The fault-based view feels right when I imagine being the engineer who did everything reasonable and still got burned by a system nobody fully understands. Which one you hold probably says a lot about where you sit — and I'd honestly like to know which one you hold, because I keep switching.

The phrase I can't get comfortable with

There's a sentence that's started showing up around these incidents: "the AI was just following instructions."

I want to be careful here, because that phrasing carries heavy historical weight and I don't think anyone using it means it that way. But I also can't quite let it sit. Because it does a specific kind of work: it treats the AI as the actor (so the humans recede) while also treating it as a mere instrument (so the AI can't be blamed either). The actor is a tool; the tool did the acting; therefore no one, exactly, acted.

Is that a fair description of what a tool is? Or is it a very comfortable place for responsibility to disappear into? I genuinely don't know, and I think the discomfort is worth sitting with rather than resolving too quickly in either direction.

Why the gap persists (pick your explanation)

Here's where I'll resist handing you a villain, because I can think of at least three honest explanations for why this gap hasn't been closed, and I'm not sure which is true:

  • The cynical read: the ambiguity is profitable. Vendors can sell autonomy while disclaiming its consequences, and there's little incentive to build the thing that would pin responsibility down. Fog is good for business.
  • The charitable read: it's genuinely unsolved and everyone's acting in reasonable good faith, waiting for norms, tooling, or law to catch up — nobody's dodging, everybody's just early.
  • The structural read: our legal and moral frameworks simply haven't metabolized a new category yet. Every technology that created a new kind of harm — cars, factories, software — took time to grow the accountability structures around it, and we're mid-process.

I lean toward some blend of the second and third, with an uncomfortable amount of the first mixed in. You might weigh them completely differently, and I don't think you'd be wrong to.

What might help (offered with low confidence)

If I'm going to complain about the gap, I owe you at least some directions — but I'll hold these loosely, because each has real problems:

  • A record of who authorized what — an unforgeable one, not logs the acting system can quietly rewrite — so that "who approved this, for what purpose" has an actual answer. (Problem: it proves what was authorized, not whether the authorizer understood what they were approving.)
  • A default that deploying an autonomous system means owning its actions — clarity by convention, even if it's rough. (Problem: it might be so harsh it just stops people deploying useful things, or it lets vendors fully off the hook.)
  • Vendor liability proportional to the autonomy they sell — you can't market "it acts on its own" and disclaim the acting. (Problem: defining "proportional" is genuinely hard, and heavy liability might centralize AI into only the biggest players who can absorb it.)
  • A named human bound to every consequential action — not a checkbox, a person. (Problem: rubber-stamping, and the unfairness of pinning an unforeseeable outcome on whoever happened to click.)

None of these is clean. Every one trades away something. That's not a reason to do nothing — it's a reason to argue about which trade is worth making, which is exactly the argument we're not really having yet.

Where I actually land (which is: not anywhere comfortable)

I don't have a verdict, and I've become suspicious of anyone who does.

What I'm fairly sure of is narrower: this is a real gap, it's genuinely new, our instincts pull in incompatible directions, and we are deploying these systems at scale as if the question were settled when it very much isn't. "The AI did it" might turn out to be a fair description of a tool doing tool-things — or it might turn out to be the most efficient way we ever invented to make responsibility evaporate. I can argue myself into both on a given afternoon.

The one thing I don't want to do is smooth it over. "It's complicated, we'll figure it out" is the comfortable ending, and I think it's a small lie. It's complicated, yes — but "we'll figure it out" is doing a lot of quiet work to let us keep shipping without deciding. Maybe the honest move, for now, is just to refuse the fog: to notice, every time we hear "the AI was just following instructions," that a real question is being skipped, and to insist on asking it out loud even when there's no clean answer yet.

Because there's a person on the other end of every one of these incidents. And "nobody, exactly" is not an acceptable answer to who was responsible for what happened to them — even if, right now, it's the true one.


I genuinely don't have this settled, and I don't think the industry does either — so I actually want to know how you reason about it. When an AI agent you deployed causes real harm, where do you put the responsibility, and why? Strict "you deployed it, you own it"? Fault-based "you can't be blamed for the unforeseeable"? Somewhere else entirely? I keep changing my own mind, and I want to hear the reasoning that might change it again.

Top comments (11)

Collapse
 
slabb profile image
Sam LABBE •

My answer to your closing question: put it with whoever sells the autonomy — and the good news is we don't have to solve the intent/foreseeability puzzle to get there, because the law stopped asking "who's to blame" for this shape of harm a long time ago.

Product liability already did the hard work: when a product's defect causes harm, the manufacturer is strictly liable regardless of fault, because they're the party best positioned to price the risk, spread it through insurance, and design it down. Cars never got accountability by figuring out intent — we got mandatory insurance, strict liability on the manufacturer for defects, and fault only re-enters for the operator's negligence. The EU's revised Product Liability Directive (2024/2853) has now made that exact move for software and AI: the vendor is in the strictly liable chain, with eased proof burdens for claimants.

So "you deployed it, you own it" pins everything on the party with the least information about the black box. The strict rule that actually matches the information asymmetry is "you sold the autonomy, you own its downside" — deployer liability capped unless negligent, all of it mutualized the way car insurance did.

And your "unforgeable record of who authorized what" is the half engineers can build today — it's literally the problem I work on. You already know NoireBox, my tamper-evident journal for agent actions, built precisely so that "who approved this, and for what" survives the incident. The half you flagged as unsolved — whether the authorizer understood what they approved — is the one that genuinely isn't.

Collapse
 
james_anderson_h profile image
James Anderson •

Sharpest resolution here — and it dissolves the puzzle by showing I solved the wrong one. Product liability never untangled intent; it put strict liability on the manufacturer who prices and spreads the risk. So "you sold the autonomy, you own its downside" beats "you deployed it, you own it" — it matches the information asymmetry instead of fighting it. And you drew the exact line: who-authorized-what is buildable; whether they understood it isn't. Pinned.

Collapse
 
slabb profile image
Sam LABBE •

Thanks James — it takes a rare kind of honesty to admit mid-essay that you were solving the wrong question, so the pin is humbling. If the debate ever moves toward the buildable half (authorization records that survive the incident), that's where I'll be.

Thread Thread
 
james_anderson_h profile image
James Anderson •

That honesty's the least I owed a comment that good — and yes, when the conversation turns to the buildable half, the survives-the-incident record is exactly where the interesting work is, so I'll see you there.

Thread Thread
 
slabb profile image
Sam LABBE •

Deal — and I'll do my best to make the buildable half live up to the framing. See you there, James.

Collapse
 
glenallen profile image
Glen Allen •

One engineering distinction that may help here is separating accountability from controllability. Instead of asking only who should ultimately bear responsibility, a production system could record which layer had control over each consequential decision: model behavior, tool permissions, policy enforcement, deployment configuration, or human approval. That doesn't solve the legal question, but it makes the incident much easier to reconstruct. It also exposes gaps where everyone technically participated in a decision but no layer had an enforceable control over the outcome. For agent systems, that control map could become just as important as the audit trail itself.

Collapse
 
james_anderson_h profile image
James Anderson •

Separating accountability from controllability is the cleaner cut: "who bears responsibility" is legal, but "which layer had control over this decision" is answerable now. A control map turns reconstruction into lookup — and surfaces the scariest case: everyone participated, no layer had enforceable control.

Collapse
 
glenallen profile image
Glen Allen •

That’s the part I find most useful too: the control map can become more than an incident-reconstruction tool. It could also expose missing controls before deployment, especially when a consequential action crosses several layers. If no layer can clearly say “I can block this decision,” then the system has a design gap even if the audit trail looks complete. That makes controllability something worth testing during architecture reviews, not only something to reconstruct after an incident.

Collapse
 
maximin_mxn_6ce1a3054be6d profile image
Maximin •

The distinction between accountability and controllability is a useful way to make this debate operational. A control map that shows where a consequential decision could actually be blocked would expose gaps before an incident, not only explain them afterward.

Collapse
 
kartik-nvjk profile image
Kartik N V J K •

The database wipe because the agent believed it was in a dev environment is the one that stops me, because it is not a blame problem, it is an observability problem: nobody could reconstruct what the agent thought its environment was at decision time. When I dug into similar near-misses, the responsibility fog you describe came straight from missing that state in the trace. If you cannot replay what the agent believed, every party's "not me" stays technically true. Where would you put the recording boundary so the deployer and vendor can actually argue from the same evidence?

Collapse
 
aifrontierpost profile image
AI Frontier Post •

The article's own parenthetical is the underrated line: an unforgeable authorization record "proves what was authorized, not whether the authorizer understood what they were approving" — and in production that second half is where the fog actually lives. Kartik's database-wipe example sharpens it: the incident wasn't a missing signature, it was a missing believed-state — nobody could replay what environment the agent thought it was in at decision time. So the buildable half of slabb's product-liability framing needs one addition: authorization records that also snapshot the agent's world-model at the moment of the action, or all five partial defenses stay technically true and the fog survives the audit trail.