DEV Community

Jiahui Miao
Jiahui Miao

Posted on

Stop Asking Your AI to Be Helpful. Ask It to Be Accountable.

By Jiahui Miao

Every agent demo ends the same way. The agent smiles, the audience claps, and nobody checks the work. The industry's favorite word for AI right now is "helpful." I've stopped using it entirely.

I'm building a personal AI for exactly one person — me — and living with it every day has rearranged my vocabulary. Helpful turned out to be the cheapest quality an agent can have. What I actually needed, and what took real work to build, was something harder: an agent that can be held accountable. Not an agent that sounds right. An agent that shows its work, carries its debts, and answers for its mistakes — on the record, by name.

Helpfulness is a vibe. Accountability is a structure.

Here's what the helpfulness paradigm actually optimizes for: agreeableness. An agent graded on helpfulness learns to be the colleague who never pushes back, never admits uncertainty, never leaves an awkward paper trail. It volunteers, it smooths, it reassures. And when it's wrong, there's nothing to audit — because helpful agents don't produce receipts. They produce vibes.

The entire benchmarking culture reinforces this. Leaderboards score agents on how pleasing their answers are. Demo culture never asks the follow-up question: show me where that number came from. A helpful agent that can't show its work isn't your assistant. It's a confident intern with a shredder.

Accountability is the opposite discipline. It doesn't ask whether the agent sounds good. It asks three questions, every day: who owns this, when is it due, and where is the evidence. If an agent can't answer all three, its output doesn't ship. That rule applies to my personal AI, and — because my company vertciti exists to operate my life — it applies to my company too. A row merely written is not done. Every commitment needs an owner, a deadline, and a definition of done, or it doesn't exist. I audit the open ones daily. My agent lives under the same ledger discipline I run my company on.

The three debts every accountable agent owes

One: named work. Every output my agent produces carries its provenance — which inputs it used, which assumptions it made, what it verified and what it didn't. Anonymous work is unauditable work, and unauditable work doesn't ship. I participate in 3GPP's work on 6G core network standards, and the one lesson every standards body learns the hard way is this: nothing ships without a trace. Who proposed it, who objected, what changed between drafts — protocols are accountability machines. Your personal AI should be one too.

Two: committed work. When my agent takes something on, the commitment enters the ledger with a name, a date, and a definition of done — and it stays open until it closes out. This is where the helpfulness paradigm fails most visibly: helpful agents volunteer constantly and silently drop half of it. Nobody notices, because there was never a record. My agent's open commitments get audited. An agent that volunteers and forgets is a liability wearing a smile.

Three: retracted work. Last week I wrote about how my agent handles memory — and the core of it was retraction. I once put a hand-tallied aggregate in a report; a machine recount that evening proved it matched no real count. I retracted it the same night and mechanized the recount so no hand-written aggregate ever ships again. Caught, retracted, mechanized — in public, on the record, the same day. An agent that cannot retract accumulates its own mistakes into convictions. Mine forgets faster than I do, and that's the point. Retraction isn't failure. It's the only honest way to stay correct over time.

The opponent

The opponent isn't a company or a model. It's a culture: the helpful-assistant maximalism that grades agents on agreeableness and never asks for the receipt. Every leaderboard that scores "helpfulness" without scoring traceability is training the next generation of agents to be charming and unaccountable. Every demo that ends in applause instead of an audit is teaching buyers to accept vibes as infrastructure.

I have the opposite test now. When I evaluate an agent — mine or anyone's — I ask three things: who owns its mistakes, what expires, and when it last retracted something on the record. If it can't answer, it's not infrastructure. It's theater.

So stop asking your AI to be helpful. Helpfulness is cheap, and the industry is already mass-producing it. Ask it to be accountable. Ask who owns the work, when the debt is due, and where the receipt is. The agents that survive that audit are the only ones worth building a life on. Mine is learning to pass it. Every day, on the record.

Top comments (0)