Last year, one of our engineers spent an afternoon asking an LLM to refactor a monitoring service. He'd paste code, get suggestions, paste more code, correct hallucinations, paste again. Four hours later, he had a partial rewrite and a browser history full of ChatGPT tabs.
The code was fine. The process was absurd.
He wasn't using the AI wrong. He was using it exactly how everyone uses it: as a conversation partner who happens to know about code. You ask, it answers, you evaluate, you ask again. It's a better Stack Overflow. That's the ceiling.
We kept bumping into that ceiling. Not because the models were bad, but because the interaction pattern was wrong. A conversation is the wrong unit of work when what you actually need is delegation.
The ask-answer trap
Here's the pattern most engineering teams have settled into with AI. Someone has a task. They open a chat window. They describe what they want, iterating through prompts until the output looks right. Then they copy the result into wherever it actually needs to go.
This works. It's also a dead end.
The problem isn't quality. GPT-4, Claude, Gemini, whatever frontier model you prefer, they all produce genuinely useful output. The problem is that every interaction starts from zero. There's no continuity. No accumulated understanding of your codebase, your preferences, your standards. Every conversation is a first date.
It also doesn't scale. If one person can wrangle an AI through a chat window, great. But try coordinating five AI-assisted workflows across a team of twenty. Who's using which context? Where do the outputs go? How do you review what an AI produced versus what a human wrote? The chat window gives you none of this.
We spent about eight months thinking about what comes after chat. Not a better prompt interface. Something structurally different.
What changes when you give an agent a name
The first thing we tried was simple: stop treating the AI as a tool you invoke and start treating it as a worker you assign things to.
That sounds like semantics. It isn't.
When you assign work to a person, certain things are implicit. They have an identity, so you know who did what. They have context about prior work. They operate within defined boundaries of authority. And when they deliver something, you review it. You accept it or send it back with notes. Over time, they learn your standards.
None of that exists in a chat window. So we built it.
An agent in our system has an identity. Not just a name, an actual authorization boundary. It can access certain tools, certain repos, certain data sources. Its capabilities are described in what we call an AgentCard, a structured declaration of what this particular agent can and cannot do. Think of it like an employee's role description, except it's machine-readable and enforced.
This turns out to matter enormously for teams. When an agent has a defined identity and capability set, you can reason about delegation. You can say "the docs agent handles API reference updates" the same way you'd say "Sarah handles API reference updates." The mental model transfers directly.
Work as a loop, not a thread
Chat is linear. You start at the top, you end at the bottom, the conversation is the artifact.
Real work isn't like that. A task emerges from a conversation, gets shaped, gets executed, comes back for review, might get rejected and reworked. It has an owner. It has deliverables. It has acceptance criteria.
We formalized this into something we call a Loop. A Loop is a work unit that typically starts as a conversation but grows into a structured delivery cycle. Someone, human or agent, owns the Loop. There are defined deliverables. And there's an explicit acceptance step where the requester reviews the output and either accepts or rejects it.
This is different from a ticket system. Tickets are created deliberately, filed into backlogs, groomed and prioritized. Loops emerge naturally from work that's already happening. You're discussing something in a channel, a task crystallizes, someone takes ownership, work happens, results come back. The structure grows from the conversation rather than being imposed on it.
The acceptance step is where it gets interesting. When you reject an agent's output with specific feedback, that's signal. That's data about what "good" looks like in your organization. Which brings us to the part that actually compounds.
Teaching an agent your taste
Most discussions about AI agents focus on capability. Can it write code? Can it search the web? Can it use tools?
Capability is table stakes at this point. The harder question is: can it learn what good looks like for your team specifically?
We built what we call a Preference system. Every time an agent's work gets reviewed, accepted or rejected, that outcome becomes part of the agent's experience. Not fine-tuning. Not RAG in the traditional sense. Structured experience that the agent can reference in future work.
Reject a code review because the agent didn't follow your team's error handling conventions? Next time, it knows. Accept a documentation draft but note that you prefer shorter sentences? That preference persists.
This is where the employee metaphor stops being a metaphor. Employees get better at their jobs through feedback. They develop an understanding of organizational standards that goes beyond what's written in the style guide. They internalize the unwritten rules.
The Preference system is our attempt to give agents the same learning mechanism. It's slow, deliberate, and grounded in real work outcomes. No one is training a model. Everyone is just doing their jobs and the agents absorb the standards through the natural review process.
After a few weeks of active use, the difference is noticeable. Agents start producing output that needs fewer revisions. They align with team conventions that nobody bothered to document because everyone just knew them. Except now the agent knows them too.
Controlling the information surface
Here's something most agent frameworks ignore entirely: who sees what.
When a human team works on something sensitive, you don't put everyone in the same room. The legal review happens separately from the technical implementation. The security audit doesn't need access to the marketing copy. Information flows are controlled by organizational structure.
Agent systems mostly don't do this. They either give every agent access to everything or they silo agents so completely that they can't collaborate at all.
We built six collaboration modes to handle different information topologies. Solo is one agent working independently. Roundtable is a group discussion where everyone sees everything. Critic is a structured review where one agent evaluates another's work. Pipeline passes output through a sequence of agents, each seeing only what the previous one produced. Split divides work into parallel tracks with limited cross-visibility. Swarm is a large-scale coordination mode for complex tasks.
The key insight is that these modes control information visibility, not just task assignment. When you set up a Critic mode review, the reviewing agent only sees the deliverable, not the full context of how it was produced. When you run a Pipeline, each stage gets clean inputs without the noise of upstream deliberation.
This might sound like overengineering. It isn't, once you're running agents on real business data. The question of "what can this agent see" becomes as important as "what can this agent do."
The knowledge that doesn't leave
Here's the compounding effect nobody talks about.
When a senior engineer leaves your company, their institutional knowledge walks out with them. The unwritten rules, the context about why certain decisions were made, the taste that made their code reviews valuable. It's gone.
When an agent accumulates experience through months of accepted and rejected work, that knowledge belongs to the organization. The Preference data, the AgentCard configurations, the evolved understanding of team standards, all of it persists independent of any individual person.
This isn't about replacing people. Senior engineers aren't leaving because AI exists. They're leaving because they got a better offer, or they burned out, or they wanted to try something new. The question is whether the organizational knowledge they accumulated over years can persist in any form after they go.
With traditional tooling, the answer is mostly no. Documentation helps. Runbooks help. But the tacit knowledge, the stuff that's too contextual to write down, evaporates.
An agent that has been trained by your team's actual work output retains that tacit knowledge in a usable form. Not perfectly. Not completely. But meaningfully.
Where we actually are
We should be honest about the current state. Agent systems today, ours included, are early. The models hallucinate. The tool use is brittle. Context windows overflow at inconvenient moments. Running agents on real production work requires supervision that partially defeats the purpose of delegation.
But the trajectory matters more than the current snapshot. The interaction pattern of identity, delegation, structured delivery, and feedback-driven learning is sound even when the underlying models are imperfect. As models improve, the organizational scaffolding around them becomes the differentiator, not the model itself.
We've been building this system at Mininglamp Technology. It's called OCTO, a human-AI collaboration workbench that implements everything described above. It supports multiple agent runtimes including OpenClaw, Claude Code, Codex, and Hermes, because we think vendor lock-in on the execution layer is a mistake. The platform manages identity and tasks. The runtimes handle execution. Data stays in your infrastructure through private deployment.
We also recently open-sourced Mano-P, an on-device GUI agent, at github.com/Mininglamp-AI/Mano-P.
If you're thinking about agent architectures, the code is at github.com/Mininglamp-OSS.
The open question
One thing keeps nagging at us. The Preference system learns from explicit review, accepted or rejected with notes. But the most important organizational knowledge is often implicit. It's the code that nobody writes because everyone knows it would violate an unspoken norm. It's the architecture decisions that seem obvious in retrospect but only because someone with fifteen years of context made them.
How do you teach an agent the things that humans learn by osmosis? We have ideas. We don't have answers yet. If you do, we'd genuinely like to hear them.
Top comments (0)