AI coding agents changed the way I work.
When I started using them, the workflow was simple:
Give an agent a task → wait → review the result → c...
For further actions, you may consider blocking this person and/or reporting abuse
"The bottleneck became me" is the whole future of this arriving early. I run a small team of AI agents too — planner, engineer, QA, security, deploy — and the moment I went parallel, the work stopped being "can the agent do it" and became "which of these six is waiting on me right now, and for what." The agents scaled fine. My ability to hold six live states in my head did not.
The reframe that helped: I stopped trying to read the agents and started making them report in a fixed shape. Each one has to end with what it did, how it verified it, and what's still open — not prose I have to parse, a structure I can scan. It doesn't reduce the number of agents; it reduces the mental cost per agent from "reconstruct what happened" to "read one status line." Supervising a team is exactly the right frame, and the thing every new manager learns is that the skill isn't doing the work — it's building the reporting that lets you not watch all of it.
The unsolved part you're pointing at — visibility across machines without drowning in it — is the real frontier. My hack is that idle agents cost nothing, so I'd rather have one waiting on a clear question than three half-running and ambiguous. Fewer live unknowns beats more throughput, at least for the one human in the loop. Great question to be sitting in; I don't think anyone's solved it yet.
This is really a mind-opening point. I think what we actually need is to build a trust boundary between ourselves and agents, rather than trying to understand every single detail of what every agent is doing.
Another great insight 😄. I agree that we probably only need to focus on a small number of important events and states. Ultimately, this comes down to building a trust relationship with AI workers, similar to how we build trust with junior team members.
We may need a new way of collaborating with AI that is designed around human attention and reporting mechanisms, rather than simply adding more agents or more information.
I have some ideas around this direction, but turning them into something that is actually practical and usable still feels like a long way to go.
"Build a trust boundary rather than understand every detail" is the right frame, and I want to push on one word in it, because it's the word I keep getting wrong myself: trust. The junior-teammate analogy is half right, and the half that's off is the important half. With a junior human, trust accrues and lets you delegate judgment — you stop checking because they've earned it. With an agent, trust can't attach to judgment the same way, because the thing you'd be trusting resets every session and respects enforcement, not intent. So the boundary isn't "I trust this agent's decisions." It's "I trust this agent to produce an artifact I can check without re-reading the reasoning." You're not delegating the judgment; you're delegating the work and keeping the receipt.
Which actually makes your instinct more workable, not less. You don't have to build the deep trust a junior earns over months — you just have to build the reporting shape that lets a shallow, per-session trust be enough. The clinical version I keep coming back to: I'd trust a nurse I've worked with for years, and I'd still expect the high-risk drug to get a second check — because trust reduces how often I supervise, it never removes the artifact on the paths where being wrong is expensive. Trust lowers the frequency of watching; it doesn't lower the floor of what has to be provable.
On attention-first design being the real frontier — agreed, and the one reframe that's helped me is that the scarce resource isn't the number of events, it's the number of judgment-type switches. Ten agents that all need the same kind of decision from me is cheap; three that each need a different kind is expensive, because the cost is reloading my context, not reading the line. So the reporting isn't just "surface fewer states," it's "batch the states that need the same shape of judgment." I don't have the practical version either — but I think that's the axis it lives on. Genuinely one of the more useful threads I've been in; you're asking the question the tooling hasn't caught up to yet.
I spent some time on vacation 😄, so sorry for the late reply. I guess this also proves that I’m actually replying myself, not using an AI to answer 😄.
“I trust this agent to produce an artifact I can check without re-reading the reasoning.”
This is a very insightful point. What we trust is not the agent’s judgment itself, but the interaction process and the verifiable artifacts it produces.
This is also why, in coding workflows, having strong guardrails is so important: tests, regression checks, acceptance tests, coverage checks, validation of test quality, and other verification mechanisms. They definitely consume more tokens compared with simply letting an agent run freely, but they create a much more reliable boundary.
I think the way humans think has something in common with how computers work — we both have a kind of locality principle. For engineers doing highly cognitive work, focusing on a single task is almost always the most efficient mode.
Human attention usually has only one (or a very small number of) active focus points. When we try to maintain too many contexts at the same time, or switch between too many mental states in a short period, we quickly become overwhelmed.
My assumption is that once we establish a proper trust relationship with AI workers, the thing we need to focus on is not the details of every action, but the right moments to intervene.
This may actually be very similar to how we train junior teammates. At the beginning, we need them to provide daily reports and operate within a clear communication process. After enough confidence is built, we no longer need to monitor every step — we mainly review outcomes.
I’m really happy to see such insightful thoughts in this thread. This has been one of the more valuable discussions I’ve had recently.
I’m going to explore this direction further, and I’d love to keep the discussion going with you.
No apology needed — the vacation is the best possible proof of authorship 😄.
The junior-teammate analogy is the one I keep reaching for too, and it holds almost all the way, then breaks at one joint that matters for exactly your question — when to intervene. With a junior, trust grows monotonically and failures are legible: competence compounds, the weak spots are predictable, and a stuck junior looks stuck — the daily report says "I couldn't figure this out." So moving from step-monitoring to outcome-review is safe, because the risk drops in a way you can see. An agent breaks both halves. Trust isn't monotonic — flawless ten times, then confidently wrong on the eleventh. And the failure is fluent: it arrives looking identical to the successes, with no "I'm stuck" tell. A stuck junior raises a hand; a stuck agent hands you a clean-looking artifact.
So one edit on "focus on the right moments to intervene": at high trust with a human you choose those moments — with an agent you can't, and locality is exactly why. The single task you're focused on is where you'll miss a fluent failure two tasks over. The moment has to be raised by the artifact, not spotted by you. That's what the guardrails you listed actually buy — not fewer tokens, but a system where the artifact signals when it needs you instead of relying on you to catch it. This week a check of mine sat green for weeks while quietly scoring six clean files as broken; I never "chose" to intervene, because it looked fine the whole time. It became an intervention moment only once I'd built the gate that could go red. The gate picks the moment now; I just answer.
One of the better threads for me too — glad to keep it going.
I think you’re absolutely right. Human learning is usually a gradual curve, while AI capability can appear to improve almost instantly.
A junior teammate usually struggles to produce a perfectly structured explanation for a wrong solution. But AI can do exactly that — it can provide a very convincing and coherent explanation for an incorrect result. If the reviewer does not fully understand the details of the approach, it is surprisingly easy to get pulled into a perfectly logical but wrong path.
Since I’m a heavy CLI user myself, I’ve been trying to build small tools to solve the problems I personally encounter in my workflow. This whole discussion has given me a lot of inspiration.
At the same time, I don’t want to start by building a universal solution for every possible problem. I’d rather start with something much simpler: reducing the need to constantly switch tabs and helping me manage multiple machines, so I can lower the mental overhead of running many completely unrelated tasks in parallel.
I don’t know whether some of the ideas we discussed here will eventually become part of that system, but I’ve learned a lot from this conversation.
Thanks again for sharing these insights. I really appreciate this discussion.
"The bottleneck became me" — this matches what I got on my own launch. Three of the four comments were about this exact thing, not about features. One person wrote: "if a workflow stalls I'm guessing where it broke."
Two things I learned building orchestration for this:
You cannot answer "is it stuck?" from inside the agent. A hung run and a slow run look identical to the process itself. I had to add an external reaper — anything PENDING or RUNNING past 30 minutes gets marked TIMEOUT. The first time it ran it found 18 dead workflow runs I did not know existed.
And status alone is useless at your scale. "Running" tells you nothing with 10 agents. What helped was tracking which step a run is on: "running, step 3 of 5, started 2 minutes ago" tells you whether to go look. Status is a state; progress is a delta. You need the delta.
Very much agree, very much agree.
If we could also bring the current task summary and progress into focus, that would be even better.
I recently built a small tool around this problem. The first step was simply bringing agent status into my attention space, so I don’t have to constantly monitor the details or switch between tabs. Otherwise, I was already close to mental overload 😄.
The next step might be exploring how to add the delta/progress layer.
But adding delta also means the system needs to understand the meaning behind different agents’ outputs, and this is where I’m struggling with a design decision.
On one hand, I want to keep it as simple as possible and just solve my own problem with a lightweight mechanism. On the other hand, your point about progress and meaningful state changes is really compelling, and part of me wants to explore that direction as well.
So I’m a bit stuck between these two paths 😄.
Great insights and a great discussion. Let’s keep sharing ideas and continue the conversation.