DEV Community

SeaOtter ( Kentaro Mori )
SeaOtter ( Kentaro Mori )

Posted on

Seven Days Managing an AI Development Team

On the first day, I planned nine tasks. Three got done.

The other six ran into familiar problems: decisions waiting on me, prompts that I had to copy into local GUIs, and uncertainty about what was actually progressing. I had delegated the work but was still carrying much of the coordination.

From October 1 to 7, 2026, I kept a daily log of running a small AI team for personal projects. I used ChatGPT dots as the coordinator, which I call “Docchee,” with other AI tools handling implementation, review, and research.

By Day 7, I was asking the team to switch priorities for an urgent application-preparation task. Between those two points, I learned more about handoffs, approvals, and visibility than I expected.

This is a field report from one person's workflow. It is not a model benchmark or evidence that the setup can safely handle a production incident.

The team I was trying to run

I work in cloud engineering, but for this experiment I took the Product Owner role: define the goal, choose priorities, approve consequential actions, and test whether the result was useful.

The working split looked like this:

Role Responsibility
Me Goals, priorities, approvals, and acceptance testing
Docchee Coordinate tasks, choose an execution environment, and report status
ChatGPT / Codex Design, review, and implementation when needed
Claude Code Implementation and fixes
Gemini / Antigravity Share prototyping, research, and other suitable work
Qwen A separate track of local and GPU-backed experiments

These were working roles, not a fully automated routing architecture. Assignments changed with the task, available usage allowance, and access to the right environment. I keep the broader workflow in multi-ai-workflow.

Days 1–2: Handoffs were my problem too

My initial assumption was that handing out more tasks would get more work moving. Instead, it produced more places where someone needed my input.

At the start, I was manually moving prompts between interfaces. CLI access was not ready. Some tasks needed a decision I had not made explicit; others needed a local action only I could perform at that point.

The first useful adjustment was simple: make the end goal clearer, and have Docchee tell me what remained instead of making me repeatedly ask.

By Day 2, a workable rhythm was emerging. Human-in-the-loop approvals remained difficult. I wanted routine work to move without interruption, while keeping meaningful control over publishing, permissions, and other consequential actions.

That required more precise task boundaries. “Keep going” was not enough information for every next step.

Days 3–4: Local access improved, but approval boundaries remained

On Day 3, I focused on integrating the team with my Windows environment and preparing CLI-based workflows. Reducing manual copy-paste was a practical improvement.

Usage allowances also became part of task assignment. A difficult design or review could justify a stronger model, while other work could go to a different tool with more capacity available.

I was excited enough in my daily post to describe this as the start of token-aware coordination. The precise state was more modest: we were using available displays and reports to make allocation decisions. Continuous usage collection and dynamic routing were still goals, not a completed control system.

Day 4 exposed another boundary. A task routed through dots to ChatGPT or Codex could still stop for a separate approval. Using tools from the same provider did not make the entire workflow share one permission state.

The useful question became: what exact operation is blocked, in which environment, and what would let it proceed? A generic “waiting for approval” label left too much work for me to reconstruct.

Day 5: Voice instructions and a 30-minute cadence worked well

On Day 5, I ran much of the workflow through voice input. For me, speaking an instruction or review request was often faster than composing it at a keyboard.

I also introduced a status update every 30 minutes across the parallel tasks:

  1. I set the goal and priority.
  2. Docchee organized the work and assigned it.
  3. I reviewed progress and blockers at the next check-in.
  4. I returned decisions, corrections, or new requests.
  5. The work continued.

Numbered items made it easy to say “prioritize item 2” or “resume item 3.” There was a catch: once completed items disappeared and the list was renumbered, a number was no longer a stable task identifier. It needed to be checked against the task name and recent context.

I asked for completed items to be reported once and then removed. That kept the next report focused on open work and decisions.

Thirty minutes is not a universal recommendation. It happened to fit my workflow. It reduced the effort of checking in without requiring me to watch each tool continuously.

Delegating work across Claude Code, Antigravity, and ChatGPT also became more practical. The amount of coordination I could hand off still depended on access and approval in each environment.

Day 6: More kinds of work meant more invisible state

The scope expanded beyond coding and documentation: website onboarding, profile updates, and short video assets. Having an agent inspect a page and work through form fields removed a lot of repetitive effort.

The difficult part was understanding stalled background work.

I could follow a task in a visible chat or browser window. CLI processes and background activity were harder to inspect. If two tasks both said they were progressing but neither returned a result, I did not know which was running, which was waiting, or whether they were interfering with each other.

My hypothesis was that some tasks were competing for shared browser sessions or execution resources. I did not establish that as the root cause of every stall. What I did observe was inadequate visibility and a need for manual intervention.

For the next iteration, I want each active task to expose:

  • Its current owner and execution environment
  • Its last confirmed progress and timestamp
  • Whether it is running or waiting
  • The specific input, permission, or resource it needs next

This is a design goal rather than a dashboard I can claim is already complete. Shared desktop control needs particular care: parallel research can be useful, while simultaneous edits in the same interface can conflict.

Day 7: An urgent request changed the queue

The final day brought an interrupt: I wanted to prioritize preparation for Anthropic's CVP application. That involved reviewing my existing work, developing security-validation evidence, and contributing to OSS.

I explicitly assigned Astra to the urgent work. Progress felt fast. It would be misleading, however, to attribute the improvement solely to the model. Priority, task definition, and the surrounding workflow had changed too; I did not run a controlled comparison.

One concrete output was a validation report on an existing Mattermost pull request. I posted a comment with the comparison results. This was a testing contribution, not an upstream merge of my code.

In parallel, we prepared a read-only connection between Mattermost and Docchee. A short tunnel startup-and-shutdown trial had been performed, but reading an actual post through that path had not been verified at the time of writing. The integration was still unfinished.

The CVP application was also still in preparation. Useful portfolio work does not automatically establish the eligibility evidence an application asks for. I want the application to describe what I have actually done, without turning planned work or ordinary OSS activity into unsupported credentials.

What Day 7 demonstrated for me was the ability to redirect a development workflow under time pressure. Production incident response remains a separate capability to test.

The operating habits I want to keep

Give the task an observable finish line

I find this handoff format useful:

Goal: What should become possible?
Scope: What can change?
Out of scope: What should be left alone?
Environment: Where will the work run and be tested?
Acceptance: What observation will establish success?
Report: Deliverable, evidence, unverified paths, decisions needed
Enter fullscreen mode Exit fullscreen mode

It reduces ambiguity before the team starts. It also makes it easier to assess a result without rereading an entire conversation.

Separate implementation from review

I use a separate reviewer for changes, including failure paths such as timeouts, cancellation, restart, and cleanup. Findings should come back with evidence or reproduction steps so the implementation agent can act on them.

A second model's agreement is not sufficient evidence by itself. The useful part is the additional inspection and testing it performs.

Report the actual stage of completion

Implemented, passed a mock test, worked in the target environment, and accepted by the user are different milestones.

The same applies to browser work: uploading a file, saving a draft, and publishing a page are separate outcomes. I need the status report to tell me which one happened.

This became especially important as the team handled more kinds of tasks. A broad “done” could hide the exact step I still needed to verify.

End the day with a short voice debrief

The most useful routine was a ten-minute conversation at the end of the day.

I talked through where I had waited, what had become easier, and what I wanted to try next. The visibility problem on Day 6 and the incident-response idea on Day 7 both became clearer through those conversations.

Short instructions helped me move work during the day. A longer conversation helped me reflect on the workflow itself. That combination was a practical form of human-in-the-loop operation for me.

Next: Test emergency workflows in a lab

My next question comes from my cloud-engineering background: how would this team respond to a server failure, a security incident, or a disaster-recovery scenario?

I want to survey existing SRE-agent approaches and design experiments across AWS, Azure, and Google Cloud. Those experiments have not started.

The plan is to use controlled environments with defined permissions and recovery conditions. I want to measure time to initial assessment and recovery, but also incorrect actions, unnecessary changes, approval delays, and missing evidence.

Parallel investigation could help. Changes to the same resource need coordination, and recovery needs a way back if the chosen action makes things worse. These are things to verify before making claims about autonomous incident handling.

What changed after one week

I started by counting tasks. By the end of the week, I was paying more attention to the handoffs between them.

As individual tasks moved faster, the surrounding constraints became easier to see: approvals, shared resources, usage allowances, and acceptance criteria. My own decisions remained part of that system.

The 30-minute check-ins and the end-of-day voice debrief are worth keeping. Next, I want better visibility into stalled work and stronger evidence about how the team behaves when something fails.

Seven days is a short experiment. It was enough to give me a much clearer idea of what to test in week two.

This article reflects my personal projects as of October 7, 2026. I used AI assistance to organize and edit the writing.

Top comments (0)