I run three or four AI coding CLIs at once: Claude Code writing, Codex reviewing, Gemini CLI somewhere else. For months I was the courier between them. Copy the diff out of one window, paste it into another, wait, copy the review back.
So I built the thing I wanted to type instead:
Build the settings page, then ask @codex to review it, and loop until it is clean.
One agent does the work, hands it to the other with an @ mention the way you would in Slack, reads the answer, fixes what it was told, and asks again.
The mention itself was the easy part. Everything around it was harder than I expected, and most of it applies to anyone gluing terminal AIs together, with or without my tool. Here is what I would tell myself before starting.
1. The screen is not the reply
My first version read the other agent's answer off its terminal. It was wrong in three ways at once.
- It is a picture, not text. The reply is wrapped at the window's width and drawn inside the tool's own frame: borders, spinners, status lines. Un-wrapping it is guesswork.
- Part of it is simply gone. Claude Code in its full-screen mode runs on the terminal's alternate screen, which keeps no scrollback at all. Whatever scrolled past is gone from your side the moment it moves.
- The previous answer can pass for this one. If the new turn has not drawn anything yet, the last thing on screen is the old reply. A script that reads "whatever is there" will happily report it.
The fix: every one of these CLIs already writes a complete record of the conversation to disk, for its own resume feature. Read that instead.
| CLI | Where the record lives (on current versions) |
|---|---|
| Claude Code | ~/.claude/projects/<project>/<session-id>.jsonl |
| Codex CLI | ~/.codex/sessions/YYYY/MM/DD/rollout-*-<session-id>.jsonl |
| Gemini CLI | ~/.gemini/tmp/<project>/chats/session-*.jsonl |
Two rules made this reliable:
- Anchor on your own prompt. Find the prompt you sent in the record (reading back the last dozen or so turns is enough), and take only what the AI said after it. The previous answer sits before your prompt, so it can never be mistaken for this one.
- Keep a fallback. A CLI that keeps no record still gets read off the screen. It is the last resort, not the plan.
2. Read those files from the end
These transcripts grow for the whole life of a conversation. On an ordinary week of work I measured 318 MB of them. The part you want is always the last thing said.
A reader that walks from the front gets slower every day you use it. Seek to the end and read backwards, line by line, until you have the turn you need.
3. Don't write a parser per CLI
Each CLI arranges its JSON differently, and rearranges it between releases. A spec per CLI is a promise to keep chasing them. Two rules hold across all three instead:
-
The message is the object that carries a
role: the record itself, or the single field it is wrapped in (messagefor Claude,payloadfor Codex). A record with no role names its speaker in its owntype(user/geminifor Gemini). -
The words are the blocks whose
typeends intext(text,output_text,input_text), or untyped blocks that carry atextand are not marked as a thought.
Everything else in a content list is machinery: tool calls, tool results, the model's own thinking. That is not what "the reply" means. These two rules have survived several releases of all three CLIs without a change.
4. Knowing when the other agent is done
Before you can read a reply, you need to know the turn has ended. A terminal gives you no API for that. What worked was layering signals, in this order of trust:
- The CLI's own hooks, where it has them. Claude Code, Codex and Gemini CLI can all run a command when a turn starts and ends.
- The window title, which several CLIs set while a turn is running. It can only tell you that a turn is running.
- Patterns on the screen, kept in a per-CLI config file rather than compiled in, because the CLI's next release will change its screen.
- Output silence, as the fallback for a CLI nobody has described yet.
Two details mattered more than the list itself:
- Take "busy" from the hook and "finished" from the screen. No CLI reports the answer to its own permission dialog. Approving, refusing and pressing Ctrl+C are things a person does to the CLI, not events the CLI emits.
- A question on screen outranks everything. An agent that looks busy while it is actually waiting for a human is the one mistake nobody goes back to check. If the agent you asked is waiting on a permission prompt, the caller should hear "it is waiting for a person", not silence.
5. Don't let a long job time out the one who asked
An AI calling a command expects an answer within a few minutes; its tool call has a timeout. A review of a large change can take longer than that.
So the call does not simply block. It waits up to 100 seconds. If the other agent is still going, the command answers:
[shikisha] STILL WORKING: <@otter> has not finished. End your turn now;
its reply will be typed into this tab when it is done
The calling agent ends its turn cleanly. When the reviewer finishes, its reply is typed into the caller's input as a new message, and the loop carries on. Nothing is lost if the caller gives up waiting, or if a person presses Esc in the meantime.
6. Decide what an agent is allowed to drive
Letting one agent prompt another is the point. Letting it type into any terminal is not. The rules I ended up with:
-
Another AI tab can be asked. A plain terminal or a web page is only driven if a person named it with
@in the message that started the work. - The app counts the rounds and stops at a limit you set. "Loop until it is clean" should not mean "loop until the subscription runs out".
-
The instructions are opt-in. The hand-off is a small command on every tab's PATH (
shikisha ask,run,do) plus a skill file that tells the CLI how to use it. The skill is written only after you agree, and one setting removes it again.
A command plus a skill turned out simpler than a per-client integration: every one of these CLIs can already run a shell command, so the same mechanism works in all of them.
Where this lives
All of this is in SHIKISHA-TERM ("shikisha" is Japanese for the conductor of an orchestra). It is a terminal that runs your AI coding CLIs side by side, shows which one is working, finished or waiting for you, and lets them @mention each other. It can also put a project on a cloud MicroVM, so an agent running without permission prompts is not running on your laptop.
Honest limits: the window is Windows only (Linux gets a headless build you watch from a browser or phone), it drives the CLIs you already have rather than replacing them, and it is free and MIT-licensed.
If you have glued terminal AIs together yourself, I would like to hear what broke for you. The part I am least sure about is the state detection. It is heuristic, and every new CLI release is a chance for it to be wrong.

Top comments (0)