I use Astra in Codex for work beyond software development. Opus writes code well, so I wanted to hand implementation to Claude Code and keep more of Astra's allowance for everything else.
The obvious workflow was to send the task, wait for the code, and check the result. I started building TaskRoute to handle that exchange without making me copy messages between two models.
I compared that route with having Astra implement the same task directly.
The code moved. Much of the work stayed.
In one early comparison, delegation reduced Astra's uncached input tokens by only 3% and took 49% longer. Both versions passed the same 45 independent checks.
An even smaller task had been worse: 29% more uncached input, despite producing less output.
The logs showed the work Astra was still doing around the handoff. Astra read large snapshots, repeated verification, and assembled the result from several files. In the 45-check comparison, the delegated route needed 10 model steps versus 7 for direct implementation.
Removing code generation had removed only part of Astra's work.
A smaller handoff, with evidence
I collected the changes, test results, and independent reviewer findings into one compact acceptance packet. Astra still had to inspect the actual result. The packet made that inspection less scattered.
Waiting needed attention too. One compact-handoff run still woke Astra 13 times while Claude worked. A later run reduced that to two waits by letting the process wait longer between returns to the model.
That change reduced repeated context processing, but it wasn't a controlled measurement of waiting alone. The reviewer output and model behavior also varied.
A new task
I then compared both routes on a new synthetic Python task: planning batches of dependent tasks, including priorities, completed tasks, missing dependencies, and cycles.
Both routes started with the same stub and contract, used fresh Astra sessions, and faced 20 frozen acceptance tests. Each implementation also added eight tests.
| Astra execution workload | Change with TaskRoute |
|---|---|
| Uncached input tokens | 59% lower |
| Output tokens | 83% lower |
| Input including cache | 44% lower |
| Elapsed time | Almost unchanged |
Both implementations passed all 20 common tests and their eight additional tests. No live retries were needed in this pair.
These percentages describe Astra's execution sessions. Claude's usage is separate. Shared experiment preparation is also separate. This was one sequential pair on a new task, with visible tests, not a broad or randomized benchmark. It does not establish equivalent savings in subscription allowance.
A missing directory can erase the gain
In an earlier run, a missing temporary directory stopped the process before review, although Claude had already produced code.
The successful repeat alone showed a 34% reduction in uncached input. Including the failed attempt reduced that saving to 7%, before counting the parent session's diagnosis and repair work.
I added a local preflight to check the required files and directories before calling the models.
Tests also found a problem in my instructions
Another comparison could not establish equal quality. Both implementations passed their own tests, but the shared checks exposed different interpretations of which error to return when several rules failed.
The task description had left that precedence ambiguous. I clarified the contract instead of treating the result as a model ranking.
What I have now
TaskRoute is a small local Codex plugin for bounded Python function changes on macOS, using an existing Claude Code setup. Claude implements and brings in a separate reviewer; Codex checks the returned work.
The repository includes the measured comparison and its limits, plus installation instructions for Codex.
I still need to measure how this affects an ordinary working week. For now, the useful change is concrete: Astra receives a result it can inspect without spending so much effort managing the exchange.
If TaskRoute is useful to you, GitHub stars are welcome.
Top comments (0)