DEV Community

orca_forge
orca_forge

Posted on Originally published at forge.workstyle.tech

AIに12人の制作チームを演じさせたら、1日半で1億6千万トークンが消えた

📝 Originally published (in Japanese) at forge.workstyle.tech.

The Codex usage screen displayed this:

ERROR: You've hit your usage limit.
try again at Sep 24th, 2026 1:36 PM.
Enter fullscreen mode Exit fullscreen mode

I couldn't use it for another five days. I had completed 38 out of 51 episodes in the video series, leaving 13 more to go.

Just before this, when asked how much I had used Codex, I replied:

Approximately 7.4 million tokens, 101 times
Enter fullscreen mode Exit fullscreen mode

Recalculating the numbers on a different device revealed this:

ai-boardroom   approximately 159,366,148 tokens   about 94% of total Codex usage
Enter fullscreen mode Exit fullscreen mode

I was off by a factor of 20. This article documents the breakdown and why I miscalculated.

What I was building

A series of vertical short videos, approximately 15 seconds each, featuring real-life executives and engineers as caricatures. Here's how one episode is created:

Script (JSON) → Synthetic Voice (TTS)
          → Caricature (FLUX / ComfyUI)
          → Lip Sync (InfiniteTalk)
          → Subtitles, screen, and props composited with ffmpeg → 1080×1920
Enter fullscreen mode Exit fullscreen mode

I changed the production workflow midway. The client's instructions were as follows:

First, you create the scenario and have Codex assemble a video production team. Within that team, conduct reviews, and have them create the script and production instructions. After you visualize it based on their output, request another review from Codex's team, address their feedback, and complete the episode once feedback is exhausted.

I assigned 12 roles to the team: producer, content writer, SNS marketer, screenwriter, director, performer (voice and acting), cinematographer, editor, motion designer, sound engineer, SNS operator, and data analyst. Plus, a proofreader. I instructed Codex to record meeting logs "verbatim, without summarizing, retaining rejected ideas and conflicts."

What worked well

Let me start with the positives. If I only discuss costs, the entire approach might be dismissed.

Mutual skepticism proved effective. During preview review meetings, the team flagged exaggerations I had inadvertently included when transcribing the script. For instance, they rejected an end card phrase ("The man who truly brought a computer to every household") as "definitive attribution of achievement," which I had carried over from an old script.

The same happened in reverse. When I had a separate session write a social media post, comparing it with the script revealed six discrepancies in two rounds:

Post draft: "The person who stated the correct answer at the beginning of the meeting"
Script:     Sam worked alone on the document, and Pichai joined on page 42 → Not at the beginning

Post draft: "They never denied it"
Script:     8.81 seconds: "You haven't reverted it." → Denying the other person's point
Enter fullscreen mode Exit fullscreen mode

The writer hadn't noticed either issue themselves.

Storing rules in a repository allowed for concise instructions. Midway through, I reduced my instructions to Codex to this:

Please complete one vertical short video, 1080×1920.
Script: scripts/v2/ep13_ellison.json
Both the process and assets are in this repository.
Verify the final output as a still image and make any necessary corrections at your discretion.
Enter fullscreen mode Exit fullscreen mode

Constraints like no company logos, duration, and not raising hands to face level were documented in docs/PRODUCTION.md. Codex read this on its own, fixing issues like overlapping text and missing face colors before returning the output.

It ran unattended. I wrote a production driver to handle one episode's workflow: "meeting → script and caricature → audio → lip sync → finishing → release prep." It ran continuously from episodes 15 to 38, taking about one hour per episode with zero failures or skips.

The problem was happening behind the scenes.

Actual usage

Codex logs each execution in ~/.codex/sessions as JSONL. Summing the token_count from the last entry of each session and categorizing by working directory yielded this:

Sessions Total Tokens
ai-boardroom 344 159,366,148
Other 5 projects 31 approx. 9,860,000

The breakdown for ai-boardroom is as follows:

Input   158,022,372
  Of which cached: 147,115,008 (approx. 93%)
Output     1,343,776
Enter fullscreen mode Exit fullscreen mode

Only 1.34 million tokens were output. 99% of the usage was input, with over 90% being cached input—essentially, resending the same context.

The model used for all executions was gpt-6-astra, never explicitly specified. The CLI default was used throughout.

Bar graph of Codex usage. Self-reported usage was approximately 7.4 million tokens, while actual usage was 159,366,148 tokens, with the majority (147,115,008) being cached input

Why I miscalculated by 20 times

I only counted this line that codex exec outputs at the end:

tokens used
84,294
Enter fullscreen mode Exit fullscreen mode

There were three issues with this approach:

1. Meetings spawned separate agents behind the scenes

When tasked with a 12-role meeting, Codex created separate agents for each role. The minutes for episode 10 actually state:

The production lead created separate agents for five roles, conducting a first round where each agent read the input, followed by a second round for script revision R2.

These child agents were logged as separate sessions, and their token usage didn't appear in the parent's tokens used. While I counted 101 sessions, there were actually 344.

2. Terminated executions don't output a summary line

Executions halted due to a 10-minute timeout or memory exhaustion don't output the final tokens used line. Only partial usage is recorded.

3. Resending everything with each move

The agent resends the entire conversation history and all files with each move. Since I had it read the script, production instructions, the last three episodes as references, and preview stills every time, the input size grew quadratically with session length. Even with caching, this counted toward usage.

Correct way to count usage

Instead of relying on CLI output, count from session logs:

import glob, json, os, collections

rows = collections.defaultdict(lambda: [0, 0])  # Sessions, Total Tokens
for f in glob.glob(os.path.expanduser("~/.codex/sessions/**/*.jsonl"), recursive=True):
    cwd, last = None, None
    for ln in open(f, encoding="utf-8", errors="replace"):
        j = json.loads(ln)
        p = j.get("payload", j)
        cwd = cwd or p.get("cwd")
        t = (p.get("info") or {}).get("total_token_usage")
        if t:
            last = t                    # Use the last value as it's cumulative
    if last:
        key = os.path.basename(cwd or "unknown")
        rows[key][0] += 1
        rows[key][1] += last["total_tokens"]

for k, (n, tok) in sorted(rows.items(), key=lambda x: -x[1][1]):
    print(f"{k:24s} {n:4d} sessions {tok:>13,}")
Enter fullscreen mode Exit fullscreen mode

total_token_usage is cumulative within a session, so only the last entry is used. Sub-agents also appear as individual sessions, ensuring nothing is missed.

Other automation pitfalls

False usage limit detection. Initially, the driver flagged "usage limit" in the output as a limit breach. When Codex read the driver's own source code, this string appeared in the output, halting it for 30 minutes. I fixed it to only check for the actual error line:

if re.search(r"^ERROR: You've hit your usage limit", text, re.M):
Enter fullscreen mode Exit fullscreen mode

Resuming the wrong session. When running two episodes concurrently, codex exec resume --last picks the last active session. If it belongs to the other episode, it edits the wrong script. I started tracking the session id: in the output and resumed by specifying the ID.

Image generation server holding memory. ComfyUI retained about 11GB of memory while idle. Running Codex simultaneously caused memory exhaustion, crashing Codex. I stopped ComfyUI when not generating images.

What to change

Choose models by task. Tasks like transcribing finalized dialogue into JSON or following set procedures to verify stills didn't require the heaviest model. Only planning meetings needed quality judgment. I'll specify models per task with codex exec -m <model>.

Avoid spawning sub-agents in meetings. I'll have all 12 roles written within a single session, as separate contexts per role weren't necessary.

Limit files read. I stopped providing the last three episodes as references every time, including only what's needed for the task.

Estimate before running. Multiply tokens per episode by the number of episodes, comparing it with the remaining weekly limit before starting unattended execution. Each episode used about 5 million tokens; 13 episodes would require 65 million tokens.

Check remaining quota before claiming "not used." Minutes before hitting the limit, I stated, "The limit hasn't been reached." I only checked the latest log, not the remaining quota.

Conclusion

The idea of having a 12-person team wasn't wrong. Meetings caught exaggerations, script discrepancies were resolved through comparison, and videos were completed with just two lines of instructions.

What was missing was awareness of what powered it. Just as more team members mean higher labor costs, more agents mean more tokens. Holding a meeting incurs charges for the note-taker too.

Top comments (0)

Some comments may only be visible to logged-in visitors. Sign in to view all comments.