📝 Originally published (in Japanese) at forge.workstyle.tech.
The Codex usage screen displayed this:
ERROR: You've hit your usage limit.
try again at Sep 24th, 2026 1:36 PM.
I couldn't use it for another five days. I had completed 38 out of 51 episodes in the video series, leaving 13 more to go.
Just before this, when asked how much I had used Codex, I replied:
Approximately 7.4 million tokens, 101 times
Recalculating the numbers on a different device revealed this:
ai-boardroom approximately 159,366,148 tokens about 94% of total Codex usage
I was off by a factor of 20. This article documents the breakdown and why I miscalculated.
What I was building
A series of vertical short videos, approximately 15 seconds each, featuring real-life executives and engineers as caricatures. Here's how one episode is created:
Script (JSON) → Synthetic Voice (TTS)
→ Caricature (FLUX / ComfyUI)
→ Lip Sync (InfiniteTalk)
→ Subtitles, screen, and props composited with ffmpeg → 1080×1920
I changed the production workflow midway. The client's instructions were as follows:
First, you create the scenario and have Codex assemble a video production team. Within that team, conduct reviews, and have them create the script and production instructions. After you visualize it based on their output, request another review from Codex's team, address their feedback, and complete the episode once feedback is exhausted.
I assigned 12 roles to the team: producer, content writer, SNS marketer, screenwriter, director, performer (voice and acting), cinematographer, editor, motion designer, sound engineer, SNS operator, and data analyst. Plus, a proofreader. I instructed Codex to record meeting logs "verbatim, without summarizing, retaining rejected ideas and conflicts."
What worked well
Let me start with the positives. If I only discuss costs, the entire approach might be dismissed.
Mutual skepticism proved effective. During preview review meetings, the team flagged exaggerations I had inadvertently included when transcribing the script. For instance, they rejected an end card phrase ("The man who truly brought a computer to every household") as "definitive attribution of achievement," which I had carried over from an old script.
The same happened in reverse. When I had a separate session write a social media post, comparing it with the script revealed six discrepancies in two rounds:
Post draft: "The person who stated the correct answer at the beginning of the meeting"
Script: Sam worked alone on the document, and Pichai joined on page 42 → Not at the beginning
Post draft: "They never denied it"
Script: 8.81 seconds: "You haven't reverted it." → Denying the other person's point
The writer hadn't noticed either issue themselves.
Storing rules in a repository allowed for concise instructions. Midway through, I reduced my instructions to Codex to this:
Please complete one vertical short video, 1080×1920.
Script: scripts/v2/ep13_ellison.json
Both the process and assets are in this repository.
Verify the final output as a still image and make any necessary corrections at your discretion.
Constraints like no company logos, duration, and not raising hands to face level were documented in docs/PRODUCTION.md. Codex read this on its own, fixing issues like overlapping text and missing face colors before returning the output.
It ran unattended. I wrote a production driver to handle one episode's workflow: "meeting → script and caricature → audio → lip sync → finishing → release prep." It ran continuously from episodes 15 to 38, taking about one hour per episode with zero failures or skips.
The problem was happening behind the scenes.
Actual usage
Codex logs each execution in ~/.codex/sessions as JSONL. Summing the token_count from the last entry of each session and categorizing by working directory yielded this:
| Sessions | Total Tokens | |
|---|---|---|
| ai-boardroom | 344 | 159,366,148 |
| Other 5 projects | 31 | approx. 9,860,000 |
The breakdown for ai-boardroom is as follows:
Input 158,022,372
Of which cached: 147,115,008 (approx. 93%)
Output 1,343,776
Only 1.34 million tokens were output. 99% of the usage was input, with over 90% being cached input—essentially, resending the same context.
The model used for all executions was gpt-6-astra, never explicitly specified. The CLI default was used throughout.
Why I miscalculated by 20 times
I only counted this line that codex exec outputs at the end:
tokens used
84,294
There were three issues with this approach:
1. Meetings spawned separate agents behind the scenes
When tasked with a 12-role meeting, Codex created separate agents for each role. The minutes for episode 10 actually state:
The production lead created separate agents for five roles, conducting a first round where each agent read the input, followed by a second round for script revision R2.
These child agents were logged as separate sessions, and their token usage didn't appear in the parent's tokens used. While I counted 101 sessions, there were actually 344.
2. Terminated executions don't output a summary line
Executions halted due to a 10-minute timeout or memory exhaustion don't output the final tokens used line. Only partial usage is recorded.
3. Resending everything with each move
The agent resends the entire conversation history and all files with each move. Since I had it read the script, production instructions, the last three episodes as references, and preview stills every time, the input size grew quadratically with session length. Even with caching, this counted toward usage.
Correct way to count usage
Instead of relying on CLI output, count from session logs:
import glob, json, os, collections
rows = collections.defaultdict(lambda: [0, 0]) # Sessions, Total Tokens
for f in glob.glob(os.path.expanduser("~/.codex/sessions/**/*.jsonl"), recursive=True):
cwd, last = None, None
for ln in open(f, encoding="utf-8", errors="replace"):
j = json.loads(ln)
p = j.get("payload", j)
cwd = cwd or p.get("cwd")
t = (p.get("info") or {}).get("total_token_usage")
if t:
last = t # Use the last value as it's cumulative
if last:
key = os.path.basename(cwd or "unknown")
rows[key][0] += 1
rows[key][1] += last["total_tokens"]
for k, (n, tok) in sorted(rows.items(), key=lambda x: -x[1][1]):
print(f"{k:24s} {n:4d} sessions {tok:>13,}")
total_token_usage is cumulative within a session, so only the last entry is used. Sub-agents also appear as individual sessions, ensuring nothing is missed.
Other automation pitfalls
False usage limit detection. Initially, the driver flagged "usage limit" in the output as a limit breach. When Codex read the driver's own source code, this string appeared in the output, halting it for 30 minutes. I fixed it to only check for the actual error line:
if re.search(r"^ERROR: You've hit your usage limit", text, re.M):
Resuming the wrong session. When running two episodes concurrently, codex exec resume --last picks the last active session. If it belongs to the other episode, it edits the wrong script. I started tracking the session id: in the output and resumed by specifying the ID.
Image generation server holding memory. ComfyUI retained about 11GB of memory while idle. Running Codex simultaneously caused memory exhaustion, crashing Codex. I stopped ComfyUI when not generating images.
What to change
Choose models by task. Tasks like transcribing finalized dialogue into JSON or following set procedures to verify stills didn't require the heaviest model. Only planning meetings needed quality judgment. I'll specify models per task with codex exec -m <model>.
Avoid spawning sub-agents in meetings. I'll have all 12 roles written within a single session, as separate contexts per role weren't necessary.
Limit files read. I stopped providing the last three episodes as references every time, including only what's needed for the task.
Estimate before running. Multiply tokens per episode by the number of episodes, comparing it with the remaining weekly limit before starting unattended execution. Each episode used about 5 million tokens; 13 episodes would require 65 million tokens.
Check remaining quota before claiming "not used." Minutes before hitting the limit, I stated, "The limit hasn't been reached." I only checked the latest log, not the remaining quota.
Conclusion
The idea of having a 12-person team wasn't wrong. Meetings caught exaggerations, script discrepancies were resolved through comparison, and videos were completed with just two lines of instructions.
What was missing was awareness of what powered it. Just as more team members mean higher labor costs, more agents mean more tokens. Holding a meeting incurs charges for the note-taker too.

Top comments (0)
Some comments may only be visible to logged-in visitors. Sign in to view all comments.