Prepared with AI assistance using my work records and a cost-estimate report. The cost comparisons are estimates.
I went back through five days of work done with AI tools and tried to put a price on it, next to what the same work might have cost without generative AI. This post walks through how I did that and what the numbers do and don't say.
A word on framing: everything below is an estimate. None of it is a measured saving. The part I think carries over to other projects is the method: describe exactly what changed, attach evidence, and count human review time before comparing costs. A simple task list is enough. You don't need to build a complicated AI system first.
What actually happened
I reviewed work from October 5–9, 2026, in Japan Standard Time. The report groups it into 35 task-record units across software changes, research, documentation, media preparation, maintenance, and administration.
Those 35 units differ a lot in size and value. Some are improvements to existing applications. Others are documents or small maintenance jobs. A count like this isn't a productivity score, and it definitely doesn't mean 35 new products were built.
Three examples:
- An existing memory app received improvements to display, storage, synchronization, backup, and search, plus documentation.
- A voice application received recording, playback, and interface repairs, with Windows device checks recorded. Several related fixes were counted together.
- A cloud experiment produced one real answer from a small AI model through a web interface. The temporary server was deleted afterward. That completed the experiment's stated stage; a continuing service was outside its scope.
For each entry I needed one sentence describing where “done” ends. Without that, a small successful test can quietly turn into a claim about a whole system.
Evidence and limits go together
Where public code changes exist, the report links them. Other entries rely on work records or my own completion confirmation. Those are different kinds of evidence, so they're labeled differently.
A few limits worth spelling out:
- Checking a code-change record does not rerun the software.
- A recorded Windows check does not show that every device works.
- One browser-based game was confirmed in Chrome, while testing on a real Safari device remained incomplete.
If you want to try this on your own project, four fields per entry go a long way:
- The specific change or deliverable.
- The evidence available, such as a code-change link or a test result.
- The environments or conditions checked.
- What remains unfinished or unverified.
Related retries belong under the same outcome. Otherwise a hard repair inflates the task count just because it took several attempts.
How much human work is that?
Next question: how much work might experienced people need to reach the same recorded scope without generative AI?
The assumptions still allow existing repositories, ordinary scripts, templates, tools, and domain knowledge. Rebuilding every application from scratch would describe a different job.
The bottom-up estimates were about 180 hours (low), 310 (central), and 510 (high). The central calculation uses exactly 309 hours.
At an assumed 160 working hours per month, 309 hours is 309 / 160 ≈ 1.93 person-months. That is a unit conversion of the estimate, not a staffing plan or a delivery date.
These were retrospective judgments, not stopwatch measurements or observed hours saved. The range reflects different assumptions about reuse, investigation, and revision. It is not a statistical confidence interval.
I also didn't treat all waiting as labor. A download or an automated test takes time without someone watching it the whole way. Dependencies matter too: a demo has to work before it can be filmed. Total estimated labor can't promise a completion date.
The cost assumptions, spelled out
At a hypothetical JPY7,500 per hour, the central human-work scenario is:
309 hours × JPY7,500 = JPY2,317,500, or about JPY2.32 million.
That hourly rate is a planning assumption, not a verified market rate.
The AI-side reference amount is JPY5,209.92, about JPY5,210. It combines five days of allocated subscription payments with one RunPod prepaid-credit purchase. The credit purchase is money added to an account; actual computing consumption during the period is unknown.
The subscription figures come from the report's billing evidence and user-reported payments, not from current advertised prices. Some allocation details are still uncertain, including the ChatGPT billing cycle. The report also uses an assumed currency conversion where needed.
The most important caveat: that JPY5,210 leaves out my instructions, approvals, and review time, along with other costs. Comparing it directly with a complete human-work estimate would hide a major part of the workflow.
So the report adds a scenario. Suppose human management and review took 40 hours at JPY7,500 per hour. That is 40 / 160 = 0.25 person-months, and the combined amount becomes JPY305,209.92, about JPY305,210. That's roughly 87% below the central hypothetical human-work cost.
![Cost scenarios comparing the central human-work estimate with allocated AI payments and an assumed 40 hours of human oversight.
Figure rendered programmatically with AI assistance. These are planning scenarios with different cost boundaries. The 87% difference is not measured return on investment, and equal output quality has not been independently verified.
The low human-work scenario is JPY892,500, using both lower hours and a lower assumed rate. That smaller baseline matters: rework, mistakes, and extra supervision can eat the apparent advantage, and their actual costs weren't measured here.
A workflow you can check
If you want to try something similar, start with a bounded task: fix a reproducible problem, update a known document, or investigate a specific question. Define the acceptance check before you hand the work off.
![A human-led workflow in which a person defines the task, AI carries out bounded work, and a person reviews evidence before accepting the result.]
Figure rendered programmatically with AI assistance. Recommended division of responsibility. This diagram is not a claim that every task during the reviewed week followed an identical process.
Keep a person responsible for production changes, confidential information, safety, and decisions that need specialist knowledge. If you can't evaluate the result yourself, bring in someone who can.
What I take from this
With the caveats above in place, these estimates favor human-led AI assistance for work with a clear scope and a result you can check: a person keeps the requirements, decisions, and final acceptance, and AI does bounded work in between. That is the setup I'd recommend trying first.
The review also shows why careful records matter more than an impressive task count. For your own trial, record both the output and the human effort needed to make it usable. A small, well-documented comparison will tell you more about what to delegate next than a big number will. One week doesn't make a case for headcount reduction or unattended operation.
Source: Full report with task boundaries, evidence, and calculation assumptions, in Japanese.

Top comments (0)