An AI API usually gives you an answer. I wanted to explore what it would look like if it also returned a record of the request’s environmental impact—and how those figures were produced.
That’s the idea behind the Impact Receipt: impact information should travel with the AI request, not live in a separate dashboard with no link to the work that generated it.
In Part 1, I argued that carbon should travel with every AI request. Here’s what I’ve built so far—and what this first version does not do.
The prototype flow
A caller supplies a model or routing identifier, a prompt, a token budget, and a latency preference. The prototype returns dispatch information alongside environmental estimates.
There’s an important boundary: it does not call Claude or another language-model provider. The model identifier is simulated, and the response echoes the supplied prompt. This prototype tests the shape of the dispatch and receipt—not a real model call, provider-reported usage, or physical measurement.
Here’s a shortened example of the response shape:
{
"id": "",
"model": "claude-sonnet",
"tokens": 800,
"device": {
"region": "eu-north-1",
"country": "Norway",
"carbonIntensity": 68,
"renewableMix": 0.97,
"waterIntensity": 0.32,
"carbonIntensitySource": "modeled",
"waterIntensitySource": "modeled"
},
"energyKwh": 0.216045,
"carbonGrams": 14.691,
"waterMl": 69.134,
"carbonSource": "modeled",
"energySource": "modeled",
"waterSource": "modeled",
"source": "synthetic"
}
These are example prototype values—not meter readings for an individual inference.
The label matters as much as the number
A field like "carbonGrams": 14.691 looks precise. Without context, a developer could read it as a physical measurement of what that request emitted.
It isn’t.
In this prototype, the energy, carbon, and water figures are modeled allocations based on modeled or seeded infrastructure data. They are not measured at the level of an individual request. The tokens field is dispatch data, not provider-reported usage, because no model provider is executing the request.
That’s why the response labels the metrics as "modeled". The root-level "source": "synthetic" describes the data path; it does not turn the estimates into measurements.
A number needs provenance before someone can decide how much to trust it.
What the fields are for
The dispatch ID gives the record an identifier. The model field records the routing label used by the prototype; it does not prove which provider ran the request. The device and region fields add location context, which matters because energy, carbon intensity, water, and latency can vary by location.
Energy, carbon, and water are separate metrics, each with its own source label. A modeled estimate can still be useful—but only if it’s clear that it’s modeled, and clear about the assumptions behind it.
The ID is an identifier in this response. I’m not claiming that the prototype already provides a way to retrieve a saved receipt later.
What this first version proves—and what it doesn’t
This is an early prototype of the receipt contract, not a complete production feature. It shows one way a dispatch response can carry environmental estimates and provenance together.
It does not yet demonstrate a real provider call, per-inference physical metering, or a verified record of what a live model processed. It also doesn’t include a cost figure, confidence labels, or a methodology version.
Those omissions matter. A fuller receipt will need to say not only what the figures are, but also what boundaries and data sources produced them, how fresh the inputs are, and where uncertainty remains.
The first lesson from building this prototype is simple: an Impact Receipt isn’t just a set of numbers. It’s numbers plus context and provenance. Without those, precision can look like certainty. With them, even a modeled estimate can be interpreted honestly.
Next: what should a complete Impact Receipt contain?
Top comments (2)
What would an inference response need to show before you'd trust its cost, carbon, or water figures?
Before I'd trust it, I'd want the allocation methodology documented and the carbon intensity source timestamped, since grid intensity varies by hour. Leading with modeled vs measured is the right call. The dispatch ID is good for auditability. What's the plan for getting from modeled to measured?