Part 3 — Designing for provenance, uncertainty, and meaningful environmental data.
The Impact Receipt Series — Part 3
Many AI APIs can report how many tokens a request used and how long it took.
But imagine they also returned an environmental impact figure.
Would that number be enough?
Not even close.
A carbon estimate without a methodology is difficult to interpret. A water figure without a measurement boundary can be misleading. A region name without a clear connection to the actual execution location may tell us less than we think.
If an AI Impact Receipt is going to be useful, it needs more than numbers.
It needs context, provenance, and an honest account of uncertainty.
First, what is an Impact Receipt?
In the earlier Part 2 article, I explored a prototype receipt contract with modeled environmental estimates.
That earlier prototype demonstrated a possible response structure; it did not establish real per-inference measurements.
Now comes the harder design question:
What information would make a receipt genuinely useful to a developer?
1. Request identity and model
A receipt should identify the request and distinguish the requested model from the model that actually executed it.
Possible fields include:
- Request ID
- Requested model
- Executed model, when verifiable
- Provider
- Timestamp
A model name supplied by the caller is not proof that a particular provider executed the request.
The receipt should make that distinction explicit.
2. Tokens and latency
Developers need to understand the computational work and performance associated with a request.
Useful fields might include:
- Input tokens
- Output tokens
- Total tokens
- End-to-end latency
- Provider-reported usage versus estimated usage
These values also need provenance.
A token budget is not the same as actual token consumption. A locally measured response time is not necessarily the same as provider processing time.
3. Cost
Cost is often the first infrastructure trade-off teams can evaluate.
A useful receipt might distinguish:
- Estimated request cost
- Provider-reported or billed cost, if available
- Currency
- Pricing source
- Pricing version or effective date
A calculated cost based on public token pricing should not be represented as a verified billing record.
4. Energy
Energy is where the measurement problem becomes especially difficult.
What exactly are we measuring?
The accelerator? The server? Networking? Cooling? An allocation of shared infrastructure?
A proposed receipt could include an energy value, its unit, the system boundary, and its evidence classification.
If energy cannot be credibly determined for an individual request, the receipt should say so.
Unavailable is a valid technical answer.
5. Carbon
A carbon figure requires more than an energy estimate.
It also depends on the electricity emissions factor and the accounting method.
A useful receipt should explain:
- The energy basis
- The carbon intensity source
- The geographic and temporal assumptions
- The accounting method
- Whether the result is measured, modeled, or estimated
Two calculations can produce different answers without either being a direct measurement of one request's emissions.
6. Water
Water deserves its own field and methodology.
It should not be treated as a simple conversion from carbon.
A receipt might need to distinguish direct cooling water, electricity-generation water, consumption versus withdrawal, and the boundary used in the estimate.
Where the inputs are missing, displaying an exact milliliter figure could create more confusion than clarity.
7. Location and grid information
Location matters, but it is easy to overstate what is known.
A provider region label does not necessarily establish the precise physical location where every part of a request was processed.
Any grid-intensity information should therefore include its source, timestamp, geographic resolution, and assumptions.
A receipt should never imply more location certainty than the evidence supports.
8. Methodology and uncertainty
This may be the most important section.
Every environmental value should have an evidence label:
Measured: Based on a defined measurement process, with a stated boundary.
Modeled: Calculated using an explicit model and assumptions.
Estimated: Approximated from available inputs or proxy information.
Unavailable: Insufficient evidence to report a defensible value.
These categories need documented definitions. They should not become marketing badges.
A mature receipt design should also identify methodology versions, input sources, freshness, and uncertainty where that uncertainty can be meaningfully characterized.
A possible receipt structure
The following JSON is a conceptual schema example, not an output from CarbonLayer's current API or Playground.
{
"receiptVersion": "proposal-0.1",
"request": {
"id": "example-request",
"modelRequested": "example-model",
"modelExecuted": null,
"provider": null
},
"usage": {
"inputTokens": null,
"outputTokens": null,
"source": "unavailable"
},
"latency": {
"milliseconds": null,
"source": "unavailable"
},
"cost": {
"amount": null,
"currency": "USD",
"source": "unavailable"
},
"environment": {
"energyKwh": null,
"carbonGrams": null,
"waterMl": null,
"source": "unavailable"
},
"methodology": {
"version": null,
"measurementBoundary": null,
"uncertainty": null
}
}
Notice how many values are null.
That's intentional.
A trustworthy schema should allow us to represent missing evidence without inventing precision.
Where CarbonLayer stands today
CarbonLayer's public API prototype returns simulated receipts.
Its separate cost-and-latency Playground explores cost and latency through model-based simulation. It also has a single-request mode that, when provider credentials are configured, can make one direct provider call. That call is not production routing.
Real per-inference environmental impact data remains unavailable.
The schema above is a direction for discussion, not a claim that these fields are already implemented.
The bigger challenge: making receipts comparable
Even if two providers returned environmental receipts tomorrow, would their numbers be comparable?
Only if we understood their boundaries, sources, definitions, and methods.
Otherwise, one provider might report accelerator energy while another includes broader facility overhead. Both might label the result "energy per inference," despite describing different things.
That's why I think the next question isn't simply which fields to add.
It's whether the industry can agree on what those fields mean.
Next in the series: Can an AI Impact Receipt Be Standardized?
In Part 4, I'll explore what a common specification might require, which definitions need agreement, and how developers could help shape a useful contract.
Because the goal isn't just to generate a receipt.
It's to make that receipt understandable, inspectable, and eventually interoperable.
Developer question: If you were designing an Impact Receipt API, which field would you insist on including—and which field would you refuse to trust without more evidence?
Explore CarbonLayer: https://carbonlayer.polsia.io
Top comments (3)
carbonlayer, this is a deeply necessary conversation. as someone building a constraint-driven ai app (koda) and obsessing over token routing between 20b and 120b models, i see the value in transparent compute tracking every day. 🐯
if i were designing this api, the field i would insist on including is
methodology.uncertaintyalongsidemodelExecuted(distinct frommodelRequested). knowing exactly which model actually ran, paired with the confidence interval of the environmental estimate, is the only way to prevent greenwashing and enable real, actionable auditing.the field i would refuse to trust without more evidence is
environment.carbonGramsorwaterMlif they lack a strict, documentedmeasurementBoundaryandgridIntensitySource. a single number without its system boundary is just marketing fluff, not engineering data. as you brilliantly pointed out, "unavailable" is a much more honest and useful technical answer than a fabricated, precise decimal.fantastic, highly principled breakdown. looking forward to part 4 on standardization! 🌍🛡️
Some comments may only be visible to logged-in visitors. Sign in to view all comments.