DEV Community

Cover image for What If Every AI Inference Came With a Transparent Impact Receipt?
CarbonLayer
CarbonLayer

Posted on

What If Every AI Inference Came With a Transparent Impact Receipt?

An AI API gives you the model’s answer. Often, it also gives you token counts. But when you’re running inference in production, you may also need to know: What did that request cost? How long did it take? And what energy, carbon, and water estimates can be associated with it?

That’s the developer experience I want CarbonLayer to help make possible: useful information attached to each inference, with enough context to understand what the numbers do—and don’t—mean.

A response might look something like this:

{
"response": "…",
"usage": {
"input_tokens": 184,
"output_tokens": 658,
"total_tokens": 842
},
"latency_ms": 731,
"impact": {
"cost_usd": {
"value": 0.0041,
"status": "estimated",
"basis": "token usage and applicable pricing"
},
"carbon_g": {
"value": 0.73,
"status": "modeled",
"basis": "estimated energy use and grid carbon intensity"
},
"water_ml": {
"value": 12.4,
"status": "modeled",
"basis": "estimated facility and electricity-related water use"
}
},
"methodology": {
"version": "example",
"data_freshness": "example",
"confidence": {
"cost": "example",
"carbon": "example",
"water": "example"
}
}
}
This is an illustrative payload—not a live CarbonLayer response, and the values are examples. The shape matters because a number without its status and basis is easy to misread.

A token count may come directly from the model provider’s response. Latency can be observed at the point where the request passes through a routing layer—but that’s not necessarily the same as the model’s own processing time. Cost might be calculated from usage and a pricing schedule; it may not include every infrastructure or business cost.

Carbon and water require even more context. In many systems, those figures are modeled estimates, not measurements taken for that individual request. A model may estimate energy use, then combine it with data such as grid carbon intensity. A water estimate may depend on assumptions about cooling and electricity generation. The result is only meaningful if developers can see the system boundary, data sources, assumptions, and uncertainty behind it.

There’s another complication: inference doesn’t always happen as one isolated request on one isolated machine. Workloads can be batched, hardware is shared, and providers may expose only some of the information needed to attribute resource use. A per-request figure can still be useful—but it may represent an allocation, not a direct reading from a meter attached to that call.

That’s why I’d want each field to carry its own label. “Measured,” “modeled,” “estimated,” and “unavailable” shouldn’t be interchangeable—and one label shouldn’t automatically apply to every metric in the response. If the region is unknown, the data is stale, or a calculation depends on a broad assumption, the receipt should say so.

The goal isn’t to make uncertain numbers look precise. It’s to give developers enough information to decide how to use them: compare workloads, spot trends, evaluate routing choices, or decide that a metric isn’t reliable enough for a particular decision.

For an impact receipt to be useful in production, I think it needs to answer a few basic questions:

What does each number represent?
Which parts were observed, and which were modeled?
What boundary and assumptions were used?
How fresh is the underlying data?
How uncertain is the estimate?
Which methodology version produced it?
That’s the direction I’m exploring with CarbonLayer: making inference economics and environmental estimates more transparent, one request at a time.

If you’re building AI systems, what would an inference response need to show before you’d trust its cost, carbon, or water figures?

Top comments (0)