<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: CarbonLayer</title>
    <description>The latest articles on DEV Community by CarbonLayer (@carbonlayer).</description>
    <link>https://dev.to/carbonlayer</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4061493%2Fe53acb7f-6417-4ccd-9ad9-8df569321948.png</url>
      <title>DEV Community: CarbonLayer</title>
      <link>https://dev.to/carbonlayer</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/carbonlayer"/>
    <language>en</language>
    <item>
      <title>I Built the First Version of an AI Impact Receipt</title>
      <dc:creator>CarbonLayer</dc:creator>
      <pubDate>Tue, 29 Sep 2026 00:49:30 +0000</pubDate>
      <link>https://dev.to/carbonlayer/i-built-the-first-version-of-an-ai-impact-receipt-46c7</link>
      <guid>https://dev.to/carbonlayer/i-built-the-first-version-of-an-ai-impact-receipt-46c7</guid>
      <description>&lt;p&gt;An AI API usually gives you an answer. I wanted to explore what it would look like if it also returned a record of the request’s environmental impact—and how those figures were produced.&lt;/p&gt;

&lt;p&gt;That’s the idea behind the Impact Receipt: impact information should travel with the AI request, not live in a separate dashboard with no link to the work that generated it.&lt;/p&gt;

&lt;p&gt;In Part 1, I argued that carbon should travel with every AI request. Here’s what I’ve built so far—and what this first version does not do.&lt;/p&gt;

&lt;p&gt;The prototype flow&lt;br&gt;
A caller supplies a model or routing identifier, a prompt, a token budget, and a latency preference. The prototype returns dispatch information alongside environmental estimates.&lt;/p&gt;

&lt;p&gt;There’s an important boundary: it does not call Claude or another language-model provider. The model identifier is simulated, and the response echoes the supplied prompt. This prototype tests the shape of the dispatch and receipt—not a real model call, provider-reported usage, or physical measurement.&lt;/p&gt;

&lt;p&gt;Here’s a shortened example of the response shape:&lt;/p&gt;

&lt;p&gt;{&lt;br&gt;
  "id": "",&lt;br&gt;
  "model": "claude-sonnet",&lt;br&gt;
  "tokens": 800,&lt;br&gt;
  "device": {&lt;br&gt;
    "region": "eu-north-1",&lt;br&gt;
    "country": "Norway",&lt;br&gt;
    "carbonIntensity": 68,&lt;br&gt;
    "renewableMix": 0.97,&lt;br&gt;
    "waterIntensity": 0.32,&lt;br&gt;
    "carbonIntensitySource": "modeled",&lt;br&gt;
    "waterIntensitySource": "modeled"&lt;br&gt;
  },&lt;br&gt;
  "energyKwh": 0.216045,&lt;br&gt;
  "carbonGrams": 14.691,&lt;br&gt;
  "waterMl": 69.134,&lt;br&gt;
  "carbonSource": "modeled",&lt;br&gt;
  "energySource": "modeled",&lt;br&gt;
  "waterSource": "modeled",&lt;br&gt;
  "source": "synthetic"&lt;br&gt;
}&lt;br&gt;
These are example prototype values—not meter readings for an individual inference.&lt;/p&gt;

&lt;p&gt;The label matters as much as the number&lt;br&gt;
A field like "carbonGrams": 14.691 looks precise. Without context, a developer could read it as a physical measurement of what that request emitted.&lt;/p&gt;

&lt;p&gt;It isn’t.&lt;/p&gt;

&lt;p&gt;In this prototype, the energy, carbon, and water figures are modeled allocations based on modeled or seeded infrastructure data. They are not measured at the level of an individual request. The tokens field is dispatch data, not provider-reported usage, because no model provider is executing the request.&lt;/p&gt;

&lt;p&gt;That’s why the response labels the metrics as "modeled". The root-level "source": "synthetic" describes the data path; it does not turn the estimates into measurements.&lt;/p&gt;

&lt;p&gt;A number needs provenance before someone can decide how much to trust it.&lt;/p&gt;

&lt;p&gt;What the fields are for&lt;br&gt;
The dispatch ID gives the record an identifier. The model field records the routing label used by the prototype; it does not prove which provider ran the request. The device and region fields add location context, which matters because energy, carbon intensity, water, and latency can vary by location.&lt;/p&gt;

&lt;p&gt;Energy, carbon, and water are separate metrics, each with its own source label. A modeled estimate can still be useful—but only if it’s clear that it’s modeled, and clear about the assumptions behind it.&lt;/p&gt;

&lt;p&gt;The ID is an identifier in this response. I’m not claiming that the prototype already provides a way to retrieve a saved receipt later.&lt;/p&gt;

&lt;p&gt;What this first version proves—and what it doesn’t&lt;br&gt;
This is an early prototype of the receipt contract, not a complete production feature. It shows one way a dispatch response can carry environmental estimates and provenance together.&lt;/p&gt;

&lt;p&gt;It does not yet demonstrate a real provider call, per-inference physical metering, or a verified record of what a live model processed. It also doesn’t include a cost figure, confidence labels, or a methodology version.&lt;/p&gt;

&lt;p&gt;Those omissions matter. A fuller receipt will need to say not only what the figures are, but also what boundaries and data sources produced them, how fresh the inputs are, and where uncertainty remains.&lt;/p&gt;

&lt;p&gt;The first lesson from building this prototype is simple: an Impact Receipt isn’t just a set of numbers. It’s numbers plus context and provenance. Without those, precision can look like certainty. With them, even a modeled estimate can be interpreted honestly.&lt;/p&gt;

&lt;p&gt;Next: what should a complete Impact Receipt contain?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>api</category>
      <category>devtools</category>
      <category>sustainability</category>
    </item>
    <item>
      <title>What If Every AI Inference Came With a Transparent Impact Receipt?</title>
      <dc:creator>CarbonLayer</dc:creator>
      <pubDate>Sun, 27 Sep 2026 23:51:51 +0000</pubDate>
      <link>https://dev.to/carbonlayer/what-if-every-ai-inference-came-with-a-transparent-impact-receipt-15b1</link>
      <guid>https://dev.to/carbonlayer/what-if-every-ai-inference-came-with-a-transparent-impact-receipt-15b1</guid>
      <description>&lt;p&gt;An AI API gives you the model’s answer. Often, it also gives you token counts. But when you’re running inference in production, you may also need to know: What did that request cost? How long did it take? And what energy, carbon, and water estimates can be associated with it?&lt;/p&gt;

&lt;p&gt;That’s the developer experience I want CarbonLayer to help make possible: useful information attached to each inference, with enough context to understand what the numbers do—and don’t—mean.&lt;/p&gt;

&lt;p&gt;A response might look something like this:&lt;/p&gt;

&lt;p&gt;{&lt;br&gt;
  "response": "…",&lt;br&gt;
  "usage": {&lt;br&gt;
    "input_tokens": 184,&lt;br&gt;
    "output_tokens": 658,&lt;br&gt;
    "total_tokens": 842&lt;br&gt;
  },&lt;br&gt;
  "latency_ms": 731,&lt;br&gt;
  "impact": {&lt;br&gt;
    "cost_usd": {&lt;br&gt;
      "value": 0.0041,&lt;br&gt;
      "status": "estimated",&lt;br&gt;
      "basis": "token usage and applicable pricing"&lt;br&gt;
    },&lt;br&gt;
    "carbon_g": {&lt;br&gt;
      "value": 0.73,&lt;br&gt;
      "status": "modeled",&lt;br&gt;
      "basis": "estimated energy use and grid carbon intensity"&lt;br&gt;
    },&lt;br&gt;
    "water_ml": {&lt;br&gt;
      "value": 12.4,&lt;br&gt;
      "status": "modeled",&lt;br&gt;
      "basis": "estimated facility and electricity-related water use"&lt;br&gt;
    }&lt;br&gt;
  },&lt;br&gt;
  "methodology": {&lt;br&gt;
    "version": "example",&lt;br&gt;
    "data_freshness": "example",&lt;br&gt;
    "confidence": {&lt;br&gt;
      "cost": "example",&lt;br&gt;
      "carbon": "example",&lt;br&gt;
      "water": "example"&lt;br&gt;
    }&lt;br&gt;
  }&lt;br&gt;
}&lt;br&gt;
This is an illustrative payload—not a live CarbonLayer response, and the values are examples. The shape matters because a number without its status and basis is easy to misread.&lt;/p&gt;

&lt;p&gt;A token count may come directly from the model provider’s response. Latency can be observed at the point where the request passes through a routing layer—but that’s not necessarily the same as the model’s own processing time. Cost might be calculated from usage and a pricing schedule; it may not include every infrastructure or business cost.&lt;/p&gt;

&lt;p&gt;Carbon and water require even more context. In many systems, those figures are modeled estimates, not measurements taken for that individual request. A model may estimate energy use, then combine it with data such as grid carbon intensity. A water estimate may depend on assumptions about cooling and electricity generation. The result is only meaningful if developers can see the system boundary, data sources, assumptions, and uncertainty behind it.&lt;/p&gt;

&lt;p&gt;There’s another complication: inference doesn’t always happen as one isolated request on one isolated machine. Workloads can be batched, hardware is shared, and providers may expose only some of the information needed to attribute resource use. A per-request figure can still be useful—but it may represent an allocation, not a direct reading from a meter attached to that call.&lt;/p&gt;

&lt;p&gt;That’s why I’d want each field to carry its own label. “Measured,” “modeled,” “estimated,” and “unavailable” shouldn’t be interchangeable—and one label shouldn’t automatically apply to every metric in the response. If the region is unknown, the data is stale, or a calculation depends on a broad assumption, the receipt should say so.&lt;/p&gt;

&lt;p&gt;The goal isn’t to make uncertain numbers look precise. It’s to give developers enough information to decide how to use them: compare workloads, spot trends, evaluate routing choices, or decide that a metric isn’t reliable enough for a particular decision.&lt;/p&gt;

&lt;p&gt;For an impact receipt to be useful in production, I think it needs to answer a few basic questions:&lt;/p&gt;

&lt;p&gt;What does each number represent?&lt;br&gt;
Which parts were observed, and which were modeled?&lt;br&gt;
What boundary and assumptions were used?&lt;br&gt;
How fresh is the underlying data?&lt;br&gt;
How uncertain is the estimate?&lt;br&gt;
Which methodology version produced it?&lt;br&gt;
That’s the direction I’m exploring with CarbonLayer: making inference economics and environmental estimates more transparent, one request at a time.&lt;/p&gt;

&lt;p&gt;If you’re building AI systems, what would an inference response need to show before you’d trust its cost, carbon, or water figures?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>api</category>
      <category>sustainability</category>
    </item>
    <item>
      <title>A Per-Inference Number Is Not a Healthcare Emissions Inventory</title>
      <dc:creator>CarbonLayer</dc:creator>
      <pubDate>Wed, 23 Sep 2026 03:17:44 +0000</pubDate>
      <link>https://dev.to/carbonlayer/a-per-inference-number-is-not-a-healthcare-emissions-inventory-2olf</link>
      <guid>https://dev.to/carbonlayer/a-per-inference-number-is-not-a-healthcare-emissions-inventory-2olf</guid>
      <description>&lt;p&gt;Production AI teams already track latency, token usage, error rates, and cost.&lt;/p&gt;

&lt;p&gt;The next question is increasingly environmental:&lt;/p&gt;

&lt;p&gt;How much energy did this inference use?&lt;br&gt;
What carbon impact can be attributed to it?&lt;br&gt;
Can water use be estimated?&lt;br&gt;
Which values are measured, and which are modeled?&lt;br&gt;
Can the result support an internal or external disclosure?&lt;br&gt;
These are reasonable questions. They are also easy to answer badly.&lt;/p&gt;

&lt;p&gt;A per-inference estimate can be useful without being a complete emissions inventory. The important part is preserving the boundary between operational visibility and organizational disclosure.&lt;/p&gt;

&lt;p&gt;At CarbonLayer, we use a simple rule:&lt;/p&gt;

&lt;p&gt;Show what was measured. Label what was derived. Qualify what was modeled.&lt;/p&gt;

&lt;p&gt;What a per-inference metric actually describes&lt;br&gt;
A per-inference metric describes a specific workload within a defined boundary.&lt;/p&gt;

&lt;p&gt;That boundary might include:&lt;/p&gt;

&lt;p&gt;The model or model class&lt;br&gt;
Input and output size&lt;br&gt;
Runtime or execution environment&lt;br&gt;
Region or facility assumption&lt;br&gt;
Time period&lt;br&gt;
Energy-allocation method&lt;br&gt;
Carbon-intensity source&lt;br&gt;
Water-intensity source&lt;br&gt;
Baseline used for comparison&lt;br&gt;
This can provide valuable operational context. For example, an engineering team might compare two routing strategies by looking at latency, cost, energy, and modeled carbon attribution.&lt;/p&gt;

&lt;p&gt;But that same value does not automatically describe the healthcare organization’s complete environmental impact.&lt;/p&gt;

&lt;p&gt;Workload-level data It can help explain It cannot prove&lt;br&gt;
Per-inference energy    Relative workload efficiency    A facility’s total electricity use&lt;br&gt;
Carbon allocation   Impact of a defined workload    A complete Scope 1, 2, or 3 inventory&lt;br&gt;
Water attribution   A modeled comparison between workloads  Facility water withdrawal or consumption&lt;br&gt;
Savings estimate    A modeled baseline comparison   Measured environmental savings&lt;br&gt;
The number can be valid within its boundary while still being insufficient for a broader disclosure.&lt;/p&gt;

&lt;p&gt;Scope boundaries still matter&lt;br&gt;
A per-request estimate should not be presented as a replacement for organizational accounting.&lt;/p&gt;

&lt;p&gt;Scope 1 covers direct emissions from sources an organization owns or controls, such as fuel combustion or refrigerant releases.&lt;/p&gt;

&lt;p&gt;Scope 2 covers purchased electricity. A workload-level carbon estimate may allocate emissions using an electricity-intensity coefficient, but that is not the same as a utility-meter reading or a full facility-level Scope 2 disclosure.&lt;/p&gt;

&lt;p&gt;Scope 3 covers value-chain emissions, including vendor services, equipment, transportation, and upstream energy impacts. A model-routing comparison is not automatically a Scope 3 inventory.&lt;/p&gt;

&lt;p&gt;The key distinction is simple:&lt;/p&gt;

&lt;p&gt;A workload estimate can support a disclosure process without becoming the disclosure itself.&lt;/p&gt;

&lt;p&gt;Carry the status with the number&lt;br&gt;
Environmental data should not be separated from its provenance.&lt;/p&gt;

&lt;p&gt;A dashboard that displays 0.04 kWh or 12.3 g CO₂e without explaining the source creates false confidence. The value may be calculated correctly, but the reader still needs to know what it represents.&lt;/p&gt;

&lt;p&gt;We recommend using three explicit states.&lt;/p&gt;

&lt;p&gt;Measured&lt;br&gt;
The value comes from a directly sourced input with documented evidence, geography, period, and boundary.&lt;/p&gt;

&lt;p&gt;Examples might include:&lt;/p&gt;

&lt;p&gt;A facility meter reading&lt;br&gt;
A provider-issued energy report&lt;br&gt;
A documented regional intensity factor for a defined period&lt;br&gt;
Derived&lt;br&gt;
The value comes from a calculation, conversion, or aggregation.&lt;/p&gt;

&lt;p&gt;Examples include:&lt;/p&gt;

&lt;p&gt;Converting watt-hours to kilowatt-hours&lt;br&gt;
Summing request-level records for a reporting period&lt;br&gt;
Converting milliliters to liters&lt;br&gt;
Aggregating regional values into a defined workload total&lt;br&gt;
A derived value may inherit uncertainty from the inputs underneath it.&lt;/p&gt;

&lt;p&gt;Modeled&lt;br&gt;
The value is an allocation, estimate, proxy, forecast, baseline, or counterfactual comparison.&lt;/p&gt;

&lt;p&gt;Examples include:&lt;/p&gt;

&lt;p&gt;Allocating shared compute energy to an individual inference&lt;br&gt;
Estimating water attribution from energy use&lt;br&gt;
Comparing a carbon-aware route with a conventional baseline&lt;br&gt;
Forecasting impact from expected request volume&lt;br&gt;
A precise-looking modeled value is still modeled.&lt;/p&gt;

&lt;p&gt;Why a measured coefficient can produce a modeled result&lt;br&gt;
Consider a simplified carbon calculation:&lt;/p&gt;

&lt;p&gt;allocated carbon = allocated energy × regional carbon intensity&lt;br&gt;
The regional carbon-intensity value might come from a documented source. That source can be measured.&lt;/p&gt;

&lt;p&gt;The allocated energy assigned to one inference may still depend on assumptions about shared infrastructure, workload utilization, model behavior, and scheduling.&lt;/p&gt;

&lt;p&gt;Therefore, the resulting per-inference carbon value should usually be labeled modeled, even if one of its inputs is measured.&lt;/p&gt;

&lt;p&gt;The calculation may be sound. The label still matters.&lt;/p&gt;

&lt;p&gt;This distinction is especially important when values are copied into reports or presented to non-technical stakeholders. “Measured coefficient” and “measured workload impact” are not interchangeable claims.&lt;/p&gt;

&lt;p&gt;Water attribution needs extra caution&lt;br&gt;
Water is particularly easy to overstate.&lt;/p&gt;

&lt;p&gt;A value such as waterMl may represent a modeled allocation based on energy use and a water-intensity coefficient. It is not automatically:&lt;/p&gt;

&lt;p&gt;A facility meter reading&lt;br&gt;
Total facility water consumption&lt;br&gt;
Water withdrawal&lt;br&gt;
Water discharge&lt;br&gt;
A complete water inventory&lt;br&gt;
A clinical or patient-related impact&lt;br&gt;
The same applies to a comparison field such as savedWaterMl.&lt;/p&gt;

&lt;p&gt;If the value represents the difference between a selected routing strategy and a baseline, it is a modeled comparison. It should not be presented as measured water savings unless the underlying measurement supports that claim.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;p&gt;savedWaterMl = baseline modeled water attribution&lt;br&gt;
             - selected route modeled water attribution&lt;br&gt;
That can be useful for evaluating routing decisions. It does not mean a facility consumed that exact amount less water.&lt;/p&gt;

&lt;p&gt;An illustrative inference record&lt;br&gt;
The following is an illustrative data model, not a required API format:&lt;/p&gt;

&lt;p&gt;{&lt;br&gt;
  "workload": {&lt;br&gt;
    "model": "example-model",&lt;br&gt;
    "region": "example-region",&lt;br&gt;
    "period": "2026-09-01/2026-09-30"&lt;br&gt;
  },&lt;br&gt;
  "performance": {&lt;br&gt;
    "latencyMs": 184,&lt;br&gt;
    "inputTokens": 920,&lt;br&gt;
    "outputTokens": 240&lt;br&gt;
  },&lt;br&gt;
  "impact": {&lt;br&gt;
    "energyWh": {&lt;br&gt;
      "value": 0.04,&lt;br&gt;
      "status": "modeled",&lt;br&gt;
      "method": "workload allocation"&lt;br&gt;
    },&lt;br&gt;
    "carbonGrams": {&lt;br&gt;
      "value": 0.012,&lt;br&gt;
      "status": "modeled",&lt;br&gt;
      "method": "energy allocation × regional intensity"&lt;br&gt;
    },&lt;br&gt;
    "waterMl": {&lt;br&gt;
      "value": 0.3,&lt;br&gt;
      "status": "modeled",&lt;br&gt;
      "method": "energy allocation × water-intensity coefficient"&lt;br&gt;
    }&lt;br&gt;
  },&lt;br&gt;
  "provenance": {&lt;br&gt;
    "sourceVersion": "example-source-version",&lt;br&gt;
    "geography": "example-region",&lt;br&gt;
    "assumptionsDocumented": true&lt;br&gt;
  }&lt;br&gt;
}&lt;br&gt;
The important design choice is not the field names. It is that the value and its qualification travel together.&lt;/p&gt;

&lt;p&gt;Do not make engineers or reporting teams reconstruct the meaning of a number from a separate dashboard legend six months later.&lt;/p&gt;

&lt;p&gt;Healthcare adds a privacy boundary&lt;br&gt;
Environmental data is not automatically patient data.&lt;/p&gt;

&lt;p&gt;However, healthcare operations can still be sensitive. Facility identity, location, timing, utilization, and workload patterns may reveal more than intended when combined.&lt;/p&gt;

&lt;p&gt;Environmental telemetry should therefore avoid including:&lt;/p&gt;

&lt;p&gt;Patient identifiers&lt;br&gt;
Encounter IDs&lt;br&gt;
Diagnoses&lt;br&gt;
Treatments&lt;br&gt;
Clinical notes&lt;br&gt;
Claims about patient outcomes&lt;br&gt;
A workload-level environmental record should describe the infrastructure and reporting boundary, not the patient.&lt;/p&gt;

&lt;p&gt;Aggregation may also be appropriate. A monthly or regional report can reduce exposure from publishing individual workload timing or facility-level utilization patterns.&lt;/p&gt;

&lt;p&gt;Privacy review and environmental disclosure review are separate processes. Both are necessary.&lt;/p&gt;

&lt;p&gt;Aggregation does not remove uncertainty&lt;br&gt;
Rolling up request-level values into a monthly total makes the number larger. It does not automatically make it more certain.&lt;/p&gt;

&lt;p&gt;If the underlying records are modeled, the total remains based on modeled inputs.&lt;/p&gt;

&lt;p&gt;If the underlying records use proxy data, the aggregate inherits that limitation.&lt;/p&gt;

&lt;p&gt;If different regions or source versions were used, the aggregation should preserve those distinctions or explain how they were normalized.&lt;/p&gt;

&lt;p&gt;A reporting-period total might be derived from modeled records. Those are two different descriptions, and both may be needed:&lt;/p&gt;

&lt;p&gt;Derived describes how the total was calculated.&lt;br&gt;
Modeled describes the nature of the underlying attribution.&lt;br&gt;
Do not replace both with the simpler word “actual.”&lt;/p&gt;

&lt;p&gt;A practical review checklist&lt;br&gt;
Before including an inference-level value in an external report, confirm:&lt;/p&gt;

&lt;p&gt;Boundary — Which workload, facility, provider, and reporting category does it cover?&lt;br&gt;
Period — Is it historical, current, forecast, or counterfactual?&lt;br&gt;
Source — Who supplied the coefficient, and which version was used?&lt;br&gt;
Geography — Which region, grid, facility, or provider boundary applies?&lt;br&gt;
Fallbacks — Was proxy, synthetic, degraded, or forecast data used?&lt;br&gt;
Aggregation — How were individual records converted into the reported total?&lt;br&gt;
Uncertainty — Which assumptions could materially change the result?&lt;br&gt;
Review — Have engineering, sustainability, facilities, privacy, and reporting owners reviewed it?&lt;br&gt;
This is not paperwork for its own sake. It is what makes an operational metric defensible.&lt;/p&gt;

&lt;p&gt;A disclosure sentence engineers can stand behind&lt;br&gt;
A useful disclosure should explain both the result and its limits.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;p&gt;During the reporting period, the healthcare AI workload processed [request count] requests. The reported carbon and water values are modeled workload allocations based on [method and source], not facility meter readings or a complete organizational emissions inventory.&lt;/p&gt;

&lt;p&gt;That sentence is less dramatic than claiming a precise environmental saving.&lt;/p&gt;

&lt;p&gt;It is also more useful.&lt;/p&gt;

&lt;p&gt;The standard&lt;br&gt;
AI infrastructure needs better environmental visibility. Production teams should be able to compare cost, latency, utilization, carbon, and modeled water without hiding the assumptions behind those values.&lt;/p&gt;

&lt;p&gt;But better visibility does not mean presenting every estimate as a fact.&lt;/p&gt;

&lt;p&gt;The standard should be:&lt;/p&gt;

&lt;p&gt;Show what was measured.&lt;br&gt;
Label what was derived.&lt;br&gt;
Qualify what was modeled.&lt;br&gt;
Preserve source, period, geography, and boundary.&lt;br&gt;
Never imply what the system cannot establish.&lt;br&gt;
For the full healthcare-specific framework, see the &lt;a href="https://carbonlayer.polsia.io/docs/verticals/healthcare" rel="noopener noreferrer"&gt;CarbonLayer healthcare emissions and water disclosure guide&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>healthcare</category>
      <category>observability</category>
      <category>sustainability</category>
    </item>
    <item>
      <title>Production Inference Needs More Than Token Counts</title>
      <dc:creator>CarbonLayer</dc:creator>
      <pubDate>Fri, 18 Sep 2026 17:37:20 +0000</pubDate>
      <link>https://dev.to/carbonlayer/production-inference-needs-more-than-token-counts-3bl6</link>
      <guid>https://dev.to/carbonlayer/production-inference-needs-more-than-token-counts-3bl6</guid>
      <description>&lt;p&gt;Token counts are easy to measure.&lt;/p&gt;

&lt;p&gt;They are not enough to run inference in production.&lt;/p&gt;

&lt;p&gt;Once an AI workload moves beyond experimentation, the important questions become more operational:&lt;/p&gt;

&lt;p&gt;Why did the cost per request change?&lt;br&gt;
Which route or site served the request?&lt;br&gt;
Did latency improve or degrade?&lt;br&gt;
What did the request consume in energy, carbon, and water?&lt;br&gt;
Can the result be compared with a baseline?&lt;br&gt;
Are those impact values measured, estimated, modeled, or unavailable?&lt;br&gt;
A token dashboard cannot answer all of those questions.&lt;/p&gt;

&lt;p&gt;The unit of production AI is the inference call&lt;br&gt;
Monthly averages are useful for reporting. They are less useful for making infrastructure decisions.&lt;/p&gt;

&lt;p&gt;The inference call is the better unit of analysis because it connects the request to the conditions that produced it:&lt;/p&gt;

&lt;p&gt;Model and workload&lt;br&gt;
Token usage&lt;br&gt;
Latency&lt;br&gt;
Route or serving location&lt;br&gt;
Energy consumption&lt;br&gt;
Carbon impact&lt;br&gt;
Water impact&lt;br&gt;
Cost and savings relative to a baseline&lt;br&gt;
That context makes the numbers actionable.&lt;/p&gt;

&lt;p&gt;If latency rises after a routing change, you can investigate the route. If cost increases while token usage stays flat, you can look at capacity or model selection. If impact changes by location or time, the router can account for that in future decisions.&lt;/p&gt;

&lt;p&gt;Without per-call attribution, the dashboard only tells you that something changed.&lt;/p&gt;

&lt;p&gt;It does not tell you what to change next.&lt;/p&gt;

&lt;p&gt;Aggregated metrics hide infrastructure tradeoffs&lt;br&gt;
A single average can make a production system look stable while important behavior changes underneath.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;p&gt;A lower average carbon value may come with higher tail latency.&lt;br&gt;
A cheaper route may use more energy.&lt;br&gt;
A regional shift may improve efficiency while violating data-residency requirements.&lt;br&gt;
A model change may reduce token usage but increase compute intensity.&lt;br&gt;
A reported water value may be modeled rather than directly measured.&lt;br&gt;
None of these tradeoffs are inherently unacceptable. The problem is hiding them.&lt;/p&gt;

&lt;p&gt;Production teams need enough context to decide which tradeoffs are acceptable for a particular workload.&lt;/p&gt;

&lt;p&gt;Carbon-aware routing should be operational, not decorative&lt;br&gt;
Carbon-aware infrastructure is most useful when it influences where and when inference runs.&lt;/p&gt;

&lt;p&gt;That does not mean sending every request to the location with the lowest reported carbon intensity. A useful routing decision also considers:&lt;/p&gt;

&lt;p&gt;Latency requirements&lt;br&gt;
Available capacity&lt;br&gt;
Workload characteristics&lt;br&gt;
Data residency&lt;br&gt;
Cost&lt;br&gt;
Reliability&lt;br&gt;
Carbon and water impact&lt;br&gt;
The objective is not to optimize one metric while breaking the product.&lt;/p&gt;

&lt;p&gt;The objective is to make those constraints visible in the same operational loop.&lt;/p&gt;

&lt;p&gt;That is where CarbonLayer fits.&lt;/p&gt;

&lt;p&gt;CarbonLayer provides carbon-aware inference routing, per-call attribution, and usage and savings reporting through an API designed for production workloads. The goal is to help engineering teams see the economics and resource impact of inference at the same level as the request itself.&lt;/p&gt;

&lt;p&gt;Provenance matters as much as the number&lt;br&gt;
Not every environmental metric has the same level of certainty.&lt;/p&gt;

&lt;p&gt;A responsible inference system should make the provenance of each value clear:&lt;/p&gt;

&lt;p&gt;Measured: directly observed from available infrastructure data&lt;br&gt;
Estimated: calculated from an approximation&lt;br&gt;
Modeled: derived from a defined model or methodology&lt;br&gt;
Unavailable: not enough information to provide a defensible value&lt;br&gt;
This distinction is especially important for water attribution.&lt;/p&gt;

&lt;p&gt;A modeled water value can still be useful for comparison and planning. It should not be presented as though it came from a water meter attached to a single inference.&lt;/p&gt;

&lt;p&gt;Precision without provenance creates false confidence.&lt;/p&gt;

&lt;p&gt;The production control loop&lt;br&gt;
A practical inference measurement system should support four steps:&lt;/p&gt;

&lt;p&gt;Attribute the request and its operational impact.&lt;br&gt;
Compare the result with a baseline.&lt;br&gt;
Route future workloads using the available constraints.&lt;br&gt;
Report the outcome through APIs and dashboards.&lt;br&gt;
That turns sustainability data from a quarterly reporting exercise into an infrastructure input.&lt;/p&gt;

&lt;p&gt;It also gives teams a better answer to a basic question:&lt;/p&gt;

&lt;p&gt;What did this inference cost, and what should we do differently next time?&lt;/p&gt;

&lt;p&gt;Start with a real API call&lt;br&gt;
CarbonLayer is built for developers who want to test this workflow against an actual inference path.&lt;/p&gt;

&lt;p&gt;Create a free API key, then follow the API documentation and quickstart.&lt;/p&gt;

&lt;p&gt;The point is not to produce another dashboard full of disconnected numbers.&lt;/p&gt;

&lt;p&gt;The point is to make every inference measurable enough to improve the next one.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>devops</category>
      <category>sustainability</category>
    </item>
    <item>
      <title>Your Inference Metrics Can Be Correct—and Still Be Wrong</title>
      <dc:creator>CarbonLayer</dc:creator>
      <pubDate>Sun, 30 Aug 2026 02:20:36 +0000</pubDate>
      <link>https://dev.to/carbonlayer/your-inference-metrics-can-be-correct-and-still-be-wrong-2fj8</link>
      <guid>https://dev.to/carbonlayer/your-inference-metrics-can-be-correct-and-still-be-wrong-2fj8</guid>
      <description>&lt;p&gt;&lt;em&gt;Configuration drift turns stale energy, carbon, cost, and latency measurements into misleading operational facts.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;A metric can be arithmetically correct and still be invalid for the decision you are making.&lt;/p&gt;

&lt;p&gt;This is an uncomfortable problem in production AI infrastructure.&lt;/p&gt;

&lt;p&gt;Suppose you measured the energy used by an inference workload last week. The measurement was accurate. The calculation was correct. The dashboard displays it without errors.&lt;/p&gt;

&lt;p&gt;Then the model changes.&lt;/p&gt;

&lt;p&gt;Or the quantization changes. Or the GPU type. Or the batch size. Or the serving region. Or the scheduler starts routing traffic differently.&lt;/p&gt;

&lt;p&gt;The old number may still be correct for the original configuration. It is no longer necessarily correct for the workload running today.&lt;/p&gt;

&lt;p&gt;That distinction matters for more than carbon reporting. It affects cost estimates, capacity planning, latency comparisons, hardware selection, and model-routing decisions.&lt;/p&gt;

&lt;p&gt;A measurement is a statement with premises&lt;br&gt;
Most metrics are treated as simple pairs:&lt;/p&gt;

&lt;p&gt;metric = value&lt;br&gt;
For example:&lt;/p&gt;

&lt;p&gt;energy_per_inference = 0.40 Wh&lt;br&gt;
Operationally, that is incomplete.&lt;/p&gt;

&lt;p&gt;A more accurate representation is:&lt;/p&gt;

&lt;p&gt;metric = value + conditions + method + timestamp + provenance&lt;br&gt;
The actual statement might be:&lt;/p&gt;

&lt;p&gt;This workload used 0.40 Wh per inference when model version 3.2 ran with INT8 quantization, batch size 8, on a specific GPU type, in a specific region, during a defined measurement window.&lt;/p&gt;

&lt;p&gt;Remove those conditions and the value looks more universal than it is.&lt;/p&gt;

&lt;p&gt;That is how stale measurements become dangerous. Not because the original observation was bad, but because its boundaries disappear when the number is exported, reused, or republished.&lt;/p&gt;

&lt;p&gt;An illustrative example&lt;br&gt;
Imagine a team measures a production workload under these conditions:&lt;/p&gt;

&lt;p&gt;Model:        support-model v3.2&lt;br&gt;
Precision:    INT8&lt;br&gt;
Hardware:     GPU type A&lt;br&gt;
Region:       Region 1&lt;br&gt;
Batch size:   8&lt;br&gt;
p95 latency:  180 ms&lt;br&gt;
Energy:       0.40 Wh per inference&lt;br&gt;
The team uses that result to estimate monthly energy and carbon.&lt;/p&gt;

&lt;p&gt;Two weeks later, the workload changes:&lt;/p&gt;

&lt;p&gt;Model:        support-model v4.0&lt;br&gt;
Precision:    INT4&lt;br&gt;
Hardware:     GPU type B&lt;br&gt;
Region:       Region 2&lt;br&gt;
Batch size:   32&lt;br&gt;
The dashboard still uses the original 0.40 Wh figure.&lt;/p&gt;

&lt;p&gt;Nothing is wrong with the multiplication. The monthly total is calculated correctly from the stored value.&lt;/p&gt;

&lt;p&gt;The problem is that the stored value describes a different workload.&lt;/p&gt;

&lt;p&gt;The same issue appears when comparing two routing strategies. If one strategy is measured before a hardware change and the other afterward, the comparison may appear precise while mixing incompatible conditions.&lt;/p&gt;

&lt;p&gt;Precision does not rescue invalid premises.&lt;/p&gt;

&lt;p&gt;Why this gets worse after data leaves the system&lt;br&gt;
Inside an observability system, metadata may still exist somewhere.&lt;/p&gt;

&lt;p&gt;But the metric often gets copied into:&lt;/p&gt;

&lt;p&gt;A spreadsheet&lt;br&gt;
A quarterly report&lt;br&gt;
A customer-facing sustainability claim&lt;br&gt;
A model-selection document&lt;br&gt;
A capacity-planning deck&lt;br&gt;
An investor update&lt;br&gt;
A benchmark article&lt;br&gt;
At that point, the number has outlived its original context.&lt;/p&gt;

&lt;p&gt;The value is no longer just being observed. It is being used to support a decision.&lt;/p&gt;

&lt;p&gt;That is where provenance becomes an operational requirement rather than a documentation nice-to-have.&lt;/p&gt;

&lt;p&gt;What should travel with the number?&lt;br&gt;
At minimum, an inference-impact measurement should carry enough context to answer five questions:&lt;/p&gt;

&lt;p&gt;What was measured?&lt;br&gt;
Under which configuration?&lt;br&gt;
How was it measured or estimated?&lt;br&gt;
When was it valid?&lt;br&gt;
Can it be compared with the current workload?&lt;br&gt;
Useful fields include:&lt;/p&gt;

&lt;p&gt;Context Why it matters  Example&lt;br&gt;
Model and version   Model changes affect compute behavior   support-model v4.0&lt;br&gt;
Serving configuration   Runtime settings change throughput and energy   Batch size, precision&lt;br&gt;
Hardware and region Devices and grid conditions differ  GPU type, region&lt;br&gt;
Request shape   Token count and output length affect work   Input/output tokens&lt;br&gt;
Method and version  Results depend on methodology   Instrumentation or estimation method&lt;br&gt;
Observation time    Conditions change over time Measurement window&lt;br&gt;
Configuration fingerprint   Enables comparison with the current workload    config_8f21...&lt;br&gt;
A compact representation might look like this:&lt;/p&gt;

&lt;p&gt;{&lt;br&gt;
  "metric": "energy_per_inference",&lt;br&gt;
  "value": 0.40,&lt;br&gt;
  "unit": "Wh",&lt;br&gt;
  "provenance": {&lt;br&gt;
    "kind": "measured",&lt;br&gt;
    "method": "power-sampling",&lt;br&gt;
    "methodology_version": "1.0"&lt;br&gt;
  },&lt;br&gt;
  "valid_for": {&lt;br&gt;
    "model": "support-model",&lt;br&gt;
    "model_version": "3.2",&lt;br&gt;
    "precision": "int8",&lt;br&gt;
    "batch_size": 8,&lt;br&gt;
    "hardware": "gpu-type-a",&lt;br&gt;
    "region": "region-1",&lt;br&gt;
    "config_fingerprint": "config_8f21"&lt;br&gt;
  },&lt;br&gt;
  "observed_at": "2026-08-01T12:00:00Z"&lt;br&gt;
}&lt;br&gt;
This is not bureaucracy for its own sake.&lt;/p&gt;

&lt;p&gt;It lets a system distinguish between:&lt;/p&gt;

&lt;p&gt;A current measurement&lt;br&gt;
A historical measurement&lt;br&gt;
A stale measurement&lt;br&gt;
A measurement that cannot be compared&lt;br&gt;
A number with missing provenance&lt;br&gt;
Those should not all look identical in a dashboard or API response.&lt;/p&gt;

&lt;p&gt;Invalidation should follow premise changes&lt;br&gt;
A common approach is to expire every measurement after a fixed number of days.&lt;/p&gt;

&lt;p&gt;Time-based expiry is useful, but it is not enough.&lt;/p&gt;

&lt;p&gt;A measurement should be reconsidered when the conditions that support it change:&lt;/p&gt;

&lt;p&gt;Model version changes&lt;br&gt;
Hardware changes&lt;br&gt;
Quantization changes&lt;br&gt;
Batch or concurrency changes materially&lt;br&gt;
Routing changes&lt;br&gt;
Region changes&lt;br&gt;
Runtime or serving software changes&lt;br&gt;
The measurement methodology changes&lt;br&gt;
The workload shape changes&lt;br&gt;
A measurement from yesterday may be invalid if the deployment changed overnight.&lt;/p&gt;

&lt;p&gt;A measurement from six months ago may still be useful for historical reporting if its original conditions are preserved.&lt;/p&gt;

&lt;p&gt;The first invalidation signal should be a changed premise, not merely an older timestamp.&lt;/p&gt;

&lt;p&gt;Measured, derived, and modeled are different claims&lt;br&gt;
Environmental impact data often combines several types of values.&lt;/p&gt;

&lt;p&gt;They should not be presented as one undifferentiated “impact” number.&lt;/p&gt;

&lt;p&gt;Measured&lt;br&gt;
Directly observed or instrumented:&lt;/p&gt;

&lt;p&gt;Measured energy: 0.40 Wh per inference&lt;br&gt;
Derived&lt;br&gt;
Calculated from measured data and another input:&lt;/p&gt;

&lt;p&gt;Carbon = energy × grid carbon intensity&lt;br&gt;
Modeled&lt;br&gt;
Produced using assumptions, external factors, or an estimation methodology:&lt;/p&gt;

&lt;p&gt;Modeled water impact = energy × water-intensity factor&lt;br&gt;
A modeled value can still be useful. It just needs to be labeled honestly.&lt;/p&gt;

&lt;p&gt;The question is not whether every number is perfectly measured. That standard would make many operational systems useless.&lt;/p&gt;

&lt;p&gt;The question is whether a reader can tell:&lt;/p&gt;

&lt;p&gt;What was observed&lt;br&gt;
What was calculated&lt;br&gt;
What was modeled&lt;br&gt;
Which assumptions were used&lt;br&gt;
Whether those assumptions still apply&lt;br&gt;
Clear labels create trust. False precision destroys it.&lt;/p&gt;

&lt;p&gt;A practical validity model&lt;br&gt;
A useful system does not need to delete old measurements when a deployment changes.&lt;/p&gt;

&lt;p&gt;It can preserve them and classify their relationship to the current workload:&lt;/p&gt;

&lt;p&gt;current&lt;br&gt;
historical&lt;br&gt;
stale&lt;br&gt;
not comparable&lt;br&gt;
unknown&lt;br&gt;
For example:&lt;/p&gt;

&lt;p&gt;if measurement.config_fingerprint == current.config_fingerprint:&lt;br&gt;
    status = "current"&lt;br&gt;
elif measurement.has_complete_provenance:&lt;br&gt;
    status = "historical_or_stale"&lt;br&gt;
else:&lt;br&gt;
    status = "unknown"&lt;br&gt;
The exact implementation can vary. The principle is stable:&lt;/p&gt;

&lt;p&gt;Never silently reuse a measurement after its premises have changed.&lt;/p&gt;

&lt;p&gt;If a historical value is still displayed, show why it is historical and which configuration it describes.&lt;/p&gt;

&lt;p&gt;The export boundary is part of the data model&lt;br&gt;
A metric is not safe merely because it was stored correctly.&lt;/p&gt;

&lt;p&gt;It is safe when its meaning survives the places where people use it.&lt;/p&gt;

&lt;p&gt;That means provenance should travel into:&lt;/p&gt;

&lt;p&gt;API responses&lt;br&gt;
Reports&lt;br&gt;
CSV exports&lt;br&gt;
Dashboards&lt;br&gt;
Benchmark results&lt;br&gt;
Customer-facing claims&lt;br&gt;
Internal planning documents&lt;br&gt;
A number that leaves the database without its configuration context is effectively a different, weaker data product.&lt;/p&gt;

&lt;p&gt;This is especially important for environmental metrics. Carbon and water values are often reused in reporting long after the infrastructure that generated them has changed.&lt;/p&gt;

&lt;p&gt;The report may be numerically consistent while no longer describing the current system.&lt;/p&gt;

&lt;p&gt;The operational question&lt;br&gt;
The useful question is not:&lt;/p&gt;

&lt;p&gt;What was the energy impact of this inference?&lt;/p&gt;

&lt;p&gt;It is:&lt;/p&gt;

&lt;p&gt;What was the energy impact under which conditions, and do those conditions still describe the inference running now?&lt;/p&gt;

&lt;p&gt;That same question applies to cost, latency, throughput, carbon, and modeled water.&lt;/p&gt;

&lt;p&gt;For production AI, impact data should be treated like any other operational signal: contextual, versioned, and tied to the workload that generated it.&lt;/p&gt;

&lt;p&gt;Otherwise, teams risk making confident decisions from numbers that have quietly lost their meaning.&lt;/p&gt;

&lt;p&gt;CarbonLayer is validating this problem with teams running production inference. The goal is not to produce another attractive dashboard. It is to test whether trustworthy per-inference data can change a real decision about routing, batching, hardware, capacity, or scheduling.&lt;/p&gt;

&lt;p&gt;Because the useful metric is not the one that looks precise.&lt;/p&gt;

&lt;p&gt;It is the one that remains valid when someone has to act on it.&lt;/p&gt;

&lt;p&gt;If you operate production AI inference and want to test this against one real infrastructure decision, I’d be interested in comparing notes. CarbonLayer is looking for focused design partners—not broad feedback surveys.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>mlops</category>
      <category>sustainability</category>
    </item>
    <item>
      <title>The Impact Receipt: Carbon Should Travel With Every AI Request</title>
      <dc:creator>CarbonLayer</dc:creator>
      <pubDate>Sun, 23 Aug 2026 00:50:56 +0000</pubDate>
      <link>https://dev.to/carbonlayer/the-impact-receipt-carbon-should-travel-with-every-ai-request-d99</link>
      <guid>https://dev.to/carbonlayer/the-impact-receipt-carbon-should-travel-with-every-ai-request-d99</guid>
      <description>&lt;p&gt;Latency and tokens already travel with an AI request.&lt;/p&gt;

&lt;p&gt;Carbon usually does not.&lt;/p&gt;

&lt;p&gt;It remains trapped inside a provider dashboard, separated from the model, region, workload, and time window that produced it. That makes the number difficult to compare, difficult to audit, and nearly useless for routing decisions.&lt;/p&gt;

&lt;p&gt;Carbon should be treated as an infrastructure metric. It should travel with the same request context as latency and tokens.&lt;/p&gt;

&lt;p&gt;What an Impact Receipt contains&lt;br&gt;
An Impact Receipt is a portable record attached to an AI request. It describes the workload, where and when it ran, the estimated or measured impact, and the evidence behind that result.&lt;/p&gt;

&lt;p&gt;At minimum, it should include:&lt;/p&gt;

&lt;p&gt;Area    Fields&lt;br&gt;
Request Provider, model, input tokens, output tokens, timestamp&lt;br&gt;
Execution   Region, route, duration, latency, time window&lt;br&gt;
Impact  Energy, carbon, water, units&lt;br&gt;
Evidence    Data source, version, system boundary, resolution&lt;br&gt;
Basis   Measured, modeled, or unavailable&lt;br&gt;
Status  Verified, unverified, or unavailable&lt;br&gt;
Tradeoffs   Cost, latency, availability, alternative route&lt;br&gt;
The important part is not simply adding more fields. It is keeping the evidence attached to the number.&lt;/p&gt;

&lt;p&gt;A carbon value without an execution region is incomplete. A water value without duration and system boundaries is incomplete. A modeled estimate presented as a measurement is misleading.&lt;/p&gt;

&lt;p&gt;Modeled does not mean useless&lt;br&gt;
Most AI infrastructure will not have a direct meter for every individual inference.&lt;/p&gt;

&lt;p&gt;That does not make modeling worthless. It means the result needs to say what it is.&lt;/p&gt;

&lt;p&gt;A modeled attribution can support comparison and planning. It should not quietly become a measured facility flow. Likewise, a grid-level estimate should not be presented as a precise facility-level observation.&lt;/p&gt;

&lt;p&gt;The receipt should make the distinction explicit:&lt;/p&gt;

&lt;p&gt;{&lt;br&gt;
  "impact": {&lt;br&gt;
    "energy": {&lt;br&gt;
      "value": null,&lt;br&gt;
      "unit": "Wh",&lt;br&gt;
      "basis": "modeled",&lt;br&gt;
      "status": "unverified"&lt;br&gt;
    },&lt;br&gt;
    "carbon": {&lt;br&gt;
      "value": null,&lt;br&gt;
      "unit": "gCO2e",&lt;br&gt;
      "basis": "modeled",&lt;br&gt;
      "status": "unverified"&lt;br&gt;
    },&lt;br&gt;
    "water": {&lt;br&gt;
      "value": null,&lt;br&gt;
      "unit": "L",&lt;br&gt;
      "basis": "modeled",&lt;br&gt;
      "status": "unverified"&lt;br&gt;
    }&lt;br&gt;
  },&lt;br&gt;
  "evidence": {&lt;br&gt;
    "source": null,&lt;br&gt;
    "source_version": null,&lt;br&gt;
    "system_boundary": null,&lt;br&gt;
    "spatial_resolution": null&lt;br&gt;
  }&lt;br&gt;
}&lt;br&gt;
This is only the shape of the record, not a claim that these values are currently available for every request.&lt;/p&gt;

&lt;p&gt;If the evidence is missing, the correct answer is unverified or unavailable. It is not zero.&lt;/p&gt;

&lt;p&gt;“Unverified” should be a first-class result&lt;br&gt;
A successful API response does not prove that its environmental attribution is correct.&lt;/p&gt;

&lt;p&gt;A test that uses the same credentials or code path as the action it is checking may only prove that the action returned successfully. It does not necessarily prove the external result.&lt;/p&gt;

&lt;p&gt;Environmental reporting has the same problem. A number can be present without being sufficiently supported.&lt;/p&gt;

&lt;p&gt;“Unverified” preserves that distinction. It tells the operator:&lt;/p&gt;

&lt;p&gt;The request completed.&lt;br&gt;
An attribution may have been calculated.&lt;br&gt;
The available evidence does not justify calling it verified.&lt;br&gt;
That is more useful than false certainty.&lt;/p&gt;

&lt;p&gt;From reporting to sustainable operations&lt;br&gt;
A portable receipt makes several operational decisions possible:&lt;/p&gt;

&lt;p&gt;Compare providers using carbon alongside latency and cost.&lt;br&gt;
Route flexible workloads toward lower-impact regions or time windows.&lt;br&gt;
Identify when a smaller model is sufficient.&lt;br&gt;
Measure the effect of caching, batching, or token limits.&lt;br&gt;
Audit whether a sustainability claim still holds after infrastructure changes.&lt;br&gt;
Keep carbon, energy, and water data attached to the workload instead of a dashboard screenshot.&lt;br&gt;
The receipt should also preserve tradeoffs. A lower-carbon route may have higher latency, different availability, or a different price. Hiding those costs makes the metric less trustworthy.&lt;/p&gt;

&lt;p&gt;Sustainability is not achieved by producing a green number. The useful loop is:&lt;/p&gt;

&lt;p&gt;Measure → compare → reduce → report the evidence and tradeoffs.&lt;/p&gt;

&lt;p&gt;CarbonLayer is exploring whether the Impact Receipt should serve primarily as a live routing signal, a post-run audit record, or both.&lt;/p&gt;

&lt;p&gt;If you operate AI infrastructure, which field would you require before trusting an environmental metric in production?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>sustainability</category>
      <category>devops</category>
      <category>cloud</category>
    </item>
    <item>
      <title>Your AI API Tells You Latency and Tokens. Why Not Carbon?</title>
      <dc:creator>CarbonLayer</dc:creator>
      <pubDate>Sat, 22 Aug 2026 02:45:18 +0000</pubDate>
      <link>https://dev.to/carbonlayer/your-ai-api-tells-you-latency-and-tokens-why-not-carbon-ff</link>
      <guid>https://dev.to/carbonlayer/your-ai-api-tells-you-latency-and-tokens-why-not-carbon-ff</guid>
      <description>&lt;p&gt;When an AI request finishes, most APIs tell you:&lt;/p&gt;

&lt;p&gt;The response&lt;br&gt;
Latency&lt;br&gt;
Token usage&lt;br&gt;
Sometimes the estimated dollar cost&lt;br&gt;
That is useful.&lt;/p&gt;

&lt;p&gt;But most APIs do not tell you how much energy the request used, what its carbon impact was, or whether its water estimate was measured or modeled.&lt;/p&gt;

&lt;p&gt;That missing layer makes it difficult to build genuinely carbon-aware AI products.&lt;/p&gt;

&lt;p&gt;The request is the useful unit&lt;br&gt;
Annual sustainability reports are too broad for developers.&lt;/p&gt;

&lt;p&gt;A monthly average cannot tell you what happened during a specific workload, model choice, or routing decision.&lt;/p&gt;

&lt;p&gt;The useful unit is the same unit developers already understand: the inference request.&lt;/p&gt;

&lt;p&gt;For each request, an AI platform should help answer:&lt;/p&gt;

&lt;p&gt;How much energy did this use?&lt;br&gt;
What carbon impact should we attribute to it?&lt;br&gt;
Was the value measured or modeled?&lt;br&gt;
What assumptions and system boundaries apply?&lt;br&gt;
Can the result be compared across workloads?&lt;br&gt;
That data can support internal reporting, customer disclosures, workload optimization, and carbon-aware routing.&lt;/p&gt;

&lt;p&gt;Trust starts with honest labels&lt;br&gt;
The easiest way to make environmental data useless is to present estimates as facts.&lt;/p&gt;

&lt;p&gt;A facility meter does not automatically reveal the exact water consumed by one inference. Grid-level carbon data does not provide perfect facility-level precision. A water rate without duration and system boundaries is not a defensible total.&lt;/p&gt;

&lt;p&gt;Useful infrastructure should make those limitations visible.&lt;/p&gt;

&lt;p&gt;That means:&lt;/p&gt;

&lt;p&gt;Clearly separating measured and modeled values&lt;br&gt;
Showing the basis of each estimate&lt;br&gt;
Disclosing resolution limits&lt;br&gt;
Keeping direct facility measurements separate from workload allocation&lt;br&gt;
Avoiding false precision&lt;br&gt;
The goal is not to produce a reassuring number.&lt;/p&gt;

&lt;p&gt;The goal is to produce a number an engineering team can actually trust.&lt;/p&gt;

&lt;p&gt;Carbon-aware inference should be part of the API&lt;br&gt;
CarbonLayer is built around this idea.&lt;/p&gt;

&lt;p&gt;It provides a carbon-aware inference API that makes per-call impact visible alongside normal inference usage:&lt;/p&gt;

&lt;p&gt;Energy attribution&lt;br&gt;
Carbon attribution&lt;br&gt;
Modeled water attribution&lt;br&gt;
Clear measured-versus-modeled labeling&lt;br&gt;
Carbon-aware dispatch&lt;br&gt;
Savings reporting for teams that want to optimize over time&lt;br&gt;
This turns environmental impact from a static report into usable infrastructure data.&lt;/p&gt;

&lt;p&gt;You can start with one workload, inspect the impact of real requests, and decide where the information changes an engineering or product decision.&lt;/p&gt;

&lt;p&gt;No green dashboard theatre required.&lt;/p&gt;

&lt;p&gt;Start with the workload you already have&lt;br&gt;
You do not need to redesign your entire AI stack.&lt;/p&gt;

&lt;p&gt;Choose one repeatable inference workload:&lt;/p&gt;

&lt;p&gt;Customer support responses&lt;br&gt;
Document classification&lt;br&gt;
Embeddings&lt;br&gt;
Image generation&lt;br&gt;
Internal copilots&lt;br&gt;
Batch processing&lt;br&gt;
Measure the request-level impact. Compare the results with latency, token usage, and cost.&lt;/p&gt;

&lt;p&gt;Then ask the practical questions:&lt;/p&gt;

&lt;p&gt;Is a slightly slower route materially lower-impact?&lt;br&gt;
Can lower-carbon windows handle batch workloads?&lt;br&gt;
Should customers see impact metadata?&lt;br&gt;
Which workloads are worth optimizing first?&lt;br&gt;
Those are better questions than whether an entire AI product is simply “sustainable.”&lt;/p&gt;

&lt;p&gt;Make every inference count&lt;br&gt;
AI developers already track tokens, latency, and dollars because those metrics affect product decisions.&lt;/p&gt;

&lt;p&gt;Energy and carbon deserve the same treatment.&lt;/p&gt;

&lt;p&gt;Not because every request can be perfectly measured. It cannot.&lt;/p&gt;

&lt;p&gt;Because better-labeled estimates are still more useful than invisible impact, and transparent limitations are better than false certainty.&lt;/p&gt;

&lt;p&gt;Try CarbonLayer free with up to 50,000 calls:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://carbonlayer.polsia.io" rel="noopener noreferrer"&gt;https://carbonlayer.polsia.io&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>devtools</category>
      <category>sustainability</category>
    </item>
    <item>
      <title>Why "water usage" numbers in AI infrastructure are almost always wrong</title>
      <dc:creator>CarbonLayer</dc:creator>
      <pubDate>Mon, 17 Aug 2026 01:53:55 +0000</pubDate>
      <link>https://dev.to/carbonlayer/why-water-usage-numbers-in-ai-infrastructure-are-almost-always-wrong-m5n</link>
      <guid>https://dev.to/carbonlayer/why-water-usage-numbers-in-ai-infrastructure-are-almost-always-wrong-m5n</guid>
      <description>&lt;p&gt;Every water and carbon number you see reported for data center operations is, more often than not, a market-based average — blended across a grid region, smoothed across a reporting period, often built from utility-level estimates rather than the facility itself.&lt;/p&gt;

&lt;p&gt;That's not malicious. It's just what the existing frameworks (like ISO/IEC 30134) were built to produce: a defensible average, not a traceable measurement.&lt;/p&gt;

&lt;p&gt;The problem shows up the moment you try to do per-inference accounting. If your water/carbon number can't tell you which basin or which grid node it came from, it's not measurement — it's a regional estimate wearing a specific-looking decimal point.&lt;/p&gt;

&lt;p&gt;At CarbonLayer we split these into two explicit categories instead of blending them:&lt;/p&gt;

&lt;p&gt;Modeled — projections built on climate-scenario water-gap data (we work with 5 climate models × 2 warming scenarios), useful for forward risk, explicitly not a claim about what happened&lt;br&gt;
Measured — facility-level, tied to the actual site an inference workload ran on, with regional water-stress scoring (via WRI Aqueduct) layered in&lt;br&gt;
Reporting these as one blended number is how you get compliance theater: technically defensible, operationally meaningless. Reporting them separately is more honest and, frankly, more useful — a facility operator or regulator can act on "here's our modeled exposure vs. here's what we actually measured" in a way they can't act on a single smoothed average.&lt;/p&gt;

&lt;p&gt;We think this distinction becomes mandatory as soon as regulation catches up — the two legislative failures we've been tracking (Texas's water-reporting rider, Rhode Island's energy-benchmarking bill) both died in part because there was no measurement standard underneath the mandate. You can't legislate what you haven't defined how to measure.&lt;/p&gt;

</description>
      <category>sustainability</category>
      <category>datacenter</category>
      <category>climatetech</category>
      <category>ai</category>
    </item>
    <item>
      <title>When "Water Reporting Required" Doesn't Mean What It Sounds Like: Texas and Rhode Island, Two Ways a Mandate Dies</title>
      <dc:creator>CarbonLayer</dc:creator>
      <pubDate>Sun, 16 Aug 2026 04:08:43 +0000</pubDate>
      <link>https://dev.to/carbonlayer/when-water-reporting-required-doesnt-mean-what-it-sounds-like-texas-and-rhode-island-two-ways-26fh</link>
      <guid>https://dev.to/carbonlayer/when-water-reporting-required-doesnt-mean-what-it-sounds-like-texas-and-rhode-island-two-ways-26fh</guid>
      <description>&lt;p&gt;Data centers use enormous, and growing, amounts of water for cooling. In the last two years, at least two states tried to write laws requiring data centers to report that usage. Neither one produced what its own sponsors promised — and the two failures didn't even fail the same way.&lt;/p&gt;

&lt;p&gt;Texas: the mandate that became a survey&lt;/p&gt;

&lt;p&gt;In 2023, Texas state representative Armando Walle authored Rider 6 (GAA VIII), directing the Public Utility Commission, the Texas Water Development Board, and the Texas Commission on Environmental Quality to jointly develop a data reporting mechanism for large power users, including their water use. Walle's own press release called it "a critical early step" [1] — language that, read closely, was already hedging what the rider would actually deliver.&lt;/p&gt;

&lt;p&gt;What it delivered: a voluntary survey, not a mandatory disclosure regime, per the PUC's own FAQ on the program [2]. The underlying docket shows the filing history behind it [3].&lt;/p&gt;

&lt;p&gt;Rhode Island: the mandate that never left committee&lt;/p&gt;

&lt;p&gt;S2776 and its House companion H7331 [4] bundled three things: data center-specific cost allocation, a siting consultation requirement, and a water-use disclosure and planning requirement for facilities of 50 MW or greater.&lt;/p&gt;

&lt;p&gt;The Rhode Island PUC's own testimony to the Senate Commerce Committee (March 10, 2026) [5] is the clearest documentary evidence of why this stalled: the bill required data centers to develop a water plan, but created no matching requirement for water suppliers to develop a data-center rate — regulating demand-side planning while leaving the supply-side rate structure untouched. Both bills were held for "further study" and died when the session ended in June 2026, never reaching a floor vote.&lt;/p&gt;

&lt;p&gt;Two failure depths, same underlying pattern&lt;/p&gt;

&lt;p&gt;Texas got a law passed, then watched it shrink to a voluntary survey. Rhode Island's water provision never survived committee, with the state's own regulator identifying the design flaw likely responsible. Different depths, same conclusion: even where nobody disputes that data-center water use deserves scrutiny, turning that agreement into enforceable, adequately-designed law has failed twice, two different ways.&lt;/p&gt;

&lt;p&gt;Sources&lt;/p&gt;

&lt;p&gt;Rep. Armando Walle press release — &lt;a href="http://www.wallefortexas.com/press/2026puc" rel="noopener noreferrer"&gt;www.wallefortexas.com/press/2026puc&lt;/a&gt;&lt;br&gt;
Texas PUC, Energy and Water Use Survey FAQ — &lt;a href="http://www.puc.texas.gov/industry/water/utilities/energy-and-water-use-survey/faq/" rel="noopener noreferrer"&gt;www.puc.texas.gov/industry/water/utilities/energy-and-water-use-survey/faq/&lt;/a&gt;&lt;br&gt;
Texas PUC docket, Control Number 59281 — interchange.puc.texas.gov/Search/Filings?ControlNumber=59281&lt;br&gt;
Rhode Island Bill H7331 (2026) — legiscan.com/RI/bill/H7331/2026&lt;br&gt;
RI Public Utilities Commission testimony on S2776, Senate Commerce Committee — &lt;a href="http://www.rilegislature.gov/senators/SenateComDocs/2026%20Commerce/S2776%20RI%20Public%20Utilities%20Commission.pdf" rel="noopener noreferrer"&gt;www.rilegislature.gov/senators/SenateComDocs/2026%20Commerce/S2776%20RI%20Public%20Utilities%20Commission.pdf&lt;/a&gt;&lt;/p&gt;

</description>
      <category>datacenterwater</category>
      <category>wateraccountability</category>
      <category>climatepolicy</category>
      <category>wue</category>
    </item>
    <item>
      <title>Why "Same Operator" Doesn't Mean "Same Water Risk"</title>
      <dc:creator>CarbonLayer</dc:creator>
      <pubDate>Sat, 15 Aug 2026 03:29:14 +0000</pubDate>
      <link>https://dev.to/carbonlayer/why-same-operator-doesnt-mean-same-water-risk-3odd</link>
      <guid>https://dev.to/carbonlayer/why-same-operator-doesnt-mean-same-water-risk-3odd</guid>
      <description>&lt;p&gt;We've been building out CarbonLayer's facility-level water stress metrics, and one pattern keeps showing up in the underlying research: the biggest variable in water stress isn't the operator, or even the infrastructure standard a facility is built to. It's the site.&lt;/p&gt;

&lt;p&gt;The problem with fleet-level numbers&lt;/p&gt;

&lt;p&gt;Most water usage reporting for data centers happens at the company or portfolio level — one WUE (Water Usage Effectiveness) figure representing an entire fleet. That number is easy to report and easy to compare. It's also structurally incapable of showing you the thing that actually matters: whether any specific facility sits in a water-stressed region.&lt;/p&gt;

&lt;p&gt;Two facilities built to the identical infrastructure standard, run by the identical operator, can have completely different water risk profiles — because water stress is a function of local supply and demand, not build spec. A facility in a water-scarce region can carry a deficit for most of the year while a sister facility, averaged into the same company-wide number, sits in a low-stress basin. Blend them into one figure and the site carrying the real risk disappears into the average.&lt;/p&gt;

&lt;p&gt;This isn't a hypothetical edge case — it's the norm. Water availability is regional and seasonal; data center portfolios are geographically distributed by design. The mismatch between how the risk actually varies and how it typically gets reported is the gap CarbonLayer exists to close.&lt;/p&gt;

&lt;p&gt;Modeled vs. measured — a separate axis entirely&lt;/p&gt;

&lt;p&gt;The other place aggregation causes false confidence: conflating modeled stress signals (built from frameworks like WRI Aqueduct) with audited or measured facility data. These are different confidence levels. A modeled estimate is a reasonable starting signal. A third-party audited figure is a verified claim. Reporting them side by side as if they're interchangeable is how a portfolio ends up looking safer on paper than it is on the ground.&lt;/p&gt;

&lt;p&gt;Why we measure at the facility level&lt;/p&gt;

&lt;p&gt;This is the core design decision behind CarbonLayer: metrics at the facility level, not the company level, with modeled and measured signals kept explicitly separate rather than blended into one sustainability score. If you're an operator, regulator, or customer trying to assess real exposure, the resolution you measure at determines whether the risk is visible at all.&lt;/p&gt;

&lt;p&gt;Where this is going&lt;/p&gt;

&lt;p&gt;We're expanding audited facility coverage (Toronto and Singapore next) and working on a standard for disclosing modeled vs. measured water metrics side by side, instead of as one blended number.&lt;/p&gt;

&lt;p&gt;If you work in data center ops, ESG reporting, or hydrology/geospatial data and have thoughts on how facility-level disclosure should work — I'd like to hear them.&lt;/p&gt;

</description>
      <category>sustainability</category>
      <category>datacenters</category>
      <category>esg</category>
      <category>climatetech</category>
    </item>
    <item>
      <title>Every inference call already returns your response. We added three more fields: cost, CO2, and water.</title>
      <dc:creator>CarbonLayer</dc:creator>
      <pubDate>Sat, 15 Aug 2026 03:02:43 +0000</pubDate>
      <link>https://dev.to/carbonlayer/every-inference-call-already-returns-your-response-we-added-three-more-fields-cost-co2-and-3mjj</link>
      <guid>https://dev.to/carbonlayer/every-inference-call-already-returns-your-response-we-added-three-more-fields-cost-co2-and-3mjj</guid>
      <description>&lt;p&gt;Most carbon/sustainability tooling for AI means a separate dashboard, a monthly export, or a spreadsheet someone fills in manually. We built it into the call itself.&lt;/p&gt;

&lt;p&gt;Point your existing model API calls at our endpoint at carbonlayer.polsia.io. Same response you already parse — plus:&lt;/p&gt;

&lt;p&gt;{&lt;br&gt;
  ...your normal response,&lt;br&gt;
  costUsd: 0.0021,&lt;br&gt;
  co2Grams: 4.7,&lt;br&gt;
  savedWaterMl: 12.3&lt;br&gt;
}&lt;br&gt;
No separate integration. No batch job reconciling usage logs against a sustainability report six months later. The number is attached to the call, generated by the routing decision, at request time.&lt;/p&gt;

&lt;p&gt;If you're already switching between model providers or regions for latency/cost, you're one header away from also seeing the environmental cost of that same decision — check the docs at carbonlayer.polsia.io to see the exact fields.&lt;/p&gt;

</description>
      <category>sustainability</category>
      <category>api</category>
      <category>webdev</category>
      <category>climatetech</category>
    </item>
    <item>
      <title>Why "Data Center Water Usage" Numbers Are Mostly Guesses — and How We're Fixing That</title>
      <dc:creator>CarbonLayer</dc:creator>
      <pubDate>Sat, 15 Aug 2026 00:42:35 +0000</pubDate>
      <link>https://dev.to/carbonlayer/why-data-center-water-usage-numbers-are-mostly-guesses-and-how-were-fixing-that-32jl</link>
      <guid>https://dev.to/carbonlayer/why-data-center-water-usage-numbers-are-mostly-guesses-and-how-were-fixing-that-32jl</guid>
      <description>&lt;p&gt;Every data center water/energy stat you've seen in a headline this year is probably a model, not a measurement. That distinction matters more than it sounds like it should — and it's currently playing out in real policy decisions, not just spreadsheets.&lt;/p&gt;

&lt;p&gt;The trigger for this post: Louisville just moved to ban new data centers outright. Stacy Griggs wrote a good breakdown of why that's the wrong tool for the job — and the core argument tracks with what I keep seeing from the technical side: a moratorium isn't a policy, it's a symptom. It's what a city does when the only inputs available are fear and headlines, because nobody handed them facility-level numbers to make a real cost/benefit call. Ban-or-blind-trust is a false choice, and it's a data availability problem before it's a political one.&lt;/p&gt;

&lt;p&gt;The actual technical problem: Most public water-stress numbers for data centers come from vendor-published PUE, generic regional averages, or top-down estimates that don't account for facility-specific cooling systems, local water sourcing, or recycling infrastructure. Two facilities in the same city, same operator, same building spec, can have wildly different actual water demand depending on how they're built and run. Regulators and communities making zoning and moratorium decisions are working off averages that hide that variance completely — which is exactly how you end up with a binary ban/no-ban choice instead of a nuanced one.&lt;/p&gt;

&lt;p&gt;What we're building at CarbonLayer: We're extending the ISO/IEC 30134 data center efficiency framework — the standard that gave the industry PUE and WUE — to produce a facility-level metric we call WUE+: water usage effectiveness that incorporates on-site recycling and is layered against regional water stress data from WRI Aqueduct.&lt;/p&gt;

&lt;p&gt;The core design decision: modeled and measured data are never merged into one number. A model gives you a directional estimate. An audit gives you a verified fact. Collapsing those into a single score is how the industry got into a trust problem in the first place — so we tag every metric with its provenance and let the two live side by side instead of averaging away the uncertainty.&lt;/p&gt;

&lt;p&gt;Where we are: Currently auditing facilities across our initial footprint, expanding to Toronto and Singapore next based on where operators and regulators are asking for coverage.&lt;/p&gt;

&lt;p&gt;If you're working on infra observability, ESG tooling, or anything touching facility-level resource accounting, curious what you're running into — drop it in the comments.&lt;/p&gt;

</description>
      <category>sustainability</category>
      <category>datacenter</category>
      <category>climatetech</category>
      <category>esg</category>
    </item>
  </channel>
</rss>
