Pipeline & Prompts | Byte size guides on DevOps, Cloud and AI
☁️ Cloud Without the Chaos #5
⚡ Byte Size Summary
- Why native Kubernetes cost attribution (namespace/project-level, via CMMO) and organizational cost-center attribution are different problems, and why one mechanism can't do both jobs
- How cost detection (fast, alert-based) and cost attribution (detailed, reconciled) are separate concerns that don't need — and shouldn't need — the same mechanism
- What actually happens to attribution when the organizational-label layer goes away, and why that's not the same as losing all visibility
The Story
Cloud cost estimates tend to look straightforward when a workload is first being planned. Size the application, estimate average utilization, account for expected growth, enable autoscaling, and build the procurement number.
The problem starts when the workload doesn't behave like the estimate.
The procurement number for this engagement looked solid on paper. The application was homegrown, autoscaling was enabled, and the infrastructure had been sized against what appeared to be a reasonable average-utilization estimate. Nothing raised a concern during review, and there were no issues at go-live.
Then the cloud spend alerts started firing.
When we looked at the usage pattern by time of day, the reason became obvious. The heaviest workload wasn't happening during normal business hours. It was happening in the evenings, overnight, and during weekends — exactly the periods that had been smoothed into an average during the original sizing exercise.
Nothing unusual was happening with the application itself. API calls, logging, request processing, and the application's normal request/response activity were all behaving as expected. The autoscaler was doing what it was designed to do: responding to increased demand by adding capacity.
The problem was that the demand was occurring at a different time, and at a different shape, than the original estimate assumed.
The resulting cloud bill was approximately 20-30% higher than the original projection. And it stayed there. The workload's actual demand curve simply didn't match the curve used to build the estimate.
That gap between what we estimate and what the workload actually does is where the cost-attribution problem starts — the third of the five dimensions worth placing deliberately.
The Problem
The person who feels this first is usually whoever owns the cloud budget. The person who has to explain it is often whoever performed the original sizing.
That leads to a fairly simple question: which workload is actually driving the cost?
On a shared Kubernetes cluster running applications for multiple teams, knowing that the cluster cost a certain amount this month doesn't answer that question. Was the increase caused by one team's production application with genuinely spiky traffic? Was it a development or test workload that wasn't scaled down? Was a logging-heavy application consuming more resources than expected? Or was part of the cost associated with shared cluster infrastructure that shouldn't be assigned to an individual application team?
A cluster-level cost number doesn't provide that level of visibility.
Without granular attribution, the available responses tend to be broad ones: reduce the autoscaling ceiling, change the overall budget, or absorb the additional cost and try to improve the estimate next time. None of those approaches really solve the underlying problem.
Before we can optimize the cost, we need to understand where the cost came from.
Why Existing Approaches Fall Short
The first instinct is usually to look at the cloud provider's native cost-management capabilities. That's a reasonable place to start.
Cloud providers can associate costs with resources using tags, resource groups, subscriptions, accounts, and other infrastructure-level dimensions. This works well when the infrastructure resource and the accountability boundary are essentially the same thing.
Kubernetes changes that relationship. The cloud provider may be billing for virtual machines or node pools, while the application team is thinking in terms of namespaces, deployments, pods, and services. The infrastructure being billed and the workload consuming that infrastructure aren't necessarily the same object.
If three application teams share the same Kubernetes cluster, the cloud provider sees the infrastructure. The application teams see their workloads. Both views are correct. Neither view, by itself, answers the complete cost question.
The Architecture
Starting with native Kubernetes cost attribution
The next layer is understanding how much of the shared infrastructure each workload actually consumes. In an OpenShift environment, Red Hat Cost Management Metrics Operator (CMMO) provides this capability by using Prometheus/Thanos usage data from the cluster.
CMMO allows cost to be distributed at the project level, where an OpenShift project corresponds to a Kubernetes namespace. This is a significant improvement over looking only at the underlying infrastructure. Instead of asking how much did this cluster cost? we can start asking how much of that infrastructure consumption belongs to each project?
For many environments, that may be sufficient. In this case, it wasn't. The technical structure of the cluster didn't map cleanly to the organizational structure. A department could own several application namespaces. A namespace didn't necessarily map to exactly one department or cost center.
Namespace-level attribution correctly answers which workload or project consumed the resources? It doesn't necessarily answer which department or cost center should own that cost? That second question required another attribution dimension.
Adding organizational context
This is where labels become useful. Rather than replacing namespace-level attribution, we add organizational metadata to the workload. Deployment objects can carry labels representing information such as department, cost center, business unit, application owner, or environment.
An admission policy ensures the required attribution information is present before a deployment reaches production or is scheduled onto a node — see Implementation, Step 2, for the specific mechanisms this can be built on.
This creates two complementary dimensions. The namespace tells us where the workload lives. The label tells us who owns it from an organizational perspective. That distinction becomes particularly useful when the organization and the Kubernetes structure don't line up one-to-one. A department that owns five application namespaces can be shown per-namespace, or the label lets those five namespaces be viewed together as a single organizational cost center.
The label isn't replacing native Kubernetes cost attribution. It's adding context to it.
The two data sources
The implementation has two primary data sources.
The first is cluster usage data. CMMO reads Prometheus/Thanos usage data and uploads it to Red Hat's cost-management service on an approximately six-hour cycle. This provides the resource-consumption side of the equation: who used what?
The second is cloud billing data. For Azure, a native Azure Cost Export provides actual-cost data. In this implementation, the daily CSV export lands in a storage account in the same resource group as the cluster, and a service principal provides the access required for the cost-management platform to read that data. This provides the billing side: what did Azure actually charge?
The cost-management system correlates Azure VM instance IDs with OpenShift nodes and uses the usage information to distribute infrastructure costs across projects. The native attribution boundary is therefore the OpenShift project or namespace. The organizational label provides an additional dimension on top of that.
Two pipelines, not one
Cost attribution isn't a real-time process. There are two separate data pipelines that need to stay active:
- Cluster usage: Prometheus/Thanos → CMMO → cost-management platform, on a roughly six-hour upload cycle
- Azure billing: Azure Cost Export → storage account → service-principal-scoped access → cost-management platform, on a daily export cycle
As a result, the attributed cost view can take up to 24 hours to reflect the current state of the environment. The dashboard isn't a live query into the cluster — it's a view of the environment after the usage and billing data have passed through their respective collection and processing cycles.
Cost detection and cost attribution are different problems
The mechanism that tells us we have a cost problem doesn't need to be the same mechanism that tells us who caused it.
In this implementation, native Azure Cost Management spend alerts provide the faster detection mechanism — that's what identified the original overrun. The alerting path operates directly against Azure billing information and isn't dependent on the CMMO or Cost Export processing cycle.
Spend alerting asks has spending crossed the threshold? Cost attribution asks where did that spending come from? The first needs to be fast. The second needs to be detailed. Trying to make one mechanism do both jobs creates unnecessary complexity.
What happens when something goes wrong
Two failure points are worth naming.
The first is the cloud billing integration. If the service principal used to access the Azure Cost Export is given broader permissions than necessary, the integration could expose cost information beyond the intended scope.
The second is the workload-label enforcement mechanism. If the admission policy is bypassed or disabled, workloads may be deployed without the required organizational labels. That doesn't break namespace-level attribution — the native project/namespace cost information keeps working. What's lost is the additional department or cost-center dimension for workloads that don't carry the required metadata.
That's another reason to treat organizational labels as an additional attribution layer, not the foundation of the entire cost model — the foundation is CMMO's native namespace attribution, which keeps functioning whether or not the label layer does.
The diagram earns its place here by making one thing visible that prose can't: two independent pipelines, on two different cadences, converging on one dashboard — with the label/admission-policy layer drawn as an overlay on native namespace attribution rather than something the whole model depends on. That's the design decision worth seeing, not just reading.
The pattern itself isn't Azure- or ARO-specific: native platform-level usage attribution, paired with an organizational label layer for the cases namespace boundaries and org charts don't agree. Amazon EKS (Elastic Kubernetes Service) has its own usage/billing correlation through Cost and Usage Reports and Kubecost-style tooling; Azure AKS (Kubernetes Service) and Google GKE (Kubernetes Engine) have their own native cost management surfaces. The mechanics below — the operator, the export cadence, the specific IAM roles — are the ARO implementation of that pattern, because that's the engagement this comes from. Swap the platform-specific pieces and the same two-pipeline shape holds.
Implementation
Prerequisites
- Azure Red Hat OpenShift (ARO) cluster with
cluster-adminocaccess - Azure subscription hosting the cluster,
azCLI logged in with rights on the cluster's resource group - Red Hat Hybrid Cloud Console access with the Cloud Administrator role (or equivalent cost-management write access)
- Kubernetes 1.30+ (or an OpenShift version that ships it) for native
ValidatingAdmissionPolicysupport — see Step 2 for the OPA Gatekeeper / Kyverno alternatives if you're not on a version that has it - Everything below (storage account, export, service principal) is scoped to the single resource group that holds the ARO cluster — not the whole subscription
Step 1 — Install the Cost Management Metrics Operator
# Verify the operator is available, then install via OperatorHub/Software Catalog
# or apply directly:
cat <<EOF | oc apply -f -
apiVersion: operators.coreos.com/v1
kind: OperatorGroup
metadata:
name: costmanagement-metrics-operator
namespace: costmanagement-metrics-operator
spec:
targetNamespaces:
- costmanagement-metrics-operator
EOF
cat <<EOF | oc apply -f -
apiVersion: operators.coreos.com/v1alpha1
kind: Subscription
metadata:
name: costmanagement-metrics-operator
namespace: costmanagement-metrics-operator
spec:
channel: stable
installPlanApproval: Automatic
name: costmanagement-metrics-operator
source: redhat-operators
sourceNamespace: openshift-marketplace
EOF
Then create the Hybrid Console service account (Settings → Identity & Access Management → Service Accounts, added to a group with the Cloud Administrator role), store its client_id/client_secret as a cluster Secret, and apply a CostManagementMetricsConfig:
oc create namespace costmanagement-metrics-operator --dry-run=client -o yaml | oc apply -f -
cat <<EOF | oc apply -f -
apiVersion: v1
kind: Secret
metadata:
name: service-account-auth-secret
namespace: costmanagement-metrics-operator
type: Opaque
stringData:
client_id: "<CLIENT_ID>"
client_secret: "<CLIENT_SECRET>"
EOF
cat <<EOF | oc apply -f -
apiVersion: costmanagement-metrics-cfg.openshift.io/v1beta1
kind: CostManagementMetricsConfig
metadata:
name: costmanagementmetricscfg
namespace: costmanagement-metrics-operator
spec:
authentication:
type: service-account
secret_name: service-account-auth-secret
packaging:
max_reports_to_store: 30
max_size_MB: 100
prometheus_config:
collect_previous_data: true
context_timeout: 120
source:
check_cycle: 1440
create_source: true
name: aro-prod-cost
upload:
upload_cycle: 360
upload_toggle: true
EOF
This engagement used the service-account auth type rather than the deprecated basic/token mode — worth calling out, since the token mode is what most quick-start examples default to. This is what produces native, project/namespace-level cost distribution from Prometheus/Thanos usage data — the baseline every other layer sits on top of.
Rollback here is a clean uninstall: delete the CostManagementMetricsConfig, then the Subscription and OperatorGroup, then the namespace. CMMO doesn't write anything back to the cluster's application workloads — removing it stops future uploads but doesn't touch anything already reported to the Hybrid Console.
Step 2 — Enforce the cost-center label at admission time
The requirement itself is simple to state: no Deployment reaches production or gets scheduled onto a node without a cost-center label present. This engagement enforced it with Kubernetes' own native admission policy — no external operator installed on the cluster.
What was actually used — native ValidatingAdmissionPolicy. CEL-based, GA from Kubernetes 1.30; confirm your OpenShift version ships it before relying on this path:
apiVersion: admissionregistration.k8s.io/v1
kind: ValidatingAdmissionPolicy
metadata:
name: "require-cost-center-policy"
spec:
failurePolicy: Fail
matchConstraints:
resourceRules:
- apiGroups: ["apps"]
apiVersions: ["v1"]
operations: ["CREATE", "UPDATE"]
resources: ["deployments"]
validations:
- expression: "has(object.metadata.labels) && 'cost-center' in object.metadata.labels"
message: "Deployments must include a 'cost-center' label."
---
apiVersion: admissionregistration.k8s.io/v1
kind: ValidatingAdmissionPolicyBinding
metadata:
name: "require-cost-center-binding"
spec:
policyName: "require-cost-center-policy"
validationActions: ["Deny"]
No new controller to install or upgrade, at the cost of being the newest and least battle-tested option on OpenShift specifically.
If your cluster isn't on a version that ships native ValidatingAdmissionPolicy, or you're already standardized on a policy engine, the same requirement maps onto either of these — shown here for readers on a different cluster, not what this engagement ran:
Option — OPA Gatekeeper, via a constraint on K8sRequiredLabels:
apiVersion: constraints.gatekeeper.sh/v1beta1
kind: K8sRequiredLabels
metadata:
name: require-cost-center-label
spec:
match:
kinds:
- apiGroups: ["apps"]
kinds: ["Deployment"]
parameters:
labels:
- key: "cost-center"
allowedRegex: "^[a-z0-9-]+$"
This assumes the underlying ConstraintTemplate for K8sRequiredLabels is already installed from Gatekeeper's constraint template library — the constraint above references it, it doesn't define it. Pin to Gatekeeper v3.x.
Option — Kyverno, via a ClusterPolicy:
apiVersion: kyverno.io/v1
kind: ClusterPolicy
metadata:
name: require-cost-center-label
spec:
validationFailureAction: Enforce
background: true
rules:
- name: check-cost-center
match:
any:
- resources:
kinds:
- Deployment
validate:
message: "The 'cost-center' label is mandatory on all Deployments."
pattern:
metadata:
labels:
cost-center: "?*"
validationFailureAction: Enforce is what actually blocks the deploy — set to Audit first if you want visibility before you start rejecting anything.
No deployment artifact reaches a node without this label, regardless of which mechanism enforces it. Without it, Cost Management still gives you namespace-level attribution from CMMO — you just lose the department-level cut for any namespace that doesn't map cleanly to one cost center.
Step 3 — Export Azure billing data and grant read access
# Storage account in the same resource group as the cluster.
# public-network-access is disabled — reached only through a private
# endpoint, not a public IP. Where a private endpoint isn't an option,
# VNet rules plus IP firewall restrictions are the documented fallback.
az storage account create \
--name "$CM_STORAGE" \
--resource-group "$ARO_RG" \
--location "$ARO_LOCATION" \
--sku Standard_LRS \
--kind StorageV2 \
--min-tls-version TLS1_2 \
--public-network-access Disabled
# Service principal for Red Hat's read access — needs BOTH roles
SP_JSON=$(az ad sp create-for-rbac \
--name sp-rh-cost-management \
--role "Storage Blob Data Reader" \
--scopes "$EXPORT_SCOPE" \
--json-auth)
az role assignment create \
--assignee "$(echo "$SP_JSON" | jq -r .clientId)" \
--role "Cost Management Reader" \
--scope "$EXPORT_SCOPE"
Store the returned clientSecret immediately — it is not retrievable again after this command returns. Put it in Azure Key Vault if secrets are managed centrally, or in the CI/CD pipeline's own secure secret store (GitHub Actions Secrets, an Azure DevOps variable group) — never in a pipeline variable that lands in shell history or a plaintext log.
# CLI export creation needs the costmanagement extension and an existing container
az extension add --name costmanagement
az storage container create \
--account-name "$CM_STORAGE" \
--name costexport \
--auth-mode login
STORAGE_ACCOUNT_ID="/subscriptions/${SUBSCRIPTION_ID}/resourceGroups/${ARO_RG}/providers/Microsoft.Storage/storageAccounts/${CM_STORAGE}"
# Recurrence start must be today or future — Azure CLI requirement
EXPORT_FROM=$(date -u -d '+1 day' +%Y-%m-%dT00:00:00Z)
EXPORT_TO=$(date -u -d '+2 years' +%Y-%m-%dT00:00:00Z)
# Daily actual-cost export, scoped to the resource group
az costmanagement export create \
--name rh-cost-export-daily \
--scope "$EXPORT_SCOPE" \
--type ActualCost \
--timeframe MonthToDate \
--storage-account-id "$STORAGE_ACCOUNT_ID" \
--storage-container costexport \
--storage-directory daily \
--recurrence Daily \
--recurrence-period from="$EXPORT_FROM" to="$EXPORT_TO" \
--schedule-status Active
If Red Hat Cost Management rejects the export schema, fall back to creating it through the Azure Portal instead — Cost Management + Billing → Cost export → Daily export, using the "Cost and usage details (actual)" template. That's the path the guide points to when the CLI-created export doesn't validate; worth knowing before spending an afternoon debugging a schema error the CLI doesn't explain.
The Cost Management Reader role matters as much as Storage Blob Data Reader — the read permission alone doesn't get you a functioning Azure integration in the Hybrid Console. Rollback is clean on this step in isolation: removing the export or the role assignment stops new billing data from landing, but doesn't touch what CMMO already reported from the cluster side.
Security Considerations
Service principal scope on the Azure billing integration. The service principal reading the daily cost export is, by definition, reading billing data that spans every namespace and every department on the cluster. If it's scoped more broadly than Storage Blob Data Reader + Cost Management Reader on the single resource group — subscription-level access, for instance — someone with access to the cost dashboard can infer more than cost: traffic volume and scaling patterns for departments other than their own. Confirmed as a real concern on this engagement; the mitigation is scoping strictly to the resource group holding the cluster, not the subscription.
Service principal credential storage. The az ad sp create-for-rbac command that provisions Red Hat's read access returns a client secret that Azure will not show again. Store it immediately in Azure Key Vault if secrets are managed centrally, or in the CI/CD pipeline's own secure secret store — GitHub Actions Secrets or an Azure DevOps variable group — so it never lands in shell history or a plaintext log.
CMMO authentication mode. This engagement used the Hybrid Console service-account auth type (client_id/client_secret stored as a cluster Secret) rather than the deprecated basic/token mode — worth stating explicitly, since token auth is what most quick-start examples default to and is the weaker of the two options for the cluster-to-console leg of this pipeline.
Read access on the cost dashboard itself is a separate control from the pipeline feeding it. Scoping the service principal correctly limits what the data pipeline can reach. It says nothing on its own about who can view the resulting cost report. This engagement handled that separately, through Red Hat Hybrid Cloud Console User Access Groups — only authorized finance and department leads had access to aggregated cost metrics, distinct from whoever manages the underlying Azure integration.
The storage account itself. The exported billing CSV sitting in the storage account is a sensitive, cross-department asset independent of anyone going through the Hybrid Console properly — someone with direct storage account access bypasses the Console's own access model entirely. This engagement secured it with Azure Private Endpoints, so the raw billing CSVs are never reachable over a public-facing IP; virtual network rules and IP firewall restrictions are the fallback where a private endpoint isn't an option.
Tradeoffs
What you gain / what you give up — granularity vs. currency. Two-layer attribution (native namespace/project plus label-based cost-center) gives you real per-application, per-department visibility that a flat cluster bill never could. What you give up is real-time accuracy: the ~24-hour cycle across CMMO's 6-hour uploads and Azure's daily cost export means the attributed-cost view is always behind. If you're making a same-day budget decision, you're making it against the faster but less granular native Azure Cost Management alert, not the project/label-attributed dashboard.
What you gain / what you give up — deployment friction vs. cost-center accuracy. Blocking unlabeled deployments at admission time is what makes department-level attribution work for namespaces that don't map cleanly to one cost center. What it costs is friction: every deployment pipeline now has a hard dependency on getting the label right, and a missing or wrong label doesn't just skew a report, it blocks a deploy. That's a deliberate tradeoff, not an accident, but it needs to be communicated to every team shipping to the cluster before it surfaces as a support ticket.
Operability tradeoff — two independent pipelines instead of one, with a graceful (not flat) failure mode. Running CMMO (cluster-side, usage-based) and the Azure Cost Export (billing-side, dollar-based) as two separate integrations that both have to stay active is more moving parts than a single source of truth would be. The upside of that separation: when the label layer goes down, attribution doesn't collapse to an even split across departments — it falls back to CMMO's native namespace/project view, which keeps working independently of the label/admission-policy layer. You lose the department-level cut, not all visibility.
What's Next
Cloud Without the Chaos #6 picks up the framework's fifth dimension directly: blast radius, tested honestly against a real failure instead of assumed on paper. This article's own two failure points — an over-scoped billing integration, and a bypassed label policy that degrades gracefully rather than catastrophically — are a small preview of that larger question.
There's no companion repository for this one. This documents a cost-attribution architecture and a set of operational decisions specific to one ARO engagement, not a general-purpose tool — the Gatekeeper/Kyverno alternatives in Step 2 are illustrative, not what was actually run here.
Written by Pipeline & Prompts | Byte size guides on DevOps, Cloud and AI

Top comments (0)