A startup with only six engineers calculated the cost of their monthly cloud bill to be $11,847. The majority of the data generated was not from the application itself, but from the monitoring tools that were observing the application. ## The recursive joke nobody's laughing at
We often see the same pattern in the industry. You create a small app. Next, you attach the equipment needed to monitor the app. And eventually, the monitoring equipment becomes more expensive than the app itself. An example of this was documented in a Medium post on August 13, 2026. One team reduced their invoice from $11,847/mo to $3,812/mo by switching to two VMs. They came from deleting the overhead that was policing four services that could run on two boxes. That's the basic idea. First, you pay to operate the device, and then you pay an additional amount to observe its operation. š
The numbers are not subtle
This is not just one disgruntled engineer. There's a pattern here that is supported by evidence. Software engineer Devrim Ozcay wrote up a January 13, 2026 recap of a 6-microservice Spring Boot system. His AWS bill was $1,850/mo. Three months in, his Datadog bill reached $3,200/mo and peaked at a $12,000 invoice due to container auto-scaling. Ozcay was straightforward about it:
Our bill for AWS infrastructure was $1,850. Our bill for Datadog was $3,200. Let that sink in for a moment. We were literally spending more money to monitor our servers than to actually run them. For example, an Aug 12, 2026, engineering report on a nine-microservice and Lambda stack found AWS compute at $18,432 and Datadog at $24,773 because of high-cardinality custom metrics. Also, on November 13, 2025, an a developer forum thread mentioned a firm that invested $52,000 in AWS costs, but paid $97,000 for observability, with $47k going to Datadog, $38k going to Splunk, and $12k going to Sentry. The reason behind it is in how it is priced. Datadog is priced per host and per custom metric. The tab autoscales with you. One SRE in that thread summed up the conversation with leadership:
I was asked by leadership why observability was so expensive. I explained that it was because Datadog has a per host cost, and we implement autoscaling. They stared at me as if I had just said something in Martian. ## The mesh tax nobody budgeted for
Having observability is important, but you've likely over-orchestrated. On January 24, 2026, performance tests on service mesh sidecars were shared by Kubernetes performance expert Nawaz Dhandala. Your average sidecar proxy for service mesh consumes 50-100MB of RAM, and 10-50m CPU per pod, additionally it adds 2-5ms of latency per network hop. In an article on August 13, 2026, Cloud architect Khimananda Oli expanded on the disadvantages of sidecar containers, stating that they can consume as much as 200MB RAM each and introduce up to 15ms latency. Do the multiplication. A request crossing five microservices can pile on 75ms of latency and burn a full gigabyte of RAM just to route traffic. This is the painful aspect. You break down the application into smaller parts for quick implementation, but then you face performance and latency issues as you integrate and reassemble those parts over a network. ## People are quietly walking it back
The industry is taking this seriously. An estimated 42% of organizations that have adopted microservices has consolidated at least some services back into larger deployable units. Yes, you're right. Actually, it wasn't about ideology at all. The main issues were related to the complexity of debugging and the high costs of network latency - which were exactly the problems the new tooling was meant to address in the first place. Here's how I see it:
ā The bill isn't a compute problem, it's an architecture problem wearing a compute costume. ā Per-host, per-metric pricing punishes the elasticity you were sold on. ā Four services on Kubernetes with a mesh is often two VMs cosplaying as "scale."
ā You can't monitor your way out of a design you didn't need. This doesn't imply that observability is negative. It just indicates that the ratio is off balance. If the observers are more expensive than the entities being observed, the solution is typically not cheaper observers, but rather reducing the number of entities being observed. First, reduce the surface area, and then, you can proceed to instrument the remaining part. Here's what I'd like to ask you: What's the highest observability-to-compute ratio you've ever encountered, and was there an audible gasp when you mentioned it?
Top comments (1)
Honest answer to your question first: our ratio is close to nil, and that is not a brag, it is scale. Four servers, with Prometheus, Loki and Alertmanager in containers on a box we already pay for. At that size the ratio you describe cannot happen, so I am not the one with the gasp story.
But your framing made me look at a different ratio, and that one was ugly.
Yesterday our deploy pipeline went red and stayed red for 33 hours. Alerting worked exactly as designed: two Telegram messages went out, both delivered, both accurate. Nobody acted on either. The alert-to-action ratio for that incident was 2:0.
Same week: four CI checks green and blind. One scheduled job with 33 passing tests that had never executed once. One test that had silently skipped 50 of its last 64 invocations and reported green every time.
None of that cost a cent in observability spend. All of it cost days.
So the question I would add to yours: has anyone measured signals-sent against signals-acted-on? A $97k Datadog bill at least announces itself. A monitoring stack that costs nothing and is never read is the same failure with no invoice attached ā and much harder to argue for fixing, precisely because nobody can point at a line item.