<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Rishit Bhowmick</title>
    <description>The latest articles on DEV Community by Rishit Bhowmick (@rishit_bmk).</description>
    <link>https://dev.to/rishit_bmk</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4071107%2Fb0f09a5d-b26e-42e7-9815-fe8a7e932d54.jpg</url>
      <title>DEV Community: Rishit Bhowmick</title>
      <link>https://dev.to/rishit_bmk</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/rishit_bmk"/>
    <language>en</language>
    <item>
      <title>Monitoring vs Observability: What's the Difference, and Where Do You Start?</title>
      <dc:creator>Rishit Bhowmick</dc:creator>
      <pubDate>Sun, 04 Oct 2026 13:10:14 +0000</pubDate>
      <link>https://dev.to/rishit_bmk/monitoring-vs-observability-whats-the-difference-and-where-do-you-start-1h6b</link>
      <guid>https://dev.to/rishit_bmk/monitoring-vs-observability-whats-the-difference-and-where-do-you-start-1h6b</guid>
      <description>&lt;p&gt;Your dashboard is green. Users are still complaining. Sound familiar?&lt;/p&gt;

&lt;p&gt;That gap, between "everything looks fine" and "something is clearly broken", is exactly where the monitoring vs observability conversation begins. If you're building on the cloud (or building anything with AI in it), it's worth understanding both. Let's keep it simple.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Monitoring&lt;/em&gt;: the questions you already know to ask&lt;/p&gt;

&lt;p&gt;Monitoring is about watching things you know can go wrong. CPU above 80%? Alert. Error rate spiking? Alert. Disk filling up? Alert.&lt;/p&gt;

&lt;p&gt;You decide the questions in advance, build dashboards and alerts around them, and get pinged when a number crosses a line. It's essential, and it works well for failures you've seen before.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Observability&lt;/em&gt;: the questions you didn't know you'd need&lt;/p&gt;

&lt;p&gt;Modern systems are made of microservices, queues, third-party APIs and now LLM calls. Failures there are often new ones. Observability is about being able to ask fresh questions of your system, without shipping new code just to find the answer.&lt;/p&gt;

&lt;p&gt;You get there with three kinds of data, usually called signals:&lt;/p&gt;

&lt;p&gt;Metrics: numbers over time (latency, error rate, requests per second)&lt;br&gt;
Logs: timestamped records of what happened&lt;br&gt;
Traces: the full journey of one request across all your services&lt;/p&gt;

&lt;p&gt;Monitoring tells you something is wrong. Observability helps you work out why. You need both.&lt;/p&gt;

&lt;p&gt;Why this is on everyone's radar right now&lt;/p&gt;

&lt;p&gt;Grafana Labs surveyed 1,300+ practitioners this year. The biggest worry wasn't tooling, it was complexity and overhead (38%), followed by signal-to-noise (34%) and cost (31%). And 77% said open source or open standards matter to their observability strategy.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Translation: teams are drowning in data, paying a lot for it, and don't want to be locked in. Which brings us to the most useful thing to know.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;OpenTelemetry: instrument once, send anywhere&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;OpenTelemetry (OTel) is the open standard for collecting metrics, logs and traces. Instead of using each vendor's own agent, you instrument your code once with OTel and then choose where the data goes. Switching tools later doesn't mean rewriting your instrumentation.&lt;/p&gt;

&lt;p&gt;It's also moving fast. A few recent bits:&lt;/p&gt;

&lt;p&gt;The Kubernetes attributes processor, which adds Kubernetes metadata to your telemetry, hit v1.0.0 in September.&lt;br&gt;
The Go Logs API and SDK reached release candidate status in August.&lt;br&gt;
The JavaScript SDK 3.0 is planned for October 15, and it merges the separate tracing packages into one (@opentelemetry/sdk-trace), so JS folks should plan a small migration.&lt;/p&gt;

&lt;p&gt;If you only pick one thing to learn this quarter, make it OTel.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fj86kcdn427lxv2aca9kh.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fj86kcdn427lxv2aca9kh.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;&lt;strong&gt;The tools&lt;/strong&gt;&lt;/em&gt;: what each one does differently&lt;/p&gt;

&lt;p&gt;Think of these as categories, not a leaderboard. Each is here for one thing it does differently.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Datadog is a managed platform. Its LLM Observability traces every step of an LLM workflow, clusters production traffic to show quality trends, and can mask sensitive data and flag prompt-injection attempts.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Grafana is the open-source-first route. Grafana Alloy, its vendor-neutral distribution of the OTel Collector, collects Prometheus metrics, OpenTelemetry data, Loki logs and profiles in one agent.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;OpenObserve is open source (Rust) and puts logs, metrics, traces, real user monitoring and LLM tracing in one tool. It accepts OpenTelemetry data and can be queried with SQL or PromQL. It runs as a single binary on S3-compatible object storage, and it claims 140x lower storage costs than Elasticsearch. That's the vendor's own figure, so test it on your data. v1.0 landed in September.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Amazon CloudWatch Omni is built around applications and teams. It shows telemetry across multiple AWS accounts, regions and other clouds like Azure, discovers services and maps dependencies automatically, and lets you investigate through natural-language chat.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Cloudflare has the edge as its angle. Its new platform builds on OpenTelemetry, W3C Trace Context and SQL, with Traces in open beta. Paid plans include 50 GB a month, then $0.25/GB ingested, with the new pricing starting December 1, 2026.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Cortex XCOR from Palo Alto Networks leads with an AI SRE agent that investigates incidents when alerts fire, plus data optimization built on Chronosphere technology. The vendor's numbers (about 75% root-cause success, 89% average data reduction) are its own, so treat them as claims.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For LLM apps and agents, tracing prompts, tool calls and token cost is now part of several of these platforms: Datadog, Grafana and CloudWatch Omni (which also has a free VS Code extension for local agent evaluation).&lt;/p&gt;

&lt;p&gt;The pattern behind all the launches&lt;/p&gt;

&lt;p&gt;Two things keep showing up: AI that investigates incidents for you, and observability for AI systems. Expect both to keep growing. But the fundamentals above haven't changed, and the AI layer is only as good as the telemetry underneath it.&lt;/p&gt;

</description>
      <category>devops</category>
      <category>monitoring</category>
      <category>sre</category>
    </item>
  </channel>
</rss>
