<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Bonree-Js</title>
    <description>The latest articles on DEV Community by Bonree-Js (@bonree-js).</description>
    <link>https://dev.to/bonree-js</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4124005%2Fb0212323-ae4e-4db3-842c-d1cbf4d57cab.jpg</url>
      <title>DEV Community: Bonree-Js</title>
      <link>https://dev.to/bonree-js</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/bonree-js"/>
    <language>en</language>
    <item>
      <title>Agentic AI Operations: Evidence-Carrying Handoffs for Multi-Agent Diagnosis</title>
      <dc:creator>Bonree-Js</dc:creator>
      <pubDate>Mon, 28 Sep 2026 09:57:18 +0000</pubDate>
      <link>https://dev.to/bonree-js/agentic-ai-operations-evidence-carrying-handoffs-for-multi-agent-diagnosis-1ko7</link>
      <guid>https://dev.to/bonree-js/agentic-ai-operations-evidence-carrying-handoffs-for-multi-agent-diagnosis-1ko7</guid>
      <description>&lt;p&gt;TL;DR — In a multi-agent operations workflow, the orchestrator's answer is only as trustworthy as the claims its sub-agents hand back. Structuring those handoffs as verifiable evidence, reasoning by elimination across the service topology, and replaying past incidents as a regression suite make multi-agent diagnosis easier to audit and improve.&lt;br&gt;
What is a multi-agent operations workbench? It is a system in which an orchestrating agent interprets an operator's request and routes parts of it to specialist agents, such as incident diagnosis, Q&amp;amp;A, reporting, or remediation.&lt;br&gt;
Written by the Bonree observability team. Bonree | Sage AI (Bonree ONE • Sage AI) is one such workbench: Bonree describes it as an AI Ops agent workbench that recognizes user intent through natural language and orchestrates the right agents, tools, and skills across unified operations and AI scenarios, home to a team of specialized agents with multi-agent collaboration for complex operational tasks. The patterns below are general design ideas, not descriptions of how any product is configured or operated.&lt;br&gt;
Keywords: multi-agent systems, LLM agents, agent reliability, agent trustworthiness, trustworthy AI, verifiable AI, root cause analysis, AIOps, AI observability, LLM observability, SRE, OpenTelemetry, DevOps&lt;br&gt;
Why handoffs are the weak point&lt;br&gt;
Discussions of multi-agent ops usually focus on routing. The harder problem is what comes back. If a specialist returns "the database looks unhealthy," the orchestrator can only repeat it, and a human reviewer cannot tell whether it came from a query, a runbook, or a guess.&lt;br&gt;
This matters more as agents are trusted with more autonomy. An orchestrator that acts on a specialist's claim without a way to check it is making a leap of faith, and that risk compounds across multi-step workflows: each additional hop is another place for an unverified claim to become the input to the next one. Evidence-carrying handoffs don't remove that risk, but they keep the chain auditable instead of just fast.&lt;br&gt;
Pattern 1: Evidence-carrying handoffs&lt;br&gt;
Ask each specialist to return a structured answer instead of prose. An illustrative structure (not a Bonree ONE data format):&lt;br&gt;
Field   Purpose&lt;br&gt;
Claim   The conclusion, in one sentence&lt;br&gt;
Confidence  Coarse (low / medium / high), not a fake-precise percentage&lt;br&gt;
Evidence    Where it came from: data source, query or reference, time range, entity, what was observed&lt;br&gt;
Ruled out   What was excluded, and which signal was checked&lt;br&gt;
Suggested next check    What to look at if the claim is wrong&lt;br&gt;
Evidence can point to OpenTelemetry traces, metrics, and logs, or any other re-runnable source.&lt;br&gt;
Pattern 2: Reason by elimination across the topology&lt;br&gt;
Checking every metric of every service is slow and noisy. Reasoning across the call chain is cheaper: if the entry service shows failures and every downstream service shows none, the fault most likely starts at the entry node, and whole downstream branches can be dropped in one step.&lt;br&gt;
Bonree has published a case of this shape from Bonree ONE • Sage AI's health analysis of a real service chain. The entry-point service showed 7 failed requests (a 0.04% error rate) while the two downstream services it called showed a zero error rate, so the analysis narrowed to the entry node instead of tracing further downstream. It separately flagged that all three services in the chain were deployed on the same host, which itself had 7 unresolved alerts, a structural single-point-of-failure risk that the per-service error rates alone did not show. Details, including the full topology reasoning and the resulting diagnostic report.&lt;br&gt;
Caveat: a zero error rate is not the same as healthy. Record which signal was checked, so latency or saturation problems are not silently excluded.&lt;br&gt;
Pattern 3: Share the tool layer, specialize the reasoning&lt;br&gt;
If every agent carries its own connectors, integration cost grows with the number of agents. Specialists work better when they differ in instructions and skills while models, tools, and knowledge come from one shared pool. Bonree ONE • Sage AI is described in the same terms: agents draw on a shared pool, so covering a new scenario does not require duplicating connections.&lt;br&gt;
Pattern 3a: Keep the shared layer framework-agnostic&lt;br&gt;
A related design question is which orchestration framework the specialists actually run on. If the shared tool and evidence layer from Pattern 3 is wired directly into one framework's APIs, every additional framework a team adopts means re-wiring those connectors again. A framework-agnostic tool contract avoids that: tool functions, schemas, and the evidence format from Pattern 1 are defined independently of any single orchestrator, with a thin adapter mapping them into whichever framework a given agent runs on.&lt;br&gt;
That contract still has to be observed consistently once it's running, and that's a separate problem: whatever framework a given specialist is built on, its tool calls, token usage, and evidence need to show up in the same shape in your traces. See the companion post on what differs across LangChain, LangGraph, Dify, and OpenClaw at the instrumentation layer, and why that's an AI Observability problem more than an orchestration one (Article 4 below).&lt;br&gt;
Pattern 4: Replay past incidents as a regression suite&lt;br&gt;
Keep closed incidents as fixtures: the alert, a topology snapshot, the time window, and the human-confirmed root cause. Replay them and track four things:&lt;br&gt;
Measure Question it answers&lt;br&gt;
Root-cause hit rate Did the final claim match the confirmed cause?&lt;br&gt;
Evidence validity   Do the cited queries reproduce what was claimed?&lt;br&gt;
Ruled-out precision Were excluded branches really innocent?&lt;br&gt;
Steps to first useful hypothesis    How quickly did the agents narrow the search?&lt;/p&gt;

&lt;p&gt;Limitations&lt;br&gt;
●Checking evidence proves a citation reproduces, not that the inference drawn from it is correct.&lt;br&gt;
●Re-running evidence needs stable, queryable tool interfaces.&lt;br&gt;
●Elimination reasoning is only as good as the topology it reads.&lt;br&gt;
●A replay suite built from a handful of incidents can overfit.&lt;br&gt;
●Framework-agnostic tool contracts add an adapter layer to maintain; for a single-framework team this is overhead until a second framework actually shows up.&lt;br&gt;
FAQ&lt;br&gt;
What is an evidence-carrying handoff? A structured sub-agent answer that includes its claim, the sources behind it, and what it ruled out, so an orchestrator or a human can verify it.&lt;br&gt;
Does multi-agent diagnosis remove human review? No. Evidence makes review faster and more specific.&lt;br&gt;
Does supporting multiple orchestration frameworks weaken the shared tool layer? Not if the tool and evidence contracts are defined independently of any framework; the framework only decides how an agent is orchestrated, not what evidence it must produce.&lt;br&gt;
Further reading&lt;br&gt;
●&lt;a href="https://en.bonree.com/en/sageai" rel="noopener noreferrer"&gt;https://en.bonree.com/en/sageai &lt;/a&gt;&lt;br&gt;
●&lt;a href="https://dev.to/bonree-js/building-ai-observability-for-the-native-stack-architecture-design-and-engineering-practice-from-nc1"&gt;https://dev.to/bonree-js/building-ai-observability-for-the-native-stack-architecture-design-and-engineering-practice-from-nc1&lt;/a&gt; &lt;br&gt;
●&lt;a href="https://en.bonree.com" rel="noopener noreferrer"&gt;https://en.bonree.com&lt;/a&gt; &lt;/p&gt;

&lt;h1&gt;
  
  
  observability #opentelemetry #devops #Bonree
&lt;/h1&gt;

</description>
      <category>ai</category>
      <category>devops</category>
      <category>llm</category>
    </item>
    <item>
      <title>Instrumenting LangChain, LangGraph, Dify, and OpenClaw: What Actually Differs at the Tracing Layer</title>
      <dc:creator>Bonree-Js</dc:creator>
      <pubDate>Mon, 28 Sep 2026 09:37:03 +0000</pubDate>
      <link>https://dev.to/bonree-js/instrumenting-langchain-langgraph-dify-and-openclaw-what-actually-differs-at-the-tracing-layer-jn0</link>
      <guid>https://dev.to/bonree-js/instrumenting-langchain-langgraph-dify-and-openclaw-what-actually-differs-at-the-tracing-layer-jn0</guid>
      <description>&lt;p&gt;The previous post walked through instrumenting an LLM call by hand: a span per model call, a histogram for token usage, a small set of fixed-cardinality labels. That works when there's one call site to wrap. It gets harder once a platform team is supporting several product teams that each picked a different agent framework, because LangChain, LangGraph, Dify, and OpenClaw don't expose the same thing to instrument in the same place.&lt;br&gt;
Disclosure: I work at Bonree, an observability vendor. Most of this post is vendor-neutral; Bonree ONE appears near the end.&lt;br&gt;
Four frameworks, four different instrumentation surfaces&lt;br&gt;
●LangChain runs in-process as a Python or JS library, and exposes a callback interface (handlers for on_llm_start, on_llm_end, on_tool_start, and similar events). Instrumentation attaches to these callbacks directly in the calling application's process.&lt;br&gt;
●LangGraph adds an explicit state machine on top: nodes, edges, and checkpoints. The natural place to attach a span is per node execution and per state transition, not per model call, since a single node can make zero, one, or several model calls.&lt;br&gt;
●Dify is mostly consumed as a hosted, low-code platform: a large part of the pipeline runs inside Dify's own engine rather than in the calling application's process, so in-process callbacks aren't available for most of it. Visibility typically comes from Dify's own execution logs or webhooks instead.&lt;br&gt;
●OpenClaw is a standalone, always-on agent runtime rather than a library you call, so instrumentation has to happen at the boundary between it and everything else, its API or webhook surface, instead of inline in application code.&lt;br&gt;
In practice, this means "add tracing" means four different things depending on the framework: wrap a callback, wrap a node function, consume an execution log, or hook a webhook. None of that is difficult in isolation. It becomes real, ongoing work when a platform team is the one expected to keep all four producing comparable data.&lt;br&gt;
Why comparable matters more than complete&lt;br&gt;
It's not enough for each framework to be traced somehow. If LangChain traces carry gen_ai.agent.name and LangGraph traces carry graph.node.name with no equivalent, per-agent cost and error-rate comparisons across frameworks silently break. The fix is the same one from the previous post: standardize on a shared attribute schema, such as OpenTelemetry's GenAI semantic conventions, and map every framework's own event model onto it, rather than letting each framework's instrumentation invent its own field names.&lt;br&gt;
What automatic instrumentation is actually buying you&lt;br&gt;
Given that mapping work, "out-of-the-box support for framework X" in an observability tool is worth more than it sounds like on a feature list. It's not just a connector, it's someone else having already done the work of mapping LangChain's callbacks, LangGraph's node/edge model, Dify's execution logs, and OpenClaw's API surface onto one consistent schema, so a platform team isn't building and maintaining four separate adapters by hand.&lt;br&gt;
Limitations&lt;br&gt;
●All four frameworks change quickly. Automatic instrumentation still needs to be re-validated against new framework versions, the same way hand-written instrumentation does.&lt;br&gt;
●A hosted platform like Dify may not expose every internal step to any external collector, in-process or not, so some signals may only ever be as detailed as what the platform's own logs provide.&lt;br&gt;
●Consistent collection solves comparability, not governance. Turning collected data into per-team cost and usage rollups is the labeling and aggregation problem covered in the previous post, not something instrumentation coverage does by itself.&lt;br&gt;
Where Bonree ONE fits&lt;br&gt;
Bonree ONE 4.0's AI Observability capability instruments Python, Node.js, and Java applications, with automatic adaptation for common model-native APIs and out-of-the-box support for LangChain, LangGraph, Dify, and OpenClaw, among other frameworks. See &lt;a href="https://en.bonree.com/en/aIObservability" rel="noopener noreferrer"&gt;https://en.bonree.com/en/aIObservability&lt;/a&gt;  and the post &lt;a href="https://dev.to/bonree-js/building-ai-observability-for-the-native-stack-architecture-design-and-engineering-practice-from-nc1"&gt;https://dev.to/bonree-js/building-ai-observability-for-the-native-stack-architecture-design-and-engineering-practice-from-nc1&lt;/a&gt;  for the underlying architecture.&lt;br&gt;
A question for readers&lt;br&gt;
If you're running agents on more than one of these frameworks, how are you keeping their telemetry comparable today: a shared schema you enforce yourselves, a vendor's auto-instrumentation, or have you not tried to reconcile them yet?&lt;/p&gt;

&lt;h1&gt;
  
  
  observability #opentelemetry #devops #Bonree
&lt;/h1&gt;

</description>
      <category>ai</category>
      <category>devops</category>
      <category>opentelemetry</category>
    </item>
    <item>
      <title>Why LLM Agent Token Costs Are Hard to Attribute, and What Telemetry Design Can Do About It</title>
      <dc:creator>Bonree-Js</dc:creator>
      <pubDate>Thu, 24 Sep 2026 08:53:49 +0000</pubDate>
      <link>https://dev.to/bonree-js/why-llm-agent-token-costs-are-hard-to-attribute-and-what-telemetry-design-can-do-about-it-iil</link>
      <guid>https://dev.to/bonree-js/why-llm-agent-token-costs-are-hard-to-attribute-and-what-telemetry-design-can-do-about-it-iil</guid>
      <description>&lt;p&gt;AI observability is the practice of collecting traces, metrics, and logs from LLM-backed applications so that latency, quality, and token cost can be explained per model call, per agent, and per conversation, not only per HTTP request.&lt;br&gt;
Disclosure: I work at Bonree, an observability vendor. Most of this post is vendor-neutral; Bonree ONE appears near the end.&lt;br&gt;
The invoice tells you the total, not the cause&lt;br&gt;
One user message to an agent can trigger several model calls, tool calls, and retries. Provider dashboards often aggregate by API key or project, which does not show which agent, which model choice, which team, or which conversation drove the spend. That gap is what telemetry design has to close.&lt;br&gt;
Design principles&lt;br&gt;
●Measure per model call, not per request. A request-level number hides which step in the chain was expensive.&lt;br&gt;
●Split by cardinality. Use metrics for totals, with a small fixed set of labels such as model, agent, team, and token type. Keep unbounded identifiers, such as conversation or user IDs, in traces, where per-conversation detail belongs. An ID used as a metric label creates a new time series for every ID.&lt;br&gt;
●Store token counts, convert to currency later. Prices change. Raw counts stay valid, and a price table applied at analysis time keeps history correct.&lt;br&gt;
●Treat prompt and completion capture as a separate decision. Content often contains customer data and needs its own review.&lt;br&gt;
●Know how your provider counts tokens. Cached and reasoning tokens may be reported differently across providers, so check what each field includes before comparing numbers.&lt;br&gt;
●Make spend attributable to a team, not only a model. Once agent, model, and token-type labels exist on the same metrics, a team-level rollup is a query, not a separate reporting pipeline, and it turns a cost question into a governance one: which teams are running agents, at what volume, and with what effect.&lt;br&gt;
Use a standard path&lt;br&gt;
OpenTelemetry's GenAI semantic conventions define shared names for LLM and agent telemetry. They are still evolving, so pin versions and re-check on upgrade. Keeping collection vendor-neutral also turns the choice of backend into a configuration decision rather than an instrumentation rewrite, which matters to any DevOps or SRE team that expects to change tools someday.&lt;br&gt;
A minimal example&lt;br&gt;
The snippet below is vendor-neutral OpenTelemetry code that illustrates the metrics-versus-traces split. It is not a Bonree ONE configuration. Tested with opentelemetry-sdk 1.44.0 and opentelemetry-semantic-conventions 0.65b0; the GenAI conventions are still evolving, so re-check attribute names when you upgrade.&lt;br&gt;
python&lt;/p&gt;

&lt;p&gt;from opentelemetry import trace, metrics&lt;/p&gt;

&lt;h1&gt;
  
  
  Assumes a TracerProvider and MeterProvider are configured elsewhere.
&lt;/h1&gt;

&lt;p&gt;tracer = trace.get_tracer("agent.instrumentation")&lt;br&gt;
meter = metrics.get_meter("agent.instrumentation")&lt;/p&gt;

&lt;p&gt;token_usage = meter.create_histogram(&lt;br&gt;
    "gen_ai.client.token.usage", unit="{token}",&lt;br&gt;
    description="tokens per call",&lt;br&gt;
)&lt;/p&gt;

&lt;p&gt;def call_llm(client, *, model, messages, agent, team, session_id):&lt;br&gt;
    with tracer.start_as_current_span(f"chat {model}") as span:&lt;br&gt;
        span.set_attribute("gen_ai.operation.name", "chat")&lt;br&gt;
        span.set_attribute("gen_ai.request.model", model)&lt;br&gt;
        span.set_attribute("gen_ai.agent.name", agent)&lt;br&gt;
        span.set_attribute("gen_ai.conversation.id", session_id)  # span only, never a metric label&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;    resp = client.chat(model=model, messages=messages)  # stand-in for your SDK call
    span.set_attribute("gen_ai.usage.input_tokens", resp.usage.input_tokens)
    span.set_attribute("gen_ai.usage.output_tokens", resp.usage.output_tokens)

    labels = {"gen_ai.request.model": model, "gen_ai.agent.name": agent, "team": team}
    token_usage.record(resp.usage.input_tokens, {**labels, "gen_ai.token.type": "input"})
    token_usage.record(resp.usage.output_tokens, {**labels, "gen_ai.token.type": "output"})
    return resp
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;Note that the conversation ID lives on the span only, while team is a small, fixed-cardinality label, so it is safe to put on the metric. In a quick local test with a stub client and two conversations across two teams, the metric produced four series (two teams × input/output), not one per conversation.&lt;br&gt;
Where Bonree ONE fits&lt;br&gt;
Bonree ONE 4.0 includes an AI Observability capability that Bonree describes as providing end-to-end tracing, span-level analysis, token usage monitoring, performance metrics, and session-level context tracking for large-model and agent environments. Bonree ONE also supports standard OpenTelemetry ingestion of traces, metrics, and logs. Whether a team-level rollup like the one above is useful on top of that depends on whether agent and team labels are already being recorded consistently, which is the instrumentation-design question this post is about, not something any specific platform does automatically. See &lt;a href="https://en.bonree.com/en/aIObservability" rel="noopener noreferrer"&gt;&lt;/a&gt;and &lt;a href="https://en.bonree.com/en/opentelemetrystandard" rel="noopener noreferrer"&gt;https://en.bonree.com/en/opentelemetrystandard&lt;/a&gt; , or the post &lt;a href="https://en.bonree.com/en/blog/aiobservabilityinpracticeinstrumentingagentchainsnotjustapicalls" rel="noopener noreferrer"&gt;https://en.bonree.com/en/blog/aiobservabilityinpracticeinstrumentingagentchainsnotjustapicalls&lt;/a&gt; . More at &lt;a href="https://en.bonree.com" rel="noopener noreferrer"&gt;https://en.bonree.com&lt;/a&gt; .&lt;br&gt;
A question for readers&lt;br&gt;
How does your team attribute agent token spend today: per API key, per agent, per team, or per conversation?&lt;/p&gt;

&lt;h1&gt;
  
  
  observability #opentelemetry #devops #Bonree
&lt;/h1&gt;

</description>
      <category>ai</category>
      <category>devops</category>
      <category>llm</category>
    </item>
    <item>
      <title>Building AI Observability for the Native Stack: Architecture Design and Engineering Practice from Bonree ONE 4.0</title>
      <dc:creator>Bonree-Js</dc:creator>
      <pubDate>Mon, 14 Sep 2026 07:10:09 +0000</pubDate>
      <link>https://dev.to/bonree-js/building-ai-observability-for-the-native-stack-architecture-design-and-engineering-practice-from-nc1</link>
      <guid>https://dev.to/bonree-js/building-ai-observability-for-the-native-stack-architecture-design-and-engineering-practice-from-nc1</guid>
      <description>&lt;p&gt;AI is changing observability and operations platforms in two directions at once, and it's easy to conflate them. The first direction is familiar: AI applications — RAG pipelines, LLM-backed customer service agents, coding copilots — are now part of the production stack, and they behave in ways traditional APM was never designed to see. The second direction is less discussed but arguably more consequential: AI is no longer only the thing being monitored. It's increasingly the thing doing the monitoring, and the thing taking action on what it finds.&lt;/p&gt;

&lt;p&gt;This article walks through how we approached both problems in Bonree ONE 4.0, the latest release of our observability and AIOps platform. Rather than a feature tour, the goal here is to explain the reasoning behind three capability areas we built — AI Observability, an AI-driven natural-language query layer called SmartAsk, and an autonomous agent workbench called Bonree ONE·Sage AI — and the production challenges that shaped their architecture.&lt;/p&gt;

&lt;p&gt;Six problems that don't show up on a status page&lt;/p&gt;

&lt;p&gt;A quick example sets the tone for all of them. An insurance customer of ours ran an AI phone agent for customer service. A caller asked how to purchase a policy; the agent misparsed the request and responded with instructions for cancelling one instead. No exception was thrown, no error rate moved, no health check failed — by every conventional signal, the call was a complete success. Hallucination and answer quality are orthogonal to uptime.&lt;/p&gt;

&lt;p&gt;That example generalizes into six patterns we kept running into, all of which shaped the architecture below:&lt;/p&gt;

&lt;p&gt;AI applications fail silently — confidently wrong output that no uptime or latency metric ever flags.&lt;br&gt;
Agent output gets trusted without being understood — code or decisions an operator can't debug when they eventually break.&lt;br&gt;
Bigger context windows don't remove the need for data governance — we've seen our own tooling get visibly worse at intent recognition once a session accumulates enough unfiltered context, well before hitting a hard token limit.&lt;br&gt;
Autonomy raises the security bar — an agent acts at machine speed, and skills that are individually safe can combine into risks nobody reviewed for.&lt;br&gt;
Cost doesn't fail gracefully — a single agentic query can burn hundreds of thousands of tokens, and unlike a slow response, an overrun is invisible until someone reads the invoice.&lt;br&gt;
Org structures haven't caught up — "AI does everything" scales to a team of one, not a team of hundreds, and how a human and an agent actually share an incident war room is still being worked out in practice.&lt;/p&gt;

&lt;p&gt;The rest of this article is about the architecture that came out of treating these as engineering problems rather than caveats: three capability areas in Bonree ONE 4.0 — AI Observability, a natural-language query layer called SmartAsk, and an autonomous agent workbench called Bonree ONE·Sage AI.&lt;/p&gt;

&lt;p&gt;AI Observability: treating AI applications as first-class citizens in the trace store&lt;/p&gt;

&lt;p&gt;The instinct when a new class of application shows up is to bolt a metric onto an existing dashboard. That doesn't hold up here, because the unit of work itself has changed. A single user request into a RAG or multi-agent application can fan out into a dozen internal steps — retrieval, several sequential or parallel model calls, tool invocations, result synthesis — each with its own latency, token cost, and independent chance of failure. Collapsing all of that into one opaque "request," the way a conventional APM trace would, throws away exactly the information needed to debug it.&lt;/p&gt;

&lt;p&gt;Show Image&lt;/p&gt;

&lt;p&gt;Collection. Instrumentation covers Python, Node.js, and Java, with automatic adaptation for common model-native APIs and out-of-the-box support for LangChain, LangGraph, Dify, and OpenClaw, among other agent frameworks. It's non-invasive — for a framework that's already supported, there's no code change required to start collecting data, which matters more than it might sound: requiring every team to hand-instrument their LangGraph pipeline before they get any observability is a real adoption barrier, not a minor inconvenience. Collection also speaks OpenTelemetry natively — Traces, Metrics, and Logs travel over OTLP — so AI telemetry isn't a second, proprietary data path sitting next to whatever teams already run for the rest of their stack.&lt;/p&gt;

&lt;p&gt;Redaction is handled at collection time, not after storage. Prompts and completions routinely contain customer data, and capturing full input/output at every span means you've effectively built a PII pipeline unless sensitive content is filtered before it lands anywhere persistent. The platform supports configurable masking rules and automatic PII identification and filtering applied at the point of capture.&lt;/p&gt;

&lt;p&gt;Processing pipeline. Once collected, data moves through real-time desensitization and format normalization, then metric extraction (performance indicators land in structured storage), then an analysis layer that runs hallucination detection and answer-quality scoring. This last stage is what turns raw traces into something closer to a quality signal rather than just a performance one — necessary precisely because, as the insurance example showed, quality problems don't otherwise register as failures at all.&lt;/p&gt;

&lt;p&gt;What you actually see. The platform exposes this data through several coordinated views rather than one dashboard trying to do everything:&lt;/p&gt;

&lt;p&gt;An application overview lists every AI service with request volume, error rate, response time, and total token consumption, so an anomalous application is visible at a glance before you drill into anything.&lt;br&gt;
Call chain analysis lists individual AI invocations with full input/output content, response time, and status, filterable by application, trace ID, user ID, or session ID.&lt;br&gt;
Call chain detail is where debugging actually happens: a Call Tree renders the trace as a time-series Gantt chart, with an "ALL" mode showing every span type (HTTP, chain, prompt, llm, parser) and an "LLM" mode that filters down to just the model-related nodes for focused analysis of the reasoning layer. A companion Call Map renders the same trace as a topology graph, with each node showing average response time, request count, and token count. Clicking any node opens a detail panel with four tabs — the input/output content and attributes, a timing breakdown, the code stack that triggered the span, and any errors or logs — which is usually enough to go from "this was slow" to a specific, fixable cause, such as a system prompt that had grown past several thousand tokens after multiple rounds of tool output were appended to context.&lt;br&gt;
A performance view aggregates response time, request count, error count, and model request count with trend comparisons against the prior day.&lt;br&gt;
A token view breaks down total consumption by model, separating input and output tokens, with trend charts over time — the level of granularity needed to answer "which model or which application is actually driving this month's bill," rather than only a top-line total.&lt;br&gt;
A model view aggregates by model rather than by application, so teams running more than one LLM in production can compare call volume, average latency, error rate, and token cost side by side and make an informed choice about which model earns its cost for a given task.&lt;br&gt;
A session view is arguably the most important one for conversational and agentic systems specifically, because the most common quality failure — degradation as context accumulates — is invisible if you only ever look at individual requests. Session analysis aggregates every trace belonging to one multi-turn conversation: total token consumption, trace count, and duration at the list level; a per-trace breakdown showing input preview, response time, and LLM/tool call counts inside a session; a waterfall view of any individual trace's internal span hierarchy, color-coded by node type (Agent, Chain, LLM, Task); and — distinctly useful for multi-agent systems — an Agent collaboration topology that shows, round by round, which node in a multi-agent execution consumed how much time and how many tokens. That last view is what makes it possible to say "the second agent in the chain is where both the latency and the token spend are concentrated" instead of only knowing that a three-hour session was expensive without knowing why.&lt;/p&gt;

&lt;p&gt;Alerts integrate directly into the same application detail view, categorized by severity (fatal, critical, warning, general, reminder) with filtering by status and rule type, so a spike surfaced in AI Observability doesn't require switching to a separate alerting tool to investigate.&lt;/p&gt;

&lt;p&gt;SmartAsk: making the data conversational without giving up rigor&lt;/p&gt;

&lt;p&gt;AI Observability answers "what is happening inside my AI applications." SmartAsk answers a related but different question: "how do I get an answer out of all my observability data — AI-related or not — without writing a query." It's a natural-language interface positioned as an expert-level Q&amp;amp;A layer over the whole platform, not just the AI-observability data described above.&lt;/p&gt;

&lt;p&gt;The core technical claim here is "deep data understanding": no data modeling or field mapping step is required before you can ask a question. The system automatically resolves the semantics of metrics, logs, traces, and events, and infers the right data source, filter conditions, and time range from the question itself, rather than requiring the user to specify them explicitly the way a dashboard query builder would. In practice this means the barrier to getting an answer drops from "know the metric name, the right PromQL syntax, and which dashboard has it" to "describe what you want to know."&lt;/p&gt;

&lt;p&gt;Rather than starting from a blank prompt every time, the platform ships with more than 30 pre-built scenarios across six categories — health inspection, fault diagnosis, performance optimization, change evaluation, ops governance, and cross-cutting "fusion" scenarios that combine several data types in one answer. These aren't templates in the sense of fixed report formats; they're curated starting points distilled from common patterns across customer deployments, meant to be used as-is or adapted with different targets, time ranges, or thresholds.&lt;/p&gt;

&lt;p&gt;Answers come back in a consistent three-part structure: an AI-generated summary in plain language, the underlying raw data, and a visualization — so a reader can trust the headline conclusion, verify it against the actual numbers, or hand the chart to someone else without redoing the analysis. Conversations persist as history that can be continued with follow-up questions, and a query that turns out to be useful more than once can be saved and reused rather than re-typed. Results export to PDF or DOC directly, and — a detail that matters operationally — a one-off query can be promoted into a permanent dashboard widget with one action, so an ad hoc investigation and a recurring monitoring need aren't two different workflows requiring two different tools.&lt;/p&gt;

&lt;p&gt;The intended audience is deliberately broad: not just the on-call engineer who already knows the query language, but developers and business stakeholders who need an answer from operational data but have never written PromQL or SQL and shouldn't need to.&lt;/p&gt;

&lt;p&gt;Bonree ONE·Sage AI: from answering questions to taking action&lt;/p&gt;

&lt;p&gt;SmartAsk gets you an answer. Bonree ONE·Sage AI is built to go further — from a natural-language description of a problem or task to actually carrying it out, using models, tools, knowledge bases, skills, and pre-built or custom agents behind a single conversational interface. The framing we use internally is that this is meant to be the difference between AI as a tool you operate and AI as a colleague that operates alongside you — a distinction that sounds like marketing language until you look at what it requires architecturally, which is considerably more than a chat UI in front of an LLM.&lt;/p&gt;

&lt;p&gt;Layered architecture. Bonree ONE·Sage AI is built as seven layers, each addressing a distinct concern:&lt;/p&gt;

&lt;p&gt;Show Image&lt;/p&gt;

&lt;p&gt;Data source layer — observability signals (logs, metrics, traces, alerts, events), operational assets (CMDB, runbooks/knowledge bases, ITSM tickets), and file assets (scripts, environment variables, credentials). All three categories are treated as inputs an agent might need, not just telemetry.&lt;br&gt;
Connector layer — the Model Context Protocol (MCP) as the standard channel for agent-to-tool and agent-to-system communication, connecting to CMDB and ticketing systems; a CLI layer that does pre-processing before data reaches the model, specifically to reduce how many tokens MCP calls consume — a detail worth noting because it's a direct response to the cost problem described earlier, not a generic engineering nicety; plus conventional API and script access.&lt;br&gt;
Security review layer — every signal and asset entering the system goes through specification validation, permission checks, and a security review before it's usable, and this is deliberately the first line of defense, not a step bolted on after something is already running.&lt;br&gt;
Security sandbox layer — this is the layer that makes autonomous execution tolerable in a production environment. It provides resource, network, filesystem, process, and multi-tenant isolation for every agent execution. Because a complex task involving multiple collaborating agents can legitimately run for minutes or, in some cases, hours, isolation has to hold for the duration of a long-running, potentially unattended job, not just for a single quick call. On top of isolation, there's ingress filtering that blocks malicious instructions before they reach an agent, egress filtering that strips sensitive information from output, execution permission controls with a human-approval gate for consequential actions, and a requirement that agent and skill creation and publishing go through supervisor review before they're available to be invoked at all.&lt;br&gt;
Scheduling and orchestration layer — managed across four dimensions: reliability (fallback strategies, error retry, circuit breaking so one flaky tool call doesn't cascade into a stuck task), validity (schema validation to keep malformed data from silently propagating between steps), agility (a primary agent that interprets intent, decomposes a task, and routes sub-tasks to the right sub-agent, monitoring and correcting course during execution), and coherence (long- and short-term memory management, including memory decay and update mechanisms, so an agent adapts to a user's patterns over time without letting stale context degrade later reasoning — the same accumulation problem described in the AI Observability section, addressed architecturally here rather than just observed).&lt;br&gt;
Interaction control layer — web UI, multi-channel access, alert-triggered execution, scheduled/periodic tasks, and session management, covering both interactive and unattended invocation.&lt;br&gt;
Application scenario layer — the top-level entry points: development and testing, release and change management, inspection and maintenance, emergency recovery, disaster-recovery drills, and ops governance.&lt;/p&gt;

&lt;p&gt;Two orchestration modes, chosen per task, not tuned on a dial. For standardized, well-understood work — a routine inspection, a change-approval flow — a fixed workflow is the right tool: explicit steps, predictable execution order, straightforward to audit against, and it fails in enumerable ways. For open-ended fault diagnosis, where the root cause is genuinely unknown at the start and the next useful action depends entirely on what the previous one revealed, a scripted workflow breaks down almost immediately, because you can't pre-write a decision tree for a failure mode you haven't seen yet. That calls for the autonomous-decision mode, where the agent reasons step by step and decides what to check next based on what it just found.&lt;/p&gt;

&lt;p&gt;We've watched this second mode play out in a real diagnostic session, and it's a useful illustration of how the security sandbox and orchestration layers work together rather than as separate concerns. An agent was asked to investigate a Java service running in a container and attempt recovery if warranted, with no information given up front about whether the environment was Kubernetes or plain Docker, or where the container lived. It probed the environment, determined it was Docker, and located the target container, then worked through memory, thread, CPU, garbage-collection, and log analysis autonomously — the kind of multi-step investigation that would otherwise require an engineer to log in and do by hand. At two specific points, where the next command would have read deep internal container state, it stopped, displayed the exact command it intended to run, and waited for explicit human confirmation before proceeding. It didn't ask permission for read-only statistics gathering. It asked only for the operations classified as having genuine potential impact — a concrete instance of the "execution permission control + human approval gate" mechanism described in the sandbox layer above, not a one-off safety feature bolted on for that specific session.&lt;/p&gt;

&lt;p&gt;Building agents, not just using them. The workbench separates three roles — administrators handle model access and platform-wide security and audit configuration; creators build skills and agents in a dedicated workspace, using either the workflow or autonomous-decision construction method, then publish for their own use or submit for review to share more broadly; users consume whatever has been published, either the built-in library or anything added from an internal resource marketplace. The same person is commonly all three at different times — building a skill in the morning and consuming someone else's in the afternoon — and the platform is built around that overlap rather than assuming rigid role separation.&lt;/p&gt;

&lt;p&gt;Out of the box, the platform includes more than 40 MCP-based tools and is compatible with external MCP servers, more than 10 ready-to-use skills including a "deep service diagnostics" skill, and a knowledge base that accepts standard document formats (Markdown, TXT, PDF) as well as API-based ingestion, so an organization's own runbooks and incident playbooks — the kind of institutional knowledge that otherwise lives in a senior engineer's head — can be folded directly into what an agent draws on. Several pre-built expert agents ship with the platform, including a terminal/host diagnostics agent and a database-specialist agent built on patterns distilled from experienced DBA troubleshooting workflows, usable directly or scheduled to run as an unattended, always-on specialist.&lt;/p&gt;

&lt;p&gt;Why this is worth the architectural complexity. The value case breaks into four tiers that build on each other. The most direct is efficiency: routine inspection, troubleshooting, reporting, and alert handling can run end-to-end through an agent, freeing operators from repetitive work and materially shortening mean time to resolution once a team can pull full-stack data with one query instead of stitching it together by hand. The second is an asset value that's easy to undervalue until you've lost it: a senior engineer's troubleshooting instinct, previously undocumented and lost when they leave or change roles, gets encoded as a reusable skill or knowledge-base entry instead — turning tacit, personal know-how into a structured, transferable asset. The third is a shift in operating model, from reactive to anticipatory: scheduled inspection combined with baseline analysis surfaces risk before it becomes an incident, and direct integration with ticketing and CMDB systems closes the loop from detection through resolution rather than stopping at notification. The fourth is more strategic than operational: full execution logging supports the audit requirements that regulated industries — finance, government — actually need, which is a precondition for deploying autonomous agents in those environments at all, not an optional nice-to-have.&lt;/p&gt;

&lt;p&gt;What's still unresolved&lt;/p&gt;

&lt;p&gt;It would be dishonest to present all of this as a solved problem, and a few things are worth naming directly.&lt;/p&gt;

&lt;p&gt;Skill-level security review is necessary but not sufficient. A skill that queries a monitoring API and a skill that restarts a process can each look completely safe in isolation and still combine into something neither reviewer anticipated once an agent chains them together at runtime in a sequence nobody explicitly tested. Static, pre-publish review catches the safety of individual components; it does not catch emergent risk from composition. The mitigations available so far — runtime ingress/egress filtering that doesn't depend on having anticipated the specific dangerous combination in advance, and treating the human-confirmation gate as a backstop rather than a complete answer — are partial, not complete, and we don't think this is a solved problem industry-wide.&lt;/p&gt;

&lt;p&gt;Cost governance is a genuine, ongoing tension rather than a one-time optimization. The internal example cited earlier — token spend growing substantially even as the explicit goal of adopting AI tooling was cost reduction — isn't unusual, and treating token consumption as an engineering-visible signal (per model, per session, per conversation turn, as described in the AI Observability section) is necessary precisely because that tension doesn't resolve itself; it has to be actively monitored and traded off against agent autonomy on purpose.&lt;/p&gt;

&lt;p&gt;And organizationally, the question of how a human on-call engineer and an autonomous diagnostic agent actually collaborate during a live incident — who has authority to decide, how escalation works when the agent's confidence is low, what a shared "war room" looks like when one participant is a model — is still being worked out in practice at most organizations we've talked to, ourselves included. Tooling can support that collaboration, but it can't yet fully define it.&lt;/p&gt;

&lt;p&gt;Where this leaves the three pieces&lt;/p&gt;

&lt;p&gt;AI Observability, SmartAsk, and Bonree ONE·Sage AI aren't three independent features so much as three layers of the same problem. Observability tells you what's actually happening — including inside the AI applications that used to be a blind spot entirely. SmartAsk makes that information reachable in natural language, without requiring every stakeholder to learn a query syntax first. And Bonree ONE·Sage AI is where understanding turns into action, under a governance model — layered isolation, staged review, and a confirmation gate scoped specifically to consequential operations — built for the fact that the system doing the acting is no longer only a human being.&lt;/p&gt;

&lt;p&gt;None of this makes AI-native operations a solved problem. But it's a different, and more specific, problem than "add a chatbot to the dashboard," and the architecture ends up looking correspondingly different once you take that seriously.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
    </item>
  </channel>
</rss>
