The most difficult AI workflows to debug are the ones that don't crash. A multi-step agent pipeline can quietly return a plausible-looking answer that's actually wrong, without ever throwing an error. One broken step early in the chain compounds into a wrong final output, and standard monitoring rarely catches it: the run finishes, the logs show green, and the problem only surfaces hours later when a downstream team flags that something's off.
This tutorial walks through a concrete case using openJiuwen Agent Core v0.1.18. You'll instrument a three-node pipeline (fetch → analyze → summarize) with OTEL-compatible spans (structured traces that record what each step did, in a format most observability tools understand). Then you will read back their output (latency, token counts, and a per-step failure flag), and pinpoint the exact node where the chain went wrong. The complete companion code is available at the tutorial repository.
Why silent failures escape standard logs
A silent failure happens when a step in an agent pipeline returns something that looks like a normal result but isn't. No exception is raised, no log line turns red, and no HTTP error code shows up, even though the output is wrong. Standard Python logging and HTTP status codes are built to catch exceptions. Agent-specific failures often don't rise to that level, so they slip straight past.
Three failure modes show up most often in production pipelines:
- Incorrect tool selection: The agent picks the wrong function for the job. That function still runs, still returns, and still fills the Observation (the agent's record of what a tool call returned) with data, just data that has nothing to do with the task. No exception fires, because nothing actually went wrong from the tool's point of view.
- Hallucinated tool arguments: The agent calls the right tool, but with parameters it made up or left out. If the tool tolerates missing input rather than rejecting it, the call succeeds and returns something empty or nonsensical.
- Context truncation: The prompt is too long for the model's context window (the amount of text it can consider at once), and the model API returns a normal HTTP 200 response with an empty list of completions. The token count comes back as zero, no error is raised, and the node exits as if nothing happened.
This matters more in a ReAct-style agent loop, which cycles through Thought, Action, and Observation, because each step can look successful on its own while the whole chain quietly drifts off course. A truncated completion still produces something that reads as a valid string, so the next Thought step picks it up and keeps going as if nothing were wrong. By the time the loop finishes, the final output is wrong, but there's still no exception and no red log line to explain why.
Sample workflow setup
You'll need Python 3.11, 3.12, or 3.13 (openjiuwen doesn't yet support 3.14). Clone the repository, change into the project directory, and install the dependencies:
git clone https://github.com/manishh/openjiuwen-debugging-an-ai-workflow-that-fails-silently-with-agent
cd openjiuwen-debugging-an-ai-workflow-that-fails-silently-with-agent
pip install -r requirements.txt
See the openJiuwen quickstart documentation for initial API key setup. The tutorial reads five environment variables: OPENJIUWEN_API_KEY, OPENJIUWEN_MODEL, OPENJIUWEN_API_BASE, OPENJIUWEN_PROVIDER, and OPENJIUWEN_SSL_VERIFY. If you leave out OPENJIUWEN_API_KEY, the code runs in offline simulation mode, which still emits and analyzes spans and only skips the live model call.
The tutorial's target pipeline is a sentiment-analysis workflow with three connected steps:
-
fetch_dataretrieves and preprocesses the source text. -
analyze_sentimentclassifies its sentiment. This is the node that simulates a silent failure. -
summarizegenerates a human-readable summary.
The Thought/Action/Observation loop above describes how a live ReActAgent reasons. The pipeline in src/main.py simulates the same failure pattern with three plain Python functions, so it runs deterministically offline without a live agent loop. The agent construction in examples/02_workflow_agent_setup.py looks like this:
card = AgentCard(
id="sentiment-analysis-pipeline",
name="sentiment-analysis-pipeline",
description="Fetch -> analyse -> summarise sentiment pipeline",
)
config = ReActAgentConfig().configure_model_client(
provider=PROVIDER,
api_key=API_KEY,
api_base=API_BASE,
model_name=MODEL,
verify_ssl=SSL_VERIFY,
)
# The agent below drives the three-node pipeline:
# node_fetch_data -> node_analyze_sentiment -> node_summarize
agent = ReActAgent(card=card).configure(config)
print(f"Agent '{agent.card.name}' created.")
ReActAgent is openjiuwen's current general-purpose agent, built around the Thought/Action/Observation loop described earlier. It decides at runtime which tool to call next, rather than following a hardcoded step order.
The older WorkflowAgent, which ran a fixed, user-defined node sequence, is deprecated as of openjiuwen 0.1.18 in favor of this AgentCard + ReActAgentConfig pattern. This tutorial's pipeline order (fetch → analyze → summarize) is enforced by the plain Python function calls in src/main.py, not by the agent, so the fixed order holds regardless of which agent class builds it.
Full-link observability in Agent Core
Full-link observability in Agent Core means that every tool call, token count, and step latency is captured as a structured, OTEL-compatible span (a self-contained record of one unit of work). Chain those spans together and you get a complete execution trace that you can query programmatically or view in Agent Studio, with no third-party APM (Application Performance Monitoring) tool required.
As of v0.1.18, Agent Core emits this telemetry natively. To capture and inspect those traces locally, you attach exporters to a TracerProvider instance. The full setup is in examples/01_setup_observability.py:
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import SimpleSpanProcessor, ConsoleSpanExporter
from opentelemetry.sdk.trace.export.in_memory_span_exporter import InMemorySpanExporter
exporter = InMemorySpanExporter()
provider = TracerProvider()
provider.add_span_processor(SimpleSpanProcessor(ConsoleSpanExporter()))
provider.add_span_processor(SimpleSpanProcessor(exporter))
# Always get the tracer from your provider, not from opentelemetry.trace.get_tracer().
tracer = provider.get_tracer("agent-core.tutorial")
SimpleSpanProcessor exports each span the moment it finishes, so the console output and the trace report you'll build later stay in the order the spans actually ran. Get your tracer from your own provider instance (provider.get_tracer(...)) rather than from the global opentelemetry.trace module. If Agent Core has already initialized the global provider, calling trace.get_tracer() returns a tracer bound to the SDK's own exporter, and your spans go somewhere you can't see them.
With the tracer in hand, wrap every step in a custom span that records four diagnostic signals. The broken analyze_sentiment node in examples/03_custom_spans.py shows what a silently failing node looks like once it’s instrumented:
def node_analyze_sentiment(data: dict) -> dict:
with tracer.start_as_current_span("tool.analyze_sentiment") as span:
t0 = time.perf_counter()
# Simulate: context window overflow caused the model to return
# an empty choices list. Status code was 200 – no error logged.
time.sleep(0.15) # slightly elevated latency: the client waited for timeout
result = {"sentiment": "", "token_count": 0, "confidence": 0.0}
span.set_attribute("agent.tool.name", "analyze_sentiment")
span.set_attribute("agent.tool.token_count", result["token_count"])
span.set_attribute("agent.step.latency_ms", round((time.perf_counter() - t0) * 1000, 2))
# This flag is the key signal in the trace: token_count==0 is a silent failure.
span.set_attribute("agent.step.silent_failure", result["token_count"] == 0)
return result
Run the script (python examples/03_custom_spans.py), and the span summary makes the failure obvious:
Span summary:
tool.fetch_data tokens= 312 latency= 40.15 ms silent_failure=False
tool.analyze_sentiment tokens= 0 latency= 150.15 ms silent_failure=True
tool.summarize tokens= 144 latency= 50.13 ms silent_failure=False
Look at the analyze_sentiment row: tokens=0 and silent_failure=True, even though the node returned normally with no exception raised. Your latency numbers will vary slightly between runs, since they come from a real time.sleep() call, but that row is the one to watch either way.
The last line of the code snippet is the one that matters most. Every node sets agent.step.silent_failure, not just the broken one, and it’s True whenever token_count == 0. That's what turns "did this succeed?" into something you can actually query. Together, the four attributes (agent.tool.name, agent.tool.token_count, agent.step.latency_ms, and agent.step.silent_failure) are the signals you'll read in the next section. Because every node writes all four, the trace report has a consistent shape to scan no matter which node broke.
How to read the trace output
The diagram below shows the full-link trace for a broken run. The root workflow.session span wraps all three tool spans, and tool.analyze_sentiment is highlighted because its agent.step.silent_failure attribute is True.
flowchart LR
A["workflow.session<br/>root span"] --> B["tool.fetch_data<br/>tokens: 312<br/>latency: 40 ms<br/>silent_failure: False"]
B --> C["tool.analyze_sentiment<br/>tokens: 0<br/>latency: 150 ms<br/>silent_failure: True"]
C --> D["tool.summarize<br/>tokens: 144<br/>latency: 50 ms<br/>silent_failure: False"]
style C fill:#ff6b6b,color:#fff
Once the workflow finishes, iterate over exporter.get_finished_spans() to read each span back; each one maps to a single node. The reporting loop in examples/04_read_trace_output.py prints the broken and fixed runs one after the other:
def _report(label: str, spans: list) -> None:
print(f"\n{'='*72}\n{label}\n{'='*72}")
print(f"{'Span':<35} {'tokens':>7} {'latency ms':>11} {'failure':>8}")
print("-" * 72)
for s in spans:
a = s.attributes or {}
flag = "◄ FAIL" if a.get("agent.step.silent_failure") else "OK"
print(
f"{s.name:<35} {str(a.get('agent.tool.token_count','-')):>7} "
f"{str(a.get('agent.step.latency_ms','-')):>11} {flag:>8}"
)
Running it prints both reports, so you can compare them line by line:
========================================================================
BROKEN RUN (inject_failure=True)
========================================================================
Span tokens latency ms failure
------------------------------------------------------------------------
tool.fetch_data 312 40.0 OK
tool.analyze_sentiment 0 150.0 ◄ FAIL
tool.summarize 144 50.0 OK
workflow.session - - OK
========================================================================
FIXED RUN (inject_failure=False)
========================================================================
Span tokens latency ms failure
------------------------------------------------------------------------
tool.fetch_data 312 40.0 OK
tool.analyze_sentiment 87 40.0 OK
tool.summarize 144 50.0 OK
workflow.session - - OK
The comparison shows the pattern to look for:
-
A healthy node shows
token_count > 0,latency_mswithin the expected range for that step, andOK. -
A broken node shows
token_count == 0because the model returned nothing, elevatedlatency_ms, and◄ FAIL. In this simulation, the latency is elevated because the client waited out a timeout before accepting the empty response.
Running src/main.py produces the same pattern from a single script, without the side-by-side comparison. Its analyse_trace() function reads from the InMemorySpanExporter and prints:
⚠ First broken span: tool.analyze_sentiment
token_count = 0
latency_ms = 150.24
workflow.node= node_analyze_sentiment
In Agent Studio's online debugging environment, the same execution graph appears visually in the session view, with the broken span flagged in the error log alongside its attributes. Because Agent Studio reads the same structured trace data Agent Core emits, you don't need to run analyse_trace() by hand once you're working against a live session there.
Each failure mode leaves its own signature in the spans:
| Failure mode | Span signature |
|---|---|
| Context truncation |
token_count == 0 with elevated latency |
| Hallucinated arguments (no validation) |
token_count > 0, but the output is semantically empty or causes a schema error in a downstream span; latency is usually normal |
| Hallucinated arguments (with the validation guard below) |
token_count == 0, normal latency, and an agent.tool.arg_error attribute naming the bad argument |
| Incorrect tool selection | An agent.tool.name that doesn't match the expected step, with output that doesn't fit the task |
Three targeted fixes from span evidence
You now have the failing span in front of you. The next step is turning that evidence into a fix. Each of the three failure modes leaves a distinct span signature, and each signature points to one specific code change.
Context truncation (token_count == 0, elevated latency)
Add a token budget check before submitting the prompt. The fix in examples/05_apply_fixes.py guards the analysis node:
MAX_PROMPT_TOKENS = 3800 # leave headroom for the model's completion budget
def node_analyze_sentiment_fixed(data: dict[str, Any]) -> dict[str, Any]:
with tracer.start_as_current_span("tool.analyze_sentiment") as span:
t0 = time.perf_counter()
prompt_text = data.get("text", "")
estimated_tokens = _token_estimate(prompt_text)
if estimated_tokens > MAX_PROMPT_TOKENS:
# Truncate to fit. In production use a summarisation sub-call instead.
chars_allowed = MAX_PROMPT_TOKENS * 4
prompt_text = prompt_text[:chars_allowed]
span.set_attribute("agent.prompt.truncated", True)
# Simulate the model call with a valid response after truncation.
time.sleep(0.04)
result = {"sentiment": "positive", "token_count": 87, "confidence": 0.91}
...
The * 4 uses the rough rule of thumb that one token is about four characters of English text. In production, replace _token_estimate with a real tokenizer (tiktoken, for example). The span's agent.prompt.truncated attribute records that truncation happened, so you can monitor for prompts that keep hitting the ceiling.
Hallucinated tool arguments (token_count > 0 but output semantically empty, or argument schema errors in downstream spans)
Validate arguments at the tool boundary and record any error in the span. The invoke_tool_with_validation function in examples/05_apply_fixes.py shows the pattern:
def invoke_tool_with_validation(
tool_name: str, args: dict[str, Any]
) -> dict[str, Any]:
with tracer.start_as_current_span(f"tool.{tool_name}") as span:
span.set_attribute("agent.tool.name", tool_name)
valid, error = _validate_tool_args(tool_name, args)
if not valid:
# Reflect the validation error back into the span so the ReAct
# Observation step can self-correct on the next iteration.
span.set_attribute("agent.tool.arg_error", error)
span.set_attribute("agent.step.silent_failure", True)
span.set_attribute("agent.tool.token_count", 0)
return {"error": error, "token_count": 0}
...
Running the script exercises both a valid call and a rejected one, so you can see the guard working on real span output:
--- Fix 2: Argument validation (valid args) ---
result={'output': 'Result from analyze_sentiment', 'token_count': 95}
--- Fix 2: Argument validation (hallucinated / missing args) ---
result={'error': "Missing required args for analyze_sentiment: ['text']", 'token_count': 0}
The rejected call still produces token_count == 0, the same signal as context truncation. The difference is that it's caught before any model call happens, its latency stays normal, and the returned error names the missing argument. That's how you tell the two failure modes apart in a real trace: check the tool's result payload and the agent.tool.arg_error attribute, not just the token count. Returning the structured error in the Observation also lets a ReActAgent loop self-correct on the next iteration rather than continuing with a bad result.
Incorrect tool selection (wrong agent.tool.name in the span, output mismatches intent)
Here the span shows a tool name that doesn't match the expected step. The fix is on the prompt side, so the agent picks the correct function on the first attempt:
- Tighten the tool description schema.
- Add a worked example of the correct tool call.
- Constrain the output format, so the agent has less room to substitute a plausible-sounding alternative.
This fix changes the prompt rather than the tool code, so it has no span-level output of its own to show, and examples/05_apply_fixes.py doesn't demonstrate it.
Verifying the pattern holds
The smoke tests in tests/smoke_test.py check that the underlying span mechanics work: that a span with token_count == 0 is picked up as silent_failure == True, and that a multi-span workflow isolates the one broken node from the healthy ones.
def test_otel_silent_failure_detection():
...
broken = [s for s in finished if s.attributes.get("agent.step.silent_failure")]
assert len(broken) == 1, "Expected one silent-failure span"
assert broken[0].attributes["agent.tool.token_count"] == 0
These tests build their own spans directly rather than importing from examples/05_apply_fixes.py. They confirm that the detection pattern itself is sound, not that any specific fix in this section works. After changing a fix, rerun the example script and compare its span report against the output shown above, then run pytest tests/smoke_test.py to confirm the detection logic underneath still works.
The output of smoke_test.py should show PASSED for five tests.
collected 5 items
tests/smoke_test.py::test_openjiuwen_sdk_imports PASSED [ 20%]
tests/smoke_test.py::test_otel_span_creation PASSED [ 40%]
tests/smoke_test.py::test_otel_silent_failure_detection PASSED [ 60%]
tests/smoke_test.py::test_otel_multi_span_workflow PASSED [ 80%]
tests/smoke_test.py::test_openjiuwen_agent_construction PASSED [100%]
Note that run_workflow(inject_failure=False) in src/main.py shows what a healthy run looks like after a fix: it returns the same token_count=87 result as the guard above, but it doesn't contain the fix logic itself. The guard, the validation, and the prompt-tightening code live in examples/05_apply_fixes.py; main.py only toggles between the broken and fixed outcomes to show the contrast in the trace report.
Closing thoughts
The debugging loop here is repeatable, and it's worth building into a habit rather than treating as a one-off fix:
- Instrument every tool node with the same four span attributes:
agent.tool.name,agent.tool.token_count,agent.step.latency_ms, andagent.step.silent_failure. - Run the workflow and read back the finished spans.
- Scan for zero-token observations and latency spikes, and the broken node names itself.
The same idea extends to multi-agent collaboration. Agent Core v0.1.18 added ReliabilityRail to the agent_teams package, which you can attach to a team of agents. Its detectors watch for repeated tool calls, model errors, tool errors, output-length problems, context-compaction issues, and agents bouncing work back and forth. When one trips, it can correct the problem locally, escalate to the team Leader, or notify you. Once you're running several agents in one session, those signals roll up into a per-session error rate, automating much of what you just did by hand.
To explore the SDK and start tracing your own agents, see the openJiuwen documentation. To see full-link observability and Agent Studio in action, contact the openJiuwen team for a guided walkthrough.
Frequently asked questions (FAQs)
How do I know which node caused a wrong output when no exception was raised?
Scan your finished spans for agent.step.silent_failure: True combined with agent.tool.token_count == 0. That combination identifies the exact node where the model returned nothing. Elevated agent.step.latency_ms alongside a zero token count narrows the cause to context truncation specifically.
What's the difference between a context truncation failure and a hallucinated argument failure in the span data?
Context truncation produces token_count == 0 with elevated latency. Without argument validation, a hallucinated argument failure typically shows token_count > 0 (the model did respond), but the output is semantically empty or triggers a schema error in a downstream span, usually with normal latency. With the validation guard in place, the bad call is rejected before the model runs, so it shows token_count == 0 with normal latency and an agent.tool.arg_error attribute that names the problem.
Do I need Agent Studio to use full-link observability, or does it work locally?
Full-link observability works entirely locally. Attach an InMemorySpanExporter and a ConsoleSpanExporter to your TracerProvider, run the workflow, call provider.force_flush(), and iterate over exporter.get_finished_spans(). Agent Studio reads the same structured trace data and shows it visually, but emitting the spans requires no external service.
Why should I get the tracer from my own provider instance rather than the global opentelemetry.trace module?
If Agent Core has already initialized the global provider, calling trace.get_tracer() returns a tracer bound to the SDK's exporter. Your custom spans then go to the wrong destination and won't appear in your InMemorySpanExporter. Always call provider.get_tracer(...) on your own TracerProvider instance.
Can I apply the same four-attribute span pattern to a multi-agent team, not just a single pipeline?
Yes. The openjiuwen.agent_teams package lets you attach ReliabilityRail (available in v0.1.18+) to a team of agents. It adds configurable detection for repeated tool calls, model errors, tool errors, output-length anomalies, context compaction, and agents bouncing work back and forth. In production, the per-session error rates from Agent Core's structured traces become the signal you alert on across the whole team.
What should I do if prompts consistently hit the token ceiling after I add the budget check?
Replace the truncation approach with a summarization sub-call that condenses the input before it reaches the analysis node. The agent.prompt.truncated span attribute tells you how often truncation fires, so you can set a threshold and trigger the sub-call automatically when the attribute appears more than a configurable number of times per session.
How does the ReActAgent loop self-correct after a hallucinated argument error?
When invoke_tool_with_validation detects an invalid argument, it returns a structured {"error": ..., "token_count": 0} dict into the Observation. The ReActAgent reads that Observation on the next Thought step and can revise its tool call with corrected arguments, instead of continuing from a bad result with no signal that anything went wrong.
Top comments (0)