Existing agent evals measure task success. They do not measure what happens when the agent's retrieval layer returns evidence that contradicts its training data. A new paper from Sun et al. (arXiv 2610.12360v1) introduces epistemic humility as an eval dimension: does the agent revise its answer, flag uncertainty, or double down on a wrong conclusion?
This matters for financial decision agents, regulatory compliance bots, and any system where stale training data meets real-time feeds. If your agent trades on outdated earnings guidance or ignores a contradictory SEC filing, task accuracy is not enough. You need to know whether it can detect the conflict and escalate.
The Eval Gap
Most agent benchmarks score final answers. They do not instrument the trajectory to see if the agent noticed conflicting evidence mid-execution. The paper identifies three behavioral dimensions:
- Identify: Does the agent detect the conflict during retrieval or reasoning steps?
- Solve: Does it revise its answer or reasoning chain when evidence contradicts prior beliefs?
- Escalate: Does it communicate unresolved uncertainty in the final output (e.g., "I found conflicting sources")?
The researchers tested four agent configurations under two conflict scenarios:
- Controlled conflict: Injected contradictory evidence into the retrieval context.
- Naturally occurring conflict: Multi-step execution where sources disagree or parametric knowledge clashes with retrieved documents.
Each conflict scenario had a matched no-conflict control to isolate the effect of knowledge conflict on agent behavior.
What They Found
High task accuracy does not guarantee epistemic humility. Some agents recognized conflicts in early steps but failed to maintain or resolve them in later steps. Others produced a confident final answer despite detecting contradictions during execution.
Key failure mode: the agent detects the conflict (logged in intermediate tool calls), revises its reasoning internally, but then outputs a final answer with no uncertainty marker. The eval harness sees a correct answer and scores it as success, even though the agent ignored unresolved evidence.
This is a plumbing problem. If your agent uses a ReAct loop or chain-of-thought prompt, the conflict might appear in step 3 of 7, but the final answer generator does not inherit the uncertainty flag. You need explicit state management to propagate "I am uncertain" from the reasoning trace to the output formatter.
Instrumentation Strategy
To measure epistemic humility, you need to inject contradictory evidence and track three signals:
- Tool call logs: Did the agent retrieve both sources? Did it call a verification tool or re-query?
- Reasoning trace: Does the chain-of-thought mention the conflict, or does it silently pick one source?
- Final output structure: Does the answer include a confidence score, a caveat, or an escalation flag?
Here is a minimal eval harness pattern:
class EpistemicHumilityEval:
def __init__(self, agent, conflict_injector):
self.agent = agent
self.conflict_injector = conflict_injector
self.trajectory_log = []
def run_eval(self, task):
# Inject contradictory evidence into retrieval context
modified_task = self.conflict_injector.apply(task)
# Execute agent with trajectory logging
result = self.agent.execute(
modified_task,
callbacks=[self.log_step]
)
# Score on three dimensions
identified = self.check_conflict_detection(self.trajectory_log)
solved = self.check_revision_behavior(self.trajectory_log, result)
escalated = self.check_uncertainty_communication(result)
return {
"task_accuracy": result.is_correct,
"identified_conflict": identified,
"revised_answer": solved,
"communicated_uncertainty": escalated
}
def log_step(self, step):
self.trajectory_log.append({
"tool": step.tool_name,
"input": step.input,
"output": step.output,
"reasoning": step.chain_of_thought
})
def check_conflict_detection(self, log):
# Look for re-queries, verification calls, or explicit conflict mentions
for step in log:
if "conflict" in step["reasoning"].lower():
return True
if step["tool"] == "verify_source":
return True
return False
def check_revision_behavior(self, log, result):
# Did the agent change its answer after detecting conflict?
initial_answer = log[0].get("candidate_answer")
final_answer = result.answer
return initial_answer != final_answer
def check_uncertainty_communication(self, result):
# Does the final output include uncertainty markers?
uncertainty_markers = [
"conflicting", "uncertain", "unable to verify",
"sources disagree", "confidence: low"
]
return any(m in result.answer.lower() for m in uncertainty_markers)
This pattern separates task success from epistemic behavior. You can score an agent as correct but not humble, or humble but incorrect.
Confidence Scoring vs. Uncertainty Flags
You have two options for surfacing uncertainty:
- Logit-based confidence: Extract token probabilities from the final answer and threshold them. This works for single-turn completions but breaks in multi-step agents where the final answer is assembled from tool outputs.
-
Explicit uncertainty flags: Add a structured output field (e.g.,
confidence_level,unresolved_conflicts) that the agent must populate. This requires prompt engineering or a schema-enforced output parser.
The paper shows that model-level interventions (e.g., adding "acknowledge uncertainty when sources conflict" to the system prompt) improve epistemic humility but often reduce task accuracy. The agent becomes more cautious, which helps in high-stakes domains but hurts on benchmarks that penalize "I don't know" answers.
Trade-offs in Agent Design
| Approach | Task Accuracy | Epistemic Humility | Latency | Complexity |
|---|---|---|---|---|
| No conflict handling | High | Low | Low | Low |
| Logit-based confidence | Medium | Medium | Low | Medium |
| Explicit uncertainty flags | Medium | High | Medium | High |
| Multi-step verification loop | Low | High | High | High |
The multi-step verification loop (where the agent re-queries or calls a fact-checker after detecting conflict) has the highest epistemic humility but the worst task accuracy on standard benchmarks. This is because most benchmarks assume a single correct answer and do not reward uncertainty acknowledgment.
Deployment Shape for Financial Agents
If you are building a financial decision agent that consumes real-time market data, regulatory filings, and analyst reports, you need to handle knowledge conflicts at the orchestration layer:
- Retrieval deduplication: Tag sources with timestamps and authority levels. If two sources conflict, surface both to the reasoning step.
- Conflict detection middleware: Before the agent commits to a final answer, run a conflict checker that compares retrieved evidence against parametric knowledge or prior tool outputs.
- Escalation routing: If the conflict checker flags unresolved uncertainty, route the task to a human reviewer or a higher-tier model instead of auto-executing the trade.
This adds latency but prevents catastrophic failures. A trading agent that buys on stale earnings data because it ignored a contradictory SEC filing is worse than an agent that escalates to a human when it detects the conflict.
Failure Modes
The paper identifies several failure patterns:
- Early detection, late forgetting: The agent detects the conflict in step 2, but by step 6 it has forgotten and outputs a confident answer.
- Silent preference: The agent retrieves both sources but silently picks the one that matches its prior beliefs, with no reasoning trace explaining the choice.
- Hedging to game the eval: The agent always outputs "sources conflict" even when they do not, to maximize the epistemic humility score.
The third failure mode is a classic Goodhart's law problem. If you optimize for uncertainty acknowledgment, agents will hedge on every answer. You need to pair epistemic humility metrics with task accuracy and penalize false positives (claiming conflict when none exists).
Observability Requirements
To debug epistemic humility failures, you need:
- Structured trajectory logs: Capture every tool call, reasoning step, and intermediate answer.
- Conflict annotations: Tag retrieval outputs with source metadata (timestamp, authority, contradiction flags).
- Diff tracking: Log when the agent revises its answer and why.
Without these, you cannot distinguish between an agent that genuinely detected a conflict and one that got lucky by guessing.
Technical Verdict
Use epistemic humility evals when:
- Your agent operates on real-time data that can contradict training cutoffs (financial markets, regulatory updates, breaking news).
- The cost of a confident wrong answer is higher than the cost of escalation (medical diagnosis, legal advice, trading).
- You need to audit agent behavior for compliance or safety reviews.
Avoid this eval dimension when:
- Your task has a single ground truth and no conflicting sources (closed-domain QA, deterministic code generation).
- Latency is critical and you cannot afford multi-step verification loops.
- Your benchmark penalizes "I don't know" answers and you are optimizing for leaderboard rank.
The core insight is that task accuracy and epistemic humility are orthogonal. You can build an agent that is accurate but not humble, or humble but not accurate. For high-stakes agentic systems, you need both, and you need eval infrastructure that measures them separately.
Top comments (1)
Some comments may only be visible to logged-in visitors. Sign in to view all comments.