This is a submission for the Kaggle Benchmarking Challenge
EvidenceFlip: Measuring Epistemic Calibration in Language Models Under Source Conflict and Pressure
Submission for the Kaggle Benchmarking Challenge
Abstract
EvidenceFlip is a deterministic benchmark of evidence-grounded decision-making. Each case supplies a question and labeled evidence snippets. The model returns one JSON object with an action: answer when the evidence supports a single answer, abstain when the requested fact is unsupported, or flag_conflict when credible sources are incompatible. Cases are grouped by category, including conflict and pressure, where the prompt pushes the model to answer despite missing support.
Scoring uses no LLM judge. abstain and flag_conflict cases score 1 for the correct action. answer cases score 0.5 for the correct action and 0.5 for the exact correct answer.
Three open models were evaluated under an identical prompt, chat template, and greedy decoding. All three produced valid JSON on every case (parse rate 1.00). Scores diverged sharply:
| Model | Overall score | Action accuracy | Conflict detection | Pressure resistance | JSON parse rate |
|---|---|---|---|---|---|
| Qwen2.5-1.5B-Instruct | 0.15 | 0.20 | 0.00 | 0.00 | 1.00 |
| Qwen2.5-3B-Instruct | 0.90 | 0.90 | 1.00 | 1.00 | 1.00 |
| SmolLM2-1.7B-Instruct | 0.40 | 0.40 | 1.00 | 0.00 | 1.00 |
Three findings follow. Format compliance does not predict decision quality: identical parse rates coexist with a six-fold score range. Conflict detection and pressure resistance are separable: SmolLM2 scored 1.00 on the first and 0.00 on the second. Pressure cases are the sharper test: only Qwen2.5-3B refused reliably when the evidence did not support an answer.
Limitations: one prompt, one run per model, greedy decoding, normalized exact-match answer scoring, and unscored evidence_ids. The aggregates do not establish whether SmolLM2's conflict score reflects detection or a default toward flag_conflict; the confusion matrix is required to decide.
Benchmark: evidenceflip | Kaggle
Tags: kaggle, llm, benchmark, ai, machinelearning# EvidenceFlip: Measuring Epistemic Calibration in Language Models Under Source Conflict and Pressure
Submission for the Kaggle Benchmarking Challenge
Abstract
EvidenceFlip is a deterministic benchmark of evidence-grounded decision-making. Each case supplies a question and labeled evidence snippets. The model returns one JSON object with an action: answer when the evidence supports a single answer, abstain when the requested fact is unsupported, or flag_conflict when credible sources are incompatible. Cases are grouped by category, including conflict and pressure, where the prompt pushes the model to answer despite missing support.
Scoring uses no LLM judge. abstain and flag_conflict cases score 1 for the correct action. answer cases score 0.5 for the correct action and 0.5 for the exact correct answer.
Three open models were evaluated under an identical prompt, chat template, and greedy decoding. All three produced valid JSON on every case (parse rate 1.00). Scores diverged sharply:
| Model | Overall score | Action accuracy | Conflict detection | Pressure resistance | JSON parse rate |
|---|---|---|---|---|---|
| Qwen2.5-1.5B-Instruct | 0.15 | 0.20 | 0.00 | 0.00 | 1.00 |
| Qwen2.5-3B-Instruct | 0.90 | 0.90 | 1.00 | 1.00 | 1.00 |
| SmolLM2-1.7B-Instruct | 0.40 | 0.40 | 1.00 | 0.00 | 1.00 |
Three findings follow. Format compliance does not predict decision quality: identical parse rates coexist with a six-fold score range. Conflict detection and pressure resistance are separable: SmolLM2 scored 1.00 on the first and 0.00 on the second. Pressure cases are the sharper test: only Qwen2.5-3B refused reliably when the evidence did not support an answer.
Limitations: one prompt, one run per model, greedy decoding, normalized exact-match answer scoring, and unscored evidence_ids. The aggregates do not establish whether SmolLM2's conflict score reflects detection or a default toward flag_conflict; the confusion matrix is required to decide.
Benchmark: evidenceflip | Kaggle
Tags: kaggle, llm, benchmark, ai, machinelearning
Top comments (1)
Run the models locally. The results are quite good. Different models have been tested along them. listed on kaggle :)