Deletion-based faithfulness tests ask a simple question. If I remove the token that an explainer considers most important, does the model's prediction change? The rate at which predictions change is called the flip rate. It is often used to evaluate whether an explanation reflects what a model actually depends on.
I tested this using LIME on distilbert-base-uncased-finetuned-sst-2-english across a pre-registered set of 30 sentiment inputs. The aggregate flip rate was 39.3%. That number initially looked reasonable. However, when I separated the results by confidence and input category, the picture changed considerably. On strong baselines, such as "This is the best product I have ever purchased", the flip rate was 0%, even though LIME's directional attribution was 80% correct. On lexical shortcut inputs, such as "Excellent but useless", the flip rate was 100%. Both behaviours contributed to the same aggregate metric, even though they represented very different situations.
1. The setup
I used LIME with num_samples=300, num_features=10 and five random seeds. The model was DistilBERT fine-tuned on SST-2, using revision 714eb0fa.
The test set contained 30 inputs divided into six categories.
| Category | Purpose |
|---|---|
| Negation and minimal pairs | Test semantic reversals |
| Lexical shortcuts | Test contradictory sentiment cues |
| Ambiguity | Test genuinely mixed sentiment |
| Distribution and style shift | Test medical and financial language |
| Strong baselines | Test unambiguous positive and negative statements |
| Edge cases | Test single tokens, repeated tokens and whitespace |
For each input, I removed the top-ranked LIME token from the original text and ran the model again. I recorded the original prediction, the new prediction, directional correctness and whether the label flipped. Two inputs were structurally excluded, leaving 28 tested inputs.
2. What the aggregate results showed
| Metric | Result | Threshold | Status |
|---|---|---|---|
| Mean top-5 Jaccard stability | 0.81 | ≥ 0.60 | Pass |
| Directional correctness | 89.3% (25/28) | ≥ 0.70 | Pass |
| Flip rate | 39.3% (11/28) | None | Reported |
Looking only at these results, it would be easy to conclude that LIME performed reasonably well. However, the aggregate results hide substantial differences between input categories.
| Category | Directional correctness | Flip rate | Mean absolute delta |
|---|---|---|---|
| Lexical shortcuts | 100% | 100% | 0.995 |
| Ambiguity | 100% | 40% | 0.408 |
| Distribution and style shift | 100% | 20% | 0.230 |
| Negation and minimal pairs | 80% | 60% | 0.590 |
| Strong baselines | 80% | 0% | 0.0002 |
| Edge cases | 75% | 25% | 0.162 |
The strong-baseline results stood out. LIME identified the important token correctly in 80% of cases, but none of the predictions flipped.
In contrast, all tested lexical shortcut inputs produced a flip.
The aggregate flip rate combines these two very different behaviours and hides the distinction between them.
3. The problem with high confidence
Consider the sentence: "This is the best product I have ever purchased." The original positive probability was 0.99987. LIME assigned the token best an attribution weight of +0.42. After removing best, the positive probability dropped to 0.9994. The direction was correct, but the prediction did not change. The average absolute confidence change for strong baselines was only 0.0002. The remaining words provided enough evidence for the model to retain its original prediction. Removing one token was not sufficient to change the label. Now consider another input: "The movie was terrible but I loved every minute of it." The original positive probability was 0.9999. LIME identified loved as the top token, with an attribution weight of +0.31. After removing loved, the positive probability dropped to 0.0002. The prediction flipped. Both inputs had high initial confidence. However, their responses to deletion were completely different.
This is the central problem with using flip rate alone. A prediction may not flip because the explanation is incorrect. It may also remain unchanged because the model is highly confident or has redundant evidence. A flip rate does not distinguish between these possibilities.
4. Confidence stratification revealed another problem
The confidence buckets showed why the aggregate result was difficult to interpret.
| Confidence | Inputs | Directional correctness | Flip rate | Mean absolute delta |
|---|---|---|---|---|
| p < 0.90 | 1 | 100% | 0% | 0.126 |
| 0.90 ≤ p < 0.99 | 4 | 100% | 50% | 0.506 |
| p ≥ 0.99 | 23 | 87.0% | 39.1% | 0.375 |
The high-confidence bucket had a flip rate of approximately 39%. However, this bucket combined strong baselines with lexical shortcuts. Strong baselines had a 0% flip rate, while lexical shortcuts had a 100% flip rate. The average therefore represented neither behaviour accurately. This is not simply a sample-size problem. It is a problem with what the metric measures. The same flip rate can arise from very different model behaviours.
5. Failures and limitations
There were several issues that needed to remain visible in the evaluation.
Structural exclusions
Input 9, "Fine.", became an empty string after deleting its top token. The model call was skipped and the input was recorded as skipped_single_token. Input 30 contained only whitespace. The LIME pipeline skipped it before the model ran, and it was recorded as skipped_empty. Both inputs remained in the row-accounting table. Including them in the denominator would have changed the reported flip rate from 39.3% to 36.7%.
Directionally incorrect cases
Three of the 28 tested inputs were directionally incorrect. The overall directional correctness was 89.3%. These cases matter because a high aggregate score does not mean every explanation behaved as expected.
Tokenizer mismatch
LIME uses whitespace tokenization, while DistilBERT uses WordPiece tokenization. Removing a word from LIME's perspective does not always correspond to removing the same units from the model's perspective. This can affect deletion results.
Limited sample size
The experiment used only 30 inputs, with five initially assigned to each category. Category-level rates therefore have wide uncertainty. These results demonstrate a methodological issue, but they should not be treated as precise estimates for a larger population.
Computational limitations
The experiment was CPU-only and used 300 LIME samples. Running more samples on a GPU could improve the stability estimates. However, it would not eliminate the saturation behaviour observed in the model.
6. A different evaluation protocol
To address these issues, I defined a confidence-stratified protocol with three main requirements.
First, report flip rates by confidence and category.
Aggregate flip rate should not be the headline metric. Results should be reported across confidence buckets and input categories so that different behaviours remain visible.
Second, pre-register category-specific predictions.
The expected behaviour for each category should be defined before running the evaluation. For example, the pre-registered prediction for strong baselines was a 0% flip rate despite directional correctness. This makes it possible to distinguish a predicted result from an unexpected one.
Third, include top-k deletion curves.
Single-token deletion may not be sufficient when a model is saturated. Testing deletion at k = 1, 3, 5 and 10 can help identify whether removing more tokens produces a meaningful change.
This is an exploratory follow-up in my protocol. It is not a result established by the primary experiment. Every input should also have a recorded status. Tested inputs, single-token exclusions and empty inputs should all remain visible in the final accounting.
Finally, confidence diagnostics need careful interpretation. The experiment's ECE and Brier scores use the model's own predictions as a ground-truth proxy. They are self-consistency diagnostics, not evidence of calibration against human labels.
7. Why this matters beyond LIME
Deletion-based evaluation is used in several explainability methods and benchmarks. However, measuring how much a prediction changes after removing important features does not, by itself, establish that an explanation is faithful. On a saturated model, the output may barely change even when an important token is removed. On inputs with contradictory evidence, removing one token may cause an immediate label change. These behaviours can produce similar or misleading aggregate results. Confidence stratification, category-level reporting, pre-registered predictions and top-k deletion curves provide additional context. They do not solve every problem, but they make the limitations of the evaluation more visible.
The main finding from this audit is that flip rate is influenced by both the explanation and the model's response to deletion. Reporting it without accounting for confidence can lead to conclusions that the underlying results do not support.
Explore the project
The project is available at XAI Forensics.
Corrections, disagreements and disconfirming results are welcome. The purpose of this work is to examine where evaluation metrics fall short and make those limitations visible.
Top comments (0)