DEV Community

Cover image for Phish or Legit: Do LLMs Know When NOT to Cry Wolf?
Vladimir Joseph
Vladimir Joseph

Posted on

Phish or Legit: Do LLMs Know When NOT to Cry Wolf?

Kaggle Benchmarking Challenge Submission

Kaggle Benchmarking Challenge Submission by Anio Joseph

What I Benchmarked

Most security benchmarks ask: "Can the model detect the threat?" That's the easy part. The real question in a security operations center is: "Can the model stay quiet when the message is actually safe?"

I built Phish or Legit to test exactly that. It's a 40-item cybersecurity triage benchmark with 20 threats and 20 legitimate messages. The twist? Many of the legitimate messages are deliberately suspicious-looking — they have links, urgency, or unknown senders, but are completely benign.

For each item, the model must:

  1. Give a probability that it is a threat (0-100).
  2. Pick the single "tell" that gives it away (A-D, or E if it's legit).

The primary metric is Balanced Accuracy (the average of Threat Recall and Legit Specificity), but I also measured Brier Score to check calibration. A model that says "100% threat" on a benign email is just as dangerous as one that misses a real threat — it causes alert fatigue.

Models Tested

I ran the benchmark against 4 frontier models from three different providers, chosen for a mix of speed, cost, and capability:

  • gemini-2.5-pro (Google)
  • gpt-5.4-2026-03-05 (OpenAI)
  • gemini-3.7-flash (Google, smaller/faster)
  • claude-sonnet-4-5-20250929 (Anthropic)

(Note: deepseek-r1-0528 was attempted but failed to execute due to a platform-level timeout, so it is not included in the final results.)

The Findings

The results were surprising and reassuring at the same time.

Model Balanced Accuracy Brier Score Tell Accuracy
gemini-2.5-pro 1.000 0.00035 97.5%
gpt-5.4-2026-03-05 1.000 0.00190 95.0%
gemini-3.7-flash 1.000* 0.00129 96.9%
claude-sonnet-4-5 0.975 0.02289 92.5%

*gemini-3.7-flash completed 32 of 40 items.

1. Classification is (almost) solved. Calibration is not.

Three models achieved a perfect 1.0 Balanced Accuracy. They never missed a threat and never falsely flagged a benign message.

But look at the Brier Score. gemini-2.5-pro is exceptionally well-calibrated (0.00035), meaning its confidence levels are almost perfectly aligned with reality. claude-sonnet-4-5, while still very accurate, has a Brier Score 65x higher. This suggests that when Claude is wrong, it is confidently wrong — a dangerous trait in security triage.

2. The "Tell" is harder than the "Verdict"

Even the best model, gemini-2.5-pro, only scored 97.5% on Tell Accuracy. It correctly identified the item as a threat but picked the wrong clue. This tells me that while models are great at the binary decision, they struggle to articulate why something is a threat, which is crucial for analyst trust and explainability.

3. The False Positive Problem is Real

claude-sonnet-4-5 had a Legit Specificity of 0.95, meaning it incorrectly flagged 1 out of 20 legitimate messages as a threat. In a real SOC, a 5% false positive rate on a high-volume triage queue is a massive operational burden. The fact that the other models avoided this is a strong signal for their use in alert triage.

4. Smaller Models are Catching Up

gemini-3.7-flash (a "flash" tier model) achieved the same perfect Balanced Accuracy as its larger sibling, with a Brier Score that was better than GPT-5.4's. This is a huge insight for practitioners: you might not need the most expensive model for this specific triage task.

What This Means

This benchmark changed how I think about evaluating LLMs for security. Accuracy is a vanity metric. A model that scores 99% accuracy but is poorly calibrated is a liability. For real-world deployment, calibration and specificity matter more than raw detection rate.

The fact that all four models handled this task with near-perfect accuracy is a testament to how far LLMs have come. The next frontier isn't can they detect? It's can they detect with the right amount of confidence, and can they explain why?

My next steps would be to expand the benchmark with multilingual items and multi-turn conversation traps (e.g., a user asking "are you sure?").

My Benchmark

Author: Anio Joseph

Kaggle Benchmark: Phish or Legit Calibration Benchmark

You can explore the full benchmark, see the leaderboard, and read the item details on Kaggle.

Dataset: Phish or Legit Benchmark

Notebook: Phish or Legit Quickstart


This submission is part of the Kaggle Benchmarking Challenge.

Top comments (1)

Collapse
 
vladimir_joseph_247918be4 profile image
Vladimir Joseph •

good job !!!!