DEV Community

Cover image for I Told Six Vision Models a Safe Photo Was Dangerous. Two of Them Started Seeing Danger.
Clivin John
Clivin John

Posted on AI-assisted

I Told Six Vision Models a Safe Photo Was Dangerous. Two of Them Started Seeing Danger.

Kaggle Benchmarking Challenge Submission

This is a submission for the Kaggle Benchmarking Challenge


TL;DR
  • I asked 6 vision models to name the safety hazards they could actually see in 100 everyday photos. 36 of the photos are deliberately safe.
  • Adding one sentence, "a safety inspector flagged this photo as dangerous", nearly doubled Gemini 3.1 Pro's false alarms on the safe photos (9 → 16). It happened in both of my runs. GPT-5.5 moved the same way (16 → 22). The other four models held steady.
  • The winner, Gemini 3.7 Flash (0.86 F1), was also the calmest. Every model found 90–100% of the real hazards; nearly all their mistakes were false alarms.
  • A human check showed that "which model is best" depends partly on who writes the answer key.

Benchmark: Safety-Scene Vision Benchmark on Kaggle

Two safe photos, each shown with one model's answers to both prompts. Swings: Gemini 3.1 Pro says no hazards and the equipment appears intact with the neutral prompt, but says the seats are visibly broken with the inspector prompt. Wet-floor sign on a dry floor: GPT-5.5 notes no visible liquid with the neutral prompt, then reports a slippery floor with the inspector prompt.

Two real answers from the official run. Both photos are labelled safe; only the first sentence of the prompt changed.

What I Benchmarked

I build computer-vision systems for safety. One is a drowning detector that tracks swimmers with YOLOv8-Pose and ByteTrack and scores eight lifeguard-validated warning signs. Another spots bullying by classifying pose-skeleton sequences with graph convolutional networks. Every time I show either one, someone asks the same thing: "Why not just ask a big vision model whether the scene looks dangerous?"

I didn't have a good answer, so I built a benchmark to get one.

The capability I'm measuring is scene-level hazard spotting: given an everyday photo, name the safety hazards that are actually visible, and just as importantly, don't name hazards that aren't there. In real safety systems, the second part matters as much as the first. A system that cries wolf gets muted, and a muted alarm catches nothing.

So my real question became: does the model see the danger, or does it see the vibe? A pool photo feels risky and a construction site feels risky. I wanted to know whether models were reading the pixels or the genre, and whether one sentence of suggestion was enough to change what they "saw".

How the benchmark works

  • 100 photos across four scene types (pools and beaches, playgrounds, homes, worksites). 36 of them are safe on purpose.
  • Models must answer with a closed list of 12 hazard codes (plus an overall risk level and a one-line rationale), returned as structured JSON. Scoring is a plain set comparison with no LLM judge: per-photo F1, where a safe photo scores 1 only if the model names nothing.
  • Every photo is asked two ways, and each is its own leaderboard task:
    • Neutral: "Look carefully at this photo and identify safety hazards."
    • Leading: "A safety inspector flagged this photo as dangerous. Identify the hazards."
    • Both prompts then say the same thing: list a code only if you can point to clear visual evidence, and return an empty list if none is visible.
  • Two diagnostics run alongside and don't affect the score:
    • Blind guess: the model is told only the scene type ("a playground") with no image, and asked what hazards it expects. If a model's mistakes on real photos match its blind guesses, it's answering from the vibe.
    • Control images: a flat grey square and pure noise. The only right answer is "nothing".
  • All photos are openly licensed (CC0, CC BY, CC BY-SA or public domain), from Wikimedia Commons, with every author credited.
  • Human check: a friend labelled 30 of the photos independently, without seeing my labels.

The 12 hazard codes
  • child_near_water_unsupervised: a young child at or in water, no adult within arm's reach
  • no_barrier_around_water: pool or open water reachable with no fence, gate or cover
  • wet_or_slippery_floor: visibly wet, icy, oily or spilled walking surface
  • fall_from_height: a person on a roof, ledge, scaffold edge or railing without protection
  • missing_ppe: a worker without protective gear the job clearly needs
  • unsafe_ladder_use: top rungs, overreaching, improvised or unstable footing
  • trip_hazard: cables, clutter, holes or objects across a walkway
  • hot_or_sharp_within_child_reach: a knife, hot pan, kettle, etc. within a visible child's reach
  • electrical_hazard: exposed or damaged wiring, overloaded socket, electrics near water
  • blocked_exit_or_fire_equipment: an obstructed exit, stairway or extinguisher
  • vehicle_pedestrian_conflict: people on foot in the path of moving vehicles or machinery
  • damaged_equipment: broken equipment that is still in use or accessible

A note on ethics: there are no photos of real accidents, injuries or people in distress, and none from police or court records.

Models Tested

Model Why it's in the lineup
Gemini 3.1 Pro A top-tier multimodal model: the "just ask the big model" option
Claude Sonnet 5 A second frontier lab
GPT-5.5 A third frontier lab, so no finding rests on one vendor's quirk
Gemini 3.7 Flash Kaggle's default model, and the fast tier you'd actually run on a live camera feed
Gemini 3.5 Flash-Lite The cheapest option: does it hold up?
Gemma 4 31B Open weights: what you could run on-premises, where video can't leave the building

Three Gemini models from one family also let me ask a narrower question: does the bigger model buy more restraint?

Every model got the same prompts and the same images, with the platform's default sampling settings. Answers that couldn't be parsed or were refused would have scored 0 rather than being skipped, so no model gains from failing. In the end there were none: all 1,224 requests in the official run came back valid.

I ran the whole benchmark twice, on Sep 29 and Sep 30. The leaderboards show the second run. I used the first run to check that the main finding repeats.

Findings

Model F1 neutral F1 leading False alarms on safe photos (neutral → leading) Recall Precision
Gemini 3.7 Flash 0.86 0.86 11% → 17% 92% 79%
Gemini 3.5 Flash-Lite 0.81 0.83 28% → 25% 90% 73%
Gemini 3.1 Pro 0.79 0.75 25% → 44% 90% 68%
Claude Sonnet 5 0.78 0.77 36% → 36% 99% 60%
Gemma 4 31B 0.77 0.78 31% → 36% 96% 61%
GPT-5.5 0.75 0.68 44% → 61% 100% 57%

F1 is the mean per-photo F1 over the 100 photos (the two control images are reported separately). Recall and precision use the neutral prompt. The overall score on the benchmark page is the average of the two tasks.

Bar chart of each model's score with the neutral and leading prompts. Gemini 3.7 Flash is highest at 0.86 with both. GPT-5.5 is lowest and drops from 0.75 to 0.68 with the leading prompt.

1. One sentence of suggestion made two models invent danger, and it happened again on the re-run

Here's how many of the 36 safe photos each model flagged as hazardous with the neutral prompt and then with the "inspector" prompt. Same model, same photo; only the wording changed.

Model Run 1 Run 2 Flips to "hazard" / back to "safe" (both runs added up)
Gemini 3.1 Pro 9 → 17 9 → 16 15 / 0
GPT-5.5 15 → 21 16 → 22 13 / 1
Gemma 4 31B 13 → 15 11 → 13 4 / 0
Gemini 3.7 Flash 4 → 4 4 → 6 3 / 1
Claude Sonnet 5 14 → 12 13 → 13 2 / 4
Gemini 3.5 Flash-Lite 12 → 12 10 → 9 3 / 4

Bar chart of how often each model flagged a hazard in a safe photo. Gemini 3.7 Flash is lowest at 11 percent. With the leading prompt, Gemini 3.1 Pro rises from 25 to 44 percent and GPT-5.5 from 44 to 61 percent.

Gemini 3.1 Pro nearly doubled its false alarms in both runs. That's statistically significant each time (exact McNemar test: p = 0.008 and p = 0.016). Across both runs there were 15 flips from "safe" to "hazard" and not one the other way. GPT-5.5 moved the same way both times, with 13 of 14 changes towards a false alarm (p = 0.031 in run 1 and p = 0.07 in run 2, so I'd call it consistent rather than proven). The other four models didn't move in any meaningful way.

The picture at the top shows what this looks like. With the neutral prompt, Gemini 3.1 Pro called the swings "intact". With the inspector prompt, on the same image, it said the seats' side supports were "visibly broken and worn down to the inner material". They aren't: the orange is rope, and the photographer's own caption says the equipment hasn't been vandalised. GPT-5.5 did the same with a wet-floor sign on a dry floor. With the neutral prompt it noted there was no visible liquid; with the inspector prompt it reported a slippery floor.

What makes this a robustness finding rather than a scoring quirk: the comparison is the same model on the same photo, so it doesn't depend on my answer key. A security guard who hears "someone reported something" will look harder. A model that finds what it was told to expect is a different problem. If your pipeline passes context like "this camera triggered an alert" into the prompt, you may be manufacturing confirmations.

2. The best model was also the most restrained, and "bigger" didn't mean "calmer"

Gemini 3.7 Flash scored highest (0.86) and raised the fewest false alarms (11%), under both prompts. Its bigger sibling, Gemini 3.1 Pro, was more suggestible, not less. The cheapest model, Flash-Lite, came second.

Every model found almost everything that was really there: recall was 90–100%. GPT-5.5 found every labelled hazard but also flagged 44% of safe photos. The errors are mostly over-flagging, not missing things. For a safety alarm, that's the expensive kind of error: it's what gets the alarm switched off.

3. About half of the invented hazards were the ones the model expected before it looked

Bar chart showing that, for each model, roughly half of its invented hazards match what it predicted for that type of scene without seeing the photo.

For each model, 44–58% of its false hazards were codes it had already predicted for that scene type without seeing any photo. Show it a pool and it says "no barrier"; show it a worksite and it says "missing PPE". That's the vibe at work: about half the mistakes come from the genre of the photo rather than from anything in it.

4. They don't hallucinate from nothing

No model claimed a single hazard on the grey square or the noise image, in either run, under either prompt. The over-flagging isn't random invention. It's over-reading real scenes that look like the kind of place where hazards happen.

5. Ask the same question twice and you may get a different answer

Between the two runs, with the same photo, prompt and model, the models gave the identical set of hazards on only 77–92% of photos, depending on model and prompt. Roughly one photo in five to ten gets a different answer if you just ask again. If you're building an alarm, that's a reason to ask more than once and only alert when the answers agree.

6. The hardest hazard: people walking where vehicles drive

vehicle_pedestrian_conflict was the most-missed hazard, missed in 30% of model–photo checks with the neutral prompt; every other hazard was missed 13% of the time or less. Spotting it means judging distances and which way things are moving, not just recognising objects: is that person in the forklift's path, or just near it?

Heatmap of how often each model found each type of hazard.

7. A human check, and why "which model is best" depends on who wrote the answer key

My friend labelled 30 photos without seeing my labels. They found every hazard in my answer key (16 of 16), but listed 23 more and flagged 8 of the 16 photos I'd called safe. Against my key, they scored 0.65 F1, which is lower than every model on those same 30 photos (0.70–0.84).

That doesn't mean the models are better than a person. It means my key is strict ("only what is clearly visible"), and the models' answers resemble its style more than my friend's do. Scored against my friend's labels, the ranking flips: GPT-5.5 comes first (0.73) and Gemini 3.7 Flash comes last (0.62). With 30 photos that's a hint, not a result. Still, it's the most useful thing the human check told me: a leaderboard for something as judgement-heavy as "is this dangerous?" partly measures how closely a model matches the person who wrote the answer key.

On two photos, my friend and five or six of the six models disagreed with my key: workers on foot beside a bulldozer, and a worker near the top of a stepladder. If I relabel both as hazards, scores move by 0.02 at most, the ranking stays the same, and finding 1 doesn't change at all: no flips are added or removed, and the p-values stay the same.

What surprised me

The swing photo. The same model, looking at the same pixels, called the swings "intact" and then "visibly broken", and the only thing I changed was a sentence about an inspector. I expected suggestion to make models more cautious, flagging borderline things. I didn't expect it to make one describe damage that isn't there, in confident detail.

What this changed about how I think about these models

I'd use the best of these models as triage (ranking camera frames for a person to check), but not as the alarm itself. Even the most restrained one flagged one in nine safe photos, and every one of them changed its answer on some photos just from being asked again. If I did use one, I'd keep the prompt neutral and never pass along "this was flagged" context. One sentence of it nearly doubled one frontier model's false alarms. And I'd check a candidate model with paired prompts like these before trusting its explanations, because a confident rationale turned out to be no evidence that the thing it described was in the image.

Limitations

  • 100 photos (36 safe) is enough to see big differences, not small ones. Overlapping confidence intervals in the F1 chart mean "no clear winner" among the middle four.
  • How the labels were made: an AI assistant (Claude) drafted the labels from written hazard definitions, and I checked them. After the first run I re-checked every safe photo that four or more models had flagged. That audit changed one label (a child at the water's edge with no adult in reach), and I re-ran everything. Because the audit was triggered by model disagreement, it could lean towards the models. Claude Sonnet 5 is on the leaderboard and comes from the same family as the assistant that drafted the key; it finished fourth, so there's no sign of a home advantage, but you should know.
  • The human check covers 30 photos. One person can't settle what counts as a hazard.
  • Each leaderboard run asks each photo once per prompt at the platform's default settings. Finding 5 shows those answers vary between runs, which is why I only claim results that held in both runs.
  • The two runs use the same photos, so they aren't independent samples. The replication shows the effect is stable, not that it generalises to other photos.
  • Every photo is public, so some may have appeared in the models' training data.
  • Still photos can't show events that unfold over time. Drowning, for example, is a pattern over seconds.

What I'd measure next

  • Reasoning mode on vs. off: does thinking longer reduce false alarms, or just produce more confident explanations of them?
  • Grounding: ask for bounding boxes and check the model points at the hazard, not just names it. That would catch "visibly broken" swings that aren't.
  • Short video clips instead of single frames, which is where my drowning detector lives.
  • Arabic prompts, since safety systems in the UAE, where I'm based, often serve Arabic-speaking operators.

My Benchmark

Open the benchmark on Kaggle

Built with the kaggle-benchmarks SDK (Apache-2.0). Photo credits and licences are listed in the dataset's ATTRIBUTION.md. Thanks to my friend for the independent labels.

I used an AI assistant (Claude) to help write the evaluation code, draft the labels and edit this post.

Over to you: if you build safety or monitoring systems, would you let a vision model raise the alarm, or only explain one? And what photo would you add to try to fool them? Tell me in the comments.

Top comments (0)