This is a submission for the Kaggle Benchmarking Challenge
What I Benchmarked
I built a Kaggle benchmark called Liquidity Event Identification to test how well large language models can identify liquidity events from controlled price-action scenarios.
The benchmark tests whether a model can distinguish between events such as liquidity sweeps, rejections, breakouts, and acceptance based on the evidence provided in each scenario.
The benchmark contains seven scenarios:
- Equal-high liquidity sweep and rejection
- Equal-low liquidity sweep and rejection
- Breakout followed by acceptance
- Brief break followed by rejection
- Near-identical scenarios where one detail changes the correct interpretation
- Conflicting multi-timeframe evidence
- A deliberately misleading explanation
The benchmark is focused on reasoning about market structure and liquidity events. It is not designed to measure trading profitability or predict real market outcomes.
Models Tested
I ran the benchmark against 20 models available through Kaggle:
- GPT-6 Astra
- GPT-5.6 Sol
- GPT-5.6 Luna
- GPT-5.5
- GPT-5.4
- GPT-5.4 Mini
- GPT-5.4 Nano
- Claude Opus 5
- Claude Opus 4.6
- Claude Sonnet 5
- Gemini 3.7 Flash
- Gemini 3 Flash Preview
- Gemini 3.1 Flash-Lite Preview
- Gemini 3.1 Pro Preview
- Gemini 3.5 Flash
- Grok 4.20
- Qwen3 Next 80B
- GPT-OSS 120B
- DeepSeek R1
- Claude Sonnet 4.6
I used models from multiple providers to compare how different models performed on the same scenarios.
Findings
16 of the 20 model runs completed successfully. All 16 completed runs scored 7/7, resulting in 112 correct answers out of 112 completed scenario evaluations.
The other four runs — Qwen3 Next 80B, GPT-OSS 120B, DeepSeek R1, and Claude Sonnet 4.6 — encountered execution errors before completing the benchmark. They were excluded from the accuracy calculation.
The main result was that the benchmark did not differentiate the completed models. Every completed model achieved a perfect score.
This suggests that the seven scenarios are not difficult enough to distinguish model performance.
The 100% result also does not indicate that these models can trade profitably or reliably interpret real-world market conditions. The scenarios are controlled and synthetic, and the benchmark measures a specific reasoning task.
For the next version, I plan to increase the difficulty with more scenarios involving ambiguity, conflicting evidence, misleading context, stronger multi-timeframe conflicts, and cases where the most obvious interpretation is incorrect.
My Benchmark
My Kaggle benchmark:
Liquidity Event Identification — Kaggle
The benchmark contains the task definition, scenarios, evaluation logic, and model evaluation results.
Top comments (0)