<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Classified</title>
    <description>The latest articles on DEV Community by Classified (@iamclassified).</description>
    <link>https://dev.to/iamclassified</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4143571%2F87426864-f60a-4480-b3e5-a0b8c21aab5c.jpg</url>
      <title>DEV Community: Classified</title>
      <link>https://dev.to/iamclassified</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/iamclassified"/>
    <language>en</language>
    <item>
      <title>LIQUIDITY EVENT IDENTIFICATION An LLM Reasoning Benchmark</title>
      <dc:creator>Classified</dc:creator>
      <pubDate>Sat, 03 Oct 2026 21:45:44 +0000</pubDate>
      <link>https://dev.to/iamclassified/liquidity-event-identificationan-llm-reasoning-benchmark-5boi</link>
      <guid>https://dev.to/iamclassified/liquidity-event-identificationan-llm-reasoning-benchmark-5boi</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/kaggle-2026-09-23"&gt;Kaggle Benchmarking Challenge&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Benchmarked
&lt;/h2&gt;

&lt;p&gt;I built a Kaggle benchmark called &lt;strong&gt;Liquidity Event Identification&lt;/strong&gt; to test how well large language models can identify liquidity events from controlled price-action scenarios.&lt;/p&gt;

&lt;p&gt;The benchmark tests whether a model can distinguish between events such as liquidity sweeps, rejections, breakouts, and acceptance based on the evidence provided in each scenario.&lt;/p&gt;

&lt;p&gt;The benchmark contains seven scenarios:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Equal-high liquidity sweep and rejection&lt;/li&gt;
&lt;li&gt;Equal-low liquidity sweep and rejection&lt;/li&gt;
&lt;li&gt;Breakout followed by acceptance&lt;/li&gt;
&lt;li&gt;Brief break followed by rejection&lt;/li&gt;
&lt;li&gt;Near-identical scenarios where one detail changes the correct interpretation&lt;/li&gt;
&lt;li&gt;Conflicting multi-timeframe evidence&lt;/li&gt;
&lt;li&gt;A deliberately misleading explanation&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The benchmark is focused on reasoning about market structure and liquidity events. It is not designed to measure trading profitability or predict real market outcomes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Models Tested
&lt;/h2&gt;

&lt;p&gt;I ran the benchmark against 20 models available through Kaggle:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;GPT-6 Astra&lt;/li&gt;
&lt;li&gt;GPT-5.6 Sol&lt;/li&gt;
&lt;li&gt;GPT-5.6 Luna&lt;/li&gt;
&lt;li&gt;GPT-5.5&lt;/li&gt;
&lt;li&gt;GPT-5.4&lt;/li&gt;
&lt;li&gt;GPT-5.4 Mini&lt;/li&gt;
&lt;li&gt;GPT-5.4 Nano&lt;/li&gt;
&lt;li&gt;Claude Opus 5&lt;/li&gt;
&lt;li&gt;Claude Opus 4.6&lt;/li&gt;
&lt;li&gt;Claude Sonnet 5&lt;/li&gt;
&lt;li&gt;Gemini 3.7 Flash&lt;/li&gt;
&lt;li&gt;Gemini 3 Flash Preview&lt;/li&gt;
&lt;li&gt;Gemini 3.1 Flash-Lite Preview&lt;/li&gt;
&lt;li&gt;Gemini 3.1 Pro Preview&lt;/li&gt;
&lt;li&gt;Gemini 3.5 Flash&lt;/li&gt;
&lt;li&gt;Grok 4.20&lt;/li&gt;
&lt;li&gt;Qwen3 Next 80B&lt;/li&gt;
&lt;li&gt;GPT-OSS 120B&lt;/li&gt;
&lt;li&gt;DeepSeek R1&lt;/li&gt;
&lt;li&gt;Claude Sonnet 4.6&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I used models from multiple providers to compare how different models performed on the same scenarios.&lt;/p&gt;

&lt;h2&gt;
  
  
  Findings
&lt;/h2&gt;

&lt;p&gt;16 of the 20 model runs completed successfully. All 16 completed runs scored &lt;strong&gt;7/7&lt;/strong&gt;, resulting in &lt;strong&gt;112 correct answers out of 112 completed scenario evaluations&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The other four runs — Qwen3 Next 80B, GPT-OSS 120B, DeepSeek R1, and Claude Sonnet 4.6 — encountered execution errors before completing the benchmark. They were excluded from the accuracy calculation.&lt;/p&gt;

&lt;p&gt;The main result was that the benchmark did not differentiate the completed models. Every completed model achieved a perfect score.&lt;/p&gt;

&lt;p&gt;This suggests that the seven scenarios are not difficult enough to distinguish model performance.&lt;/p&gt;

&lt;p&gt;The 100% result also does not indicate that these models can trade profitably or reliably interpret real-world market conditions. The scenarios are controlled and synthetic, and the benchmark measures a specific reasoning task.&lt;/p&gt;

&lt;p&gt;For the next version, I plan to increase the difficulty with more scenarios involving ambiguity, conflicting evidence, misleading context, stronger multi-timeframe conflicts, and cases where the most obvious interpretation is incorrect.&lt;/p&gt;

&lt;h2&gt;
  
  
  My Benchmark
&lt;/h2&gt;

&lt;p&gt;My Kaggle benchmark:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.kaggle.com/benchmarks/tasks/iamclassified/liquidity-event-identification" rel="noopener noreferrer"&gt;Liquidity Event Identification — Kaggle&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The benchmark contains the task definition, scenarios, evaluation logic, and model evaluation results.&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>kagglechallenge</category>
      <category>ai</category>
      <category>machinelearning</category>
    </item>
  </channel>
</rss>
