DEV Community

Cover image for Strands Decider 2B vs LLM as a judge: a keyword regex tied the 2B
Efrain Garay
Efrain Garay

Posted on Originally published at efraingaray.com

Strands Decider 2B vs LLM as a judge: a keyword regex tied the 2B

I wanted a cheap judge for my video pipeline: given a shot description, does it have a concrete subject, a visible action and an explicit camera decision, or is it generic? Strands Decider 2B is built for that kind of yes/no call, so I measured it.

  • 75 shot descriptions: 63 captured from my pipeline (22 from production, 41 from a weak model running the same step) and 12 adversarial, with the three features annotated before running any judge.
  • With the choice primitive the 2B reaches AUC 0.91 in 34 ms per judgment. Claude Opus reaches 0.99. A regex that only looks for framing words reaches 0.90.
  • With its default primitive, noul, it caught none of the 31 generic shots at the 0.5 threshold. The v21 checkpoint does not improve it.
  • Repeated scenes in a shot list, my real failure, are the easy part: the 2B, Opus and Haiku got 13 of 13 lists right.
  • Shots, labels, every judge's output and the scripts are public: a CC-BY-4.0 dataset and a gist that reproduces every figure.

Read the full measurements → https://efraingaray.com/en/blog/juez-decision-2b/

Top comments (0)