I wanted a cheap judge for my video pipeline: given a shot description, does it have a concrete subject, a visible action and an explicit camera decision, or is it generic? Strands Decider 2B is built for that kind of yes/no call, so I measured it.
- 75 shot descriptions: 63 captured from my pipeline (22 from production, 41 from a weak model running the same step) and 12 adversarial, with the three features annotated before running any judge.
- With the
choiceprimitive the 2B reaches AUC 0.91 in 34 ms per judgment. Claude Opus reaches 0.99. A regex that only looks for framing words reaches 0.90. - With its default primitive,
noul, it caught none of the 31 generic shots at the 0.5 threshold. The v21 checkpoint does not improve it. - Repeated scenes in a shot list, my real failure, are the easy part: the 2B, Opus and Haiku got 13 of 13 lists right.
- Shots, labels, every judge's output and the scripts are public: a CC-BY-4.0 dataset and a gist that reproduces every figure.
Read the full measurements → https://efraingaray.com/en/blog/juez-decision-2b/
Top comments (0)