My pipeline had a quality gate. It worked. That was the problem.
It checked structure: is the file the right size, does it open, does it contain what it should. By every one of those measures, the assets passing through were fine. Then I looked at them myself and scored a batch 2 out of 10. Ugly, generic, forgettable.
The gate had done its job and the output was still unsellable, because structure is not quality.
Layer one: the file is valid
The first layer is mechanical, and it should be. It answers questions with yes or no.
- Is the resolution above the minimum?
- Is the background actually transparent, or is it white pixels pretending?
- Does the vector open with editable paths?
- Is the file under the size cap?
This layer is cheap, fast, and never wrong about what it measures. It catches broken files. It does not catch boring ones, and it was never going to.
Layer two: judge the art, not the file
The second layer asks a different kind of question: looking at this, would a person want it? That is a judgement call, so I handed it to a vision model with a specific brief. You are a senior art director reviewing this for a marketplace. Score it.
This is where it got interesting, and where I nearly broke it.
The calibration trap
My first version of the prompt was too aggressive. It said, in effect, that the model should reject most AI generated art, because most of it is bad.
The result was worse than no gate at all. The model started flagging genuinely good assets, 8 out of 10 material, as mangled and unusable. It had learned from my prompt that its job was to reject, so it rejected.
I rewrote the prompt to be neutral and to anchor the scores explicitly. A 7 means this specific thing. A 3 means this other thing. No editorial stance, just a rubric. The gate became stable: it passed what a human would pass and flagged what a human would flag.
One more thing: different art needs different rubrics
The same model that judged my icon sets kept rejecting flat texture assets with the note "flat, no focal point".
The model was right about the observation and wrong about the conclusion. Flat texture is supposed to be flat. There is no focal point by design. A single generic rubric was punishing a whole category for being what it was.
So the rubrics split by asset type. Icons get judged as icons. Textures get judged as textures. Seasonal art gets its own criteria. Same model, different lens.
The result
A two layer gate: one mechanical, one perceptual. The mechanical layer never sleeps and never lies. The perceptual layer catches the thing the mechanical layer cannot see. Neither is enough alone.
If your pipeline has a quality gate that passes everything, it is probably checking structure. That is half the job. The other half is judging the output, and the trap there is calibration, not capability.
The scripts, the rubrics, and the calibration notes are in the Pipeline Starter Kit:
Top comments (0)