On 2025-06-03T14:22:18Z, the log line looked like any other reject:
gate=cover_admission result=reject score=0.75 threshold=0.60 model=cover-v3
I’m Fantasy Elf, and I work on AI safety. This is the story of a threshold that lied to us.
We were running 18 generated covers through a four-gate admission check. Seventeen came back rejected. The first gate measured shadow ratio—the amount of dark area in the image. The covers themselves were generated with a parameter called shadow_ratio: 0.6. The gate had a threshold of threshold: 0.60. Same number. Two different meanings.
In the generator, 0.6 was a target for the renderer’s color grading. It told the renderer “make the dark areas about sixty percent of the mood.” In the gate, 0.6 was the maximum dark area an approved cover was allowed to have. The renderer produced a batch of text covers with dark areas at 60% by its own internal math. But the gate measured pixels after semi-transparent feathering on the edges, so its number read closer to 0.75. Every cover except one got flagged as unacceptable.
We spent half a day regenerating images, tweaking contrast, adjusting the palette, rerunning the same four gates over and over, trying to push one number down. Nobody questioned the threshold itself. Then someone asked the obvious question: “Why is the gate threshold the same as the generation parameter?” That’s when we realized the two values were never the same thing. We had copied a renderer setting into a safety gate and called it a standard.
The lesson wasn’t complicated: detection thresholds and generation targets need to be defined separately. If you copy the generator’s parameter into the safety gate, you’re letting the thing you’re measuring set the ruler. The gate was supposed to be an independent judge. We had made it a mirror.
The first thing we did was split the two values in the config. The first attempt was stupid. We created gen.shadow_ratio and gate.shadow_ratio_max, then set both to 0.6 because that’s what the old config had. Nothing changed. The split only existed on paper. The gate still read 0.75, the covers still failed, and we still had no idea why 0.6 was the number in the first place.
So we made the gate require a comment on every threshold explaining why that number was safe. Not why it produced a nice mood. Why it was safe. If the comment said something like “matches generation parameter,” we sent it back. That forced us to treat each threshold as an independent decision. It caught several other places where we were using aesthetic targets as safety criteria. One gate had a saturation threshold that was just copied from a style preset. We didn’t even know why the style preset had that number. That whole check got rewritten.
Next we added a pre-check. Before running all 18 covers through the gates, we run a small sample of 5 through the same detection pipeline and compare the gate’s measurements with the generator’s internal stats. The first version printed the difference to a log. Nobody read it. Then we made it fail loudly: if the average difference between the generator’s value and the gate’s measurement is above a tolerance, the batch stops before the first cover is officially judged. That is what caught the feathering problem on the next iteration. The generator said 0.58, the gate measured 0.74, the batch stopped immediately, and we finally saw what we had been missing for half a day.
We also tried a more ambitious approach: a full histogram comparison on every image before admission. It was too heavy. A medium batch took 40+ seconds just to hash and compare, longer than the generation itself. And because it was slow, we only enabled it for scheduled batches, which meant an ad-hoc afternoon run slipped through without it. That was a “safety measure too heavy to run is not a safety measure” moment. We dropped the full histogram and kept the small sample pre-check plus the threshold comments.
The same lesson showed up in agent safety. We now apply it to every agent that calls an external tool: default is deny, and the agent must explicitly declare which resource it wants to touch. The OIDC token carries an audience field that pins it to a specific resource and a specific action. Credentials from one system are never valid in another. When an agent calls another agent, every outbound request carries a trace header and a depth counter. We started with a depth limit of 3, then raised it to 5 after measuring that the average legit multi-hop chain was around 2.7. Five didn’t cut off normal calls, but it gave us room to notice a runaway loop before it spent forever. The depth limit did not always live in code. It used to sit in a runtime config file, and one day we found a loop running 47 levels deep because the file had been edited and nobody could tie the change to a release. We moved the limit back into code. Changing it now requires a release.
We’re still living with a compromise. The pre-check runs by default, but we only fail hard on the critical covers. For the non-critical ones, we log the numbers and move on. It hasn’t been fully validated—we haven’t collected enough edge cases to say whether that’s the right call. But at least every run now ends with a line like rejected=0 sampled=5 mean_diff=0.15. If we don’t even record the mean difference, the next time 17 out of 18 covers are rejected, we might not notice until after lunch.
So if you’re running any kind of admission gate for generated content, check your thresholds first. Make sure the detection standard didn’t come from the generation parameters. That’s the whole story. That’s the thing that cost us half a day and a string of “why is this cover failing” meetings.
We’ve all stepped on this kind of hole before. The fix is boring: two numbers, two names, two reasons. And a log line that shows both.
Top comments (0)