The Judge That Never Guesses
How an AI pipeline that fixes its own bugs still refuses to trust its own opinion.
Every autonomous system eventually faces the same temptation: let the thing that did the work also decide whether the work was any good. It's efficient. It's tidy. It's also how you end up with a pipeline that reports success on a run that quietly made everything worse. Klyro was built to resist that temptation on purpose, and the reasoning behind it is worth walking through in full.
Don't Let the Model Grade Its Own Homework
There's an obvious shortcut Klyro doesn't take: asking the same LLM that proposed a fix whether the fix actually worked. It would take one extra prompt to wire up, and it would be easy to trust right up until the first time it's wrong in a way nobody catches. Language models are persuasive by design. Ask one to defend its own patch and it will find a reason the numbers look fine, even when they don't.
So the verdict on every run comes from deterministic code, never a model call. No prompt, no interpretation, no room for a confident-sounding rationalization. Just a fixed set of numeric thresholds applied to numbers a load test actually produced. The Investigator gets to be creative, proposing whatever fix it thinks will help. The Evaluator does not get that luxury, and that asymmetry is the whole point.
The Trap of Self-Grading
It's worth being specific about why this matters. An LLM asked to evaluate its own output isn't lying, exactly. It's doing what it was trained to do: produce a plausible, coherent answer to the question it was asked. "Did this help?" is a question with a plausible-sounding "yes" available almost regardless of the data, especially once the model has already committed, in an earlier turn, to believing the fix was a good idea.
That failure mode doesn't show up in a demo. It shows up three weeks later, when someone finally cross-checks a "validated" run against the raw metrics and finds the p95 latency barely moved. By then the bad verdict has already shaped a decision. Klyro closes that gap by never letting the question reach a model in the first place.
Changed Is Not the Same as Validated
The Evaluator draws a specific, unforgiving line between two things that sound similar and aren't. Performance can change after a patch without that change meaning anything worth celebrating. A fix might shave milliseconds off one endpoint while quietly pushing error rates up somewhere else, and "changed" would technically be true.
"Validated" is a stricter claim, and Klyro only makes it when three conditions hold at once:
| Signal | Requirement |
|---|---|
| p95 latency | Improves by at least 10% |
| Error rate | Moves by no more than 0.5 percentage points |
| CPU utilization | Stays at or under 95% |
Miss any one of those and the run reports a change, not a validated optimization. That distinction, one word doing a lot of quiet work, is where a system that's honest about what an LLM's fix actually accomplished lives. It's the difference between "something happened" and "something worth trusting happened."
Same Workload, Same Data, Every Time
None of those thresholds mean anything if the before and after runs aren't genuinely comparable. A p95 improvement measured against a lighter workload or a warmer cache isn't evidence of anything except that the test changed, and a system that lets its comparison drift is a system quietly grading on a curve without admitting it.
Klyro treats that risk the way a controlled experiment would. Task CPU, memory, and replica count stay identical across both runs. The exact same k6 workload hits both. The database gets the same discipline: db-init runs a full DROP and CREATE SCHEMA followed by a fresh seed before every single run, never an INSERT-only reset that could leave leftover rows from a previous attempt quietly nudging the numbers one way or another. If the "after" run looks better, it has to be because the patch made it better. There's no easier lane for it to have found instead.
The Patch Has to Prove It Is What It Claims to Be
The trust story doesn't stop at measurement. It starts earlier, at the moment a patch is proposed at all. Before any patch from the Investigator gets anywhere near a rebuild, it has to carry the original_sha256 of the file it targets, checked against that file exactly as it exists right now. A mismatch means the patch doesn't apply, full stop, no partial credit and no best-effort merge.
That check is paired with something even simpler: a three-file allowlist the Investigator is boxed into, covering only the files a performance fix should ever need to touch. Together, those two guardrails mean the thing the Evaluator eventually judges is provably the exact change that was proposed, applied to the exact file it claimed to target. Nothing smuggled in. Nothing silently dropped. By the time a number reaches the Evaluator, its provenance is no longer a question anyone has to take on faith.
Why the Guardrails Compound
Looked at individually, the sha256 check, the file allowlist, and the deterministic Evaluator each solve a narrow problem. Looked at together, they form a chain of custody that runs from the moment a fix is imagined to the moment a verdict is printed. A patch can't drift from what it claims to be. A verdict can't be shaped by the same reasoning that produced the patch. Each link only has to hold its own weight, which is exactly what makes the whole chain trustworthy instead of merely convincing.
Two Different Kinds of Trust
Trusting an LLM to notice a slow database query and propose a fix is one kind of trust. It's a bet on creativity and pattern recognition, and it's fine to lose that bet occasionally, because a bad patch just fails to validate. Trusting a system to tell you, honestly, whether that fix actually helped is a completely different kind of trust, and conflating the two is how systems end up reporting success that doesn't survive a second look.
Klyro keeps those jobs in separate hands on purpose. An LLM proposes. A guard verifies the proposal is what it claims to be. Plain deterministic code renders the verdict. None of the three steps has to trust the others' judgment, only their outputs, and that's precisely why the whole pipeline can be trusted at all.
Top comments (0)