As part of my work at Google, we are publishing a suite of Agent Skills for Google products and technologies on GitHub. These agent skills are desi...
For further actions, you may consider blocking this person and/or reporting abuse
The ceiling rule has a measurable form that saved us from ourselves: a prompt earns its place in the suite only while the baseline still fails it - and one failing baseline attempt proves less than it feels like. A single baseline failure only bounds "solvable without the tool" at ≤95 %; the bound falls with independent attempts (1 − 0.05^(1/N)), so three baseline failures push it to ≤63 %, five to ≤45 %. We ran this admission rule over an existing suite and it quietly removed cases that had been measuring model capability rather than the tool. "Write harder prompts" becomes enforceable once admission is a measurement instead of an author's impression.
And a sixth rule we learned the expensive way, since your part 2 is about scorers: prove the suite can say no. Every grader gets a known-bad twin - a deliberately wrong output that must fail - plus a recorded date of the last time the suite actually rejected something. A grader that has never failed and a grader that silently stopped running print the same green, and without the twin you cannot tell trust from decoration.
Evaluation gets much more credible when the test set includes the failures that changed a production decision. I would keep a small regression suite of real bad outputs alongside aggregate scores, then make every model or prompt change earn its way through both.
Two things I'd add, both about what makes a score trustworthy over time rather than on the day you ran it.
A result is only as portable as the rig that produced it. Harness version, dataset snapshot, prompt revision, model version — if those aren't pinned to the number then you can't honestly compare this week's 82% to last month's 79%, and people absolutely will. Most eval "regressions" I've watched teams chase turned out to be someone quietly editing the test set. Version the suite the way you version the code and treat a change to it as a change that needs review, otherwise the number is a self-reported claim rather than evidence.
The unit test analogy is a good one but it breaks in a place worth naming. Unit tests are deterministic. Evals are sampled. A suite that passes 97% of the time isn't passing, it's a measurement with an interval around it, and gating on it like a boolean means you spend months chasing noise and then start ignoring the signal when it finally matters. Worth deciding up front how many runs a number needs before anyone is allowed to act on it.
Coming from payments, where nobody trusts a figure they can't reconstruct six months later, the interesting artifact was never the score. It's whether you can rerun it and get the same one.
The ceiling effect is the failure mode I see most when people evaluate coding agent skills. If the base model already knows the answer, the skill looks useful and the eval teaches you nothing about whether those tokens were earned. The part I would steal immediately is grading the final sandbox state instead of whether the agent walked through your preferred sequence of tool calls.
The prompt-grader mismatch caught us badly on a project last year. We had a tool-calling agent where the rubric penalized it for making more than two tool calls, but we'd never told the agent that in the task prompt. The agent was being creative and thorough, the grader was marking it wrong, and for two weeks we thought the model had regressed. The fix was embarrassingly simple, but the two weeks of confusion wasn't.
"Grade the destination, not the journey" is the right rule when the destination exists. The case I am in is the one where it does not yet: a construction system where the field outcome is months away and the first loop has not run. What we grade instead is the claim. Every output carries MEASURED, MODELED or ASPIRATIONAL with the evidence behind it, and the eval checks whether the label is right, not whether the number is. On the ceiling effect, our baseline problem is the mirror of yours. No model has deconstruction data, so the tool cannot fail to differentiate from pre-training, which makes overfitting to our own curated examples the thing to watch rather than an easy baseline.
Framing an eval as an action plus a scorer that asserts success is the shift that moved my hit rate, because loose "does this look right" checks never fail loudly enough to trust. The token-efficiency angle is underrated too: I cut eval-set cost a lot just by pruning cases that never change verdict across model swaps. Which of the five rules do you see teams break most often?
the 'evaluate plans instead of tasks' point is the one most teams skip first. we burned a sprint on multistep execution evals before realizing we had no step visibility, signal was always 'final output wrong', never which call failed.
switching to plan evals gave us granular failure modes in one pass. execution layer became separate, different harness.
ceiling effect hit us too — first eval set had 87% baseline accuracy and we declared success for 3 weeks before a user surfaced the failure.
are these generally single pass, or do you run multiturn evals for complex orchestration?
Great write-up based on the experience we gained while building google/skills.
Been burned by shiny evals that hid real flaws. Trust in AI evaluations needs transparent metrics and real user feedback - yes, even the messy ones. Curious what others think.