They Automated Everything Into CI/CD. The Pipeline Took Four Hours and Still Missed the Real Problem.
A team decided, reasonably enough on paper, that if AI testing mattered, it should all live in the CI/CD pipeline like everything else they automated. Every category, functional checks, security probes, fairness sampling, even subjective tone and helpfulness scoring, got wired into the same pipeline that used to run in minutes. It started taking close to four hours per run. Engineers began skipping it for minor changes, then for changes that weren't so minor. And the subjective quality scoring, the part that genuinely needed human judgment, got automated into a rubric-based score that technically ran fast and technically produced a number, and that number meant less than anyone wanted to admit, because some kinds of judgment don't compress cleanly into an automated check no matter how badly you want them to.
The actual question isn't whether to automate AI testing in CI/CD. It's what specifically belongs there, at what stage, and what genuinely shouldn't be forced into full automation just because everything else in the pipeline is. Here's how I'd make that call.
What Belongs at Commit Time: Fast, Cheap, Deterministic Checks
Every commit should trigger the fastest, cheapest category of AI-specific checks, the ones with a genuinely deterministic answer, schema and structured output validation, basic smoke tests confirming core functionality still works at all, checks that don't require multiple sampled runs to mean something. These need to run in something close to the same time budget as your existing unit tests, because if they don't, engineers will start treating them as optional the same way the team above did.
This tier deliberately excludes anything statistical or judgment-based. Save those for later stages. Commit-time checks exist to catch an obvious, structural break immediately, not to comprehensively validate quality on every single change.
What Belongs at Pull Request Time: Broader, Still Bounded
When a PR opens, run a broader but still time-bounded suite, targeted regression against a representative subset of your reference set, basic security checks against known adversarial patterns, enough coverage to catch a real regression before merge without asking every contributor to wait an hour for feedback on a small change. This is the tier most teams get wrong in one of two directions, either running almost nothing here and pushing everything to a slower nightly cycle, which delays real feedback on problems that could have been caught before merge, or running the full expensive suite here and making every PR miserable to work with.
The right size for this tier is whatever gives a contributor meaningful, fast signal on the change they actually made, not comprehensive coverage of everything the system does.
What Belongs on a Scheduled, Slower Cadence
Full statistical sampling, fairness testing needing real statistical power across matched examples, expensive multi-run consistency checks, anything requiring enough repeated runs or a wide enough example set that running it on every commit would be genuinely wasteful, belongs on a nightly or otherwise scheduled cadence, not gating every individual change. This is where the deeper, more expensive validation actually happens, and putting it here specifically protects the faster tiers from becoming too slow to actually use.
The trade-off is real and worth naming directly: a regression introduced early in the day might not surface until the nightly run catches it, hours later. That's an acceptable trade for keeping commit and PR feedback fast, as long as the nightly findings actually get acted on quickly once they arrive, not left sitting in a report nobody reads until the next incident forces someone to go looking.
What Shouldn't Be Fully Automated, and How to Build the Human Gate In Properly
This is the part the team in the opening story got wrong, and it's worth being honest about. Some evaluation genuinely needs human judgment, subjective quality calls, genuinely novel edge cases nobody's built a rubric for yet, anything where compressing judgment into an automated score would quietly launder a real gap in confidence. Forcing this into full automation doesn't actually solve the problem. It hides it behind a number that looks objective and isn't.
The right move is building a real human-in-the-loop gate directly into the pipeline, not routing around it. This means the pipeline can flag a change as needing human review and genuinely pause there, rather than either blocking automatically or, worse, faking a pass with an automated proxy score nobody fully trusts. A pipeline that's honest about where automation's confidence actually runs out is more useful than one that pretends everything can be scored.
Managing the Real Dollar Cost of AI-Specific CI Checks
This deserves its own explicit mention because it's a genuinely new constraint traditional CI never had to deal with. Many AI-specific checks carry a real, per-run cost, an API call to a judge model, compute for multiple sampled runs, and that cost is not the near-zero marginal cost of running a traditional deterministic test. Automating everything without thinking about this can produce a genuinely expensive pipeline that either burns budget fast or gets quietly throttled by whoever controls the spend, undermining the testing program in a way that has nothing to do with testing methodology.
Budget deliberately by tier, matching cost to how often each tier actually runs, cheap deterministic checks on every commit, moderate cost on PRs, the more expensive statistical runs reserved for the scheduled cadence where the cost is amortized across a day rather than paid on every single change.
A Visual Breakdown of What Goes Where

A Practical Checklist
- Commit-time checks are limited to fast, deterministic validation, kept in a time budget close to existing unit tests
- Pull request checks are sized to give contributors fast, meaningful feedback on their specific change, not comprehensive system-wide coverage
- Statistical sampling, fairness testing, and expensive consistency checks run on a scheduled cadence, with findings acted on quickly once they land
- A genuine human-in-the-loop gate exists in the pipeline for subjective judgment, rather than routing everything through an automated proxy score
- AI-specific check costs are budgeted deliberately by tier, matching real dollar cost to how frequently each tier actually runs
What I'd Actually Tell a Team Building This Out
Automating AI testing into CI/CD isn't a binary choice between doing it and not doing it. It's a series of specific decisions about what genuinely belongs at each speed and cost tier, and an honest acknowledgment that some things were never going to compress into a fast automated check no matter how much pressure there was to make the pipeline look comprehensive. The team that tried to automate everything didn't end up more rigorous. They ended up with a pipeline nobody trusted and a subjective quality score that meant less than the number implied.
Building CI/CD integration that actually respects these distinctions, fast where speed matters, thorough where thoroughness matters, honestly human where judgment can't be faked, is core to how PrimeQA Solutions approaches AI Testing Services engagements, because a four-hour pipeline nobody runs protects nothing, and a five-minute pipeline that quietly automated away real judgment protects even less.

Top comments (0)