How mutation testing, complexity metrics, and other pipelines goodies could reduce your team’s review fatigue
There is a new kind of tired in software engineering.
One that’s grown fat from AI code, and now lives in your PR reviews.
You’ve been tagged to review three PRs today, and you’ve got a few of your own to raise before your workday ends.
PRs are now generated, occasioned with vague method names, and sometimes found holding obscure language functions that require looking up to be sure they actually do what you think they should. Agent-led business logic is scattered, known conventions are crossed in the worst places they could be, and boundaries are so sloppy a junior would blush hearing the number of “WTF”s you unconsciously muttered trying to understand why the code was written that way.
Developers like yourself now dance this context-switching tango between calls, cognitively side-stepping into operational troubleshooting while cyclically churning through the third AI revision of the same PR again - it’s the second one you’ve seen of its kind today – and you’ve still got your own agent’s work to review.
This has been our day-in-day-out for months, as it has been for devs all around the world. You and your ‘agentic’ teammates yearn for the time when you didn’t have to prompt your way out of resolving a bug fix.
You long for the days when you honed your own deductive reasoning while practising muscle-memory refactoring and design to handle prickly bugs or daunting features by hand.
But for most of us with AI usage mandates, we mostly say farewell to writing code by hand at work. It’ll be missed.
Welcome Dante 🤘
Most developers I know don’t like living in review fatigue hell.
I bet you’ve captured every grievance you’ve seen in PRs and meetings and SKILL.md-efied them into agent files galore. You’ve shared with your colleagues and paired up to play with the fanciest models; you’ve collectively tried all the latest AI corralling to improve the quality of what you’re reading so you can get through it faster.
How many times have we heard, “but if you just use this one hack I use,” or “this one super prompt strategy that cures cancer and feeds the world’s hungry babies”, your codebase will be cured! We hear this in every tech video now; however, as the industry matures, the gospel changes not so long after you’ve heard it.
You’ve adjusted to this new world. You know what your days will be like now. The PRs have improved a bit from all this skill-writing effort… sort of.
You can’t out-prompt the quality problem
Agent fallibility is the ever-stale ha-ha we see in the memes in the reels we see after work, and jest over with colleagues before meetings and at dinner conversations with friends.

The internal monologue I swear LLMs secretly have
The internal monologue I swear LLMs secretly have
The non-deterministic and unbridled toddler-like attention span of today’s LLMs costs you and your colleagues daily in mental bandwidth.
They don’t always do as they’re asked and they forget what they’re doing.
You lack agency yet bear responsibility for their work when you are often the last buck before their code hits production. Checking becomes exhausting.
So what now? Is it all doom and gloom? No!
What we’re trying
If I’ve learnt anything from working with agents, it is that we humans have context windows/bandwidths too, and we need to protect them too.
We’ve tried a number of things in my team to reduce the mental cost of reviewing mainly AI-generated code. Extensive skills and new workflows resulted in smaller, bite-sized PRs (for human and AI context windows alike), AI review bots, mob and pair programming, all focused around extreme programming principles.
These strategies have helped catch issues and promote quality and intervene with“AI-weirdness” issues earlier, thus making PRs simpler.
The most complex work is handled by seniors on the team, who can model the code in their heads beforehand and make the agent implement it “to that model”.
However, all these flows still depend on human effort - the review fatigue remains - it’s just shifted to earlier points because non-deterministic quality checks are unreliable, and themselves require human checking.
So, what did we do?
We wanted to see if we could utilise tools that could deterministically offload some of the checking we humans were doing into our CI/CD pipeline (something not practically demonstrated about much online for our ecosystem). The hope was that agents and PR authors could read these reports, before more human reviewers get involved.
Essentially enriching our automated feedback loop.
We shopped around and landed on the following for this experiment:
- Test coverage reports - what % of each line in your codebase has a test that touches it, and in theory, is protecting it from unintentional change. The more good coverage you have, the better your safety net.
- Mutation testing - indicates the quality of said test coverage. How “really good” that 96% score you got on that rickety-looking legacy method you want to change is. Often shows you the holes in said safety net.
- C.R.A.P. scoring - the Change Risk Anti-Patterns score takes your test coverage report for a given method and looks at how Cyclomatically Complex (basically: how nasty to understand) it is. The higher the score, the riskier the code is to change, and therefore the more likely your system is change-averse. High complexity + low coverage = high CRAP score. High coverage + low complexity = low CRAP score. Shows you where the sleeping dragons lie in your system and gives you a beginning talking point with your peers to address them 🐲.
- Static analysis - used to find convention-breaking issues, coding errors, and vulnerabilities based on rulesets you can configure based on your language’s preferred conventions and your team’s collective preferences too! Can be added rule-by-rule to gradually improve legacy or specific areas of the code (e.g. new code gets full ruleset, whilst legacy gets baby step rules).
- Automated refactoring - this one is lovely; can be configured rule-by-rule for legacy or set to max for new code to automatically apply tiny refactors and convention changes as your team sees fit. Some can even aid in language and framework upgrades. Dry run in the pipeline. But beware, some rules can break code if you blindly apply their suggested changes (cover your code first, and agree on the rules).
∴ developers reading PRs can be certain that the agreed-and-safe-to-have-on-as-rules conventions are in place BEFORE they read any code. They don’t need to check hard that the PR has good coverage and that the complexity and ease of changing the codebase haven’t increased, hopefully decreased.
These tools can also aid in strategic “tidy-first” refactoring and characterising (via tests) of change-adverse (risky) areas of the code in discrete PRs BEFORE the new feature work is implemented. This makes the feature PRs more laser-focused, simpler, and less fatiguing to read. i.e. this let’s you more easily break up the work.
How we did it
It took a bored developer’s weekend (hello 👋) and an afternoon of the team’s time to get all this work into our pipeline. After showing the results to our manager, and the prior week’s retro revealed how fried we all felt at times, we got buy-in to add these “wins” to our codebase because their benefits were obvious. The coming weeks will be the true test, but the PRs we’re already seeing are much nicer!
We focused on our backend PHP system, since front-end code is notoriously harder to test. We wanted a bigger bang for our buck in this experiment.
We already had our tests running across multiple test runner jobs, splitting them by runtime to even out their finishing times to approximately the same time. This way, you’re not left disproportionately waiting for one test job while the rest finish minutes earlier.
PHPUnit, and indeed many of the XUnit test frameworks come pre-packed with coverage report options that can be exported in many formats.
We wanted to introduce Mutation testing first as our test suite is our biggest safety net asset when making changes to our system, and with agents now writing tests, we wanted to keep its quality as high as possible.
We chose Infection for our mutation test runner given its infamy and its well-maintained and stable nature.
One thorn, however, was its requirement of the full test suite’s coverage data - so it can determine which tests to run and therefore identify which mutations (synthetic bugs it introduces) are “killed” (or critically not killed) by our tests. We didn’t want to rerun the entire test suite AGAIN to generate this report as that would more than double the pipeline’s runtime (from 14 min to 30+) in one step, and thus degrade our productivity.
Instead, I investigated and got PHPCOV working to create a Merge Coverage step after all our test runners have completed successfully. It hoovers up each partial coverage report and merges them into one complete report file the new Mutation testing step can use.
These complete reports from the Merge Coverage step are then used like so:
Below are snippets of how we did this in CircleCI:
jobs:
backend-tests:
parallelism: 10 # (or whatever number floats your goat)
steps:
- run:
name: Run tests and write coverage
command: |
mkdir -p ~/test-results/coverage
XDEBUG_MODE=off php -d pcov.enabled=1 artisan test \
$(php artisan test --list-test-files | sed -n 's/^ - //p' | circleci tests split --split-by=timings) \
--order-by=random \
--random-order-seed=${CIRCLE_BUILD_NUM} \
--log-junit ~/test-results/junit-${CIRCLE_NODE_INDEX}.xml \
--coverage-php ~/test-results/coverage/job-${CIRCLE_NODE_INDEX}.cov
- store_test_results:
path: ~/test-results
- store_artifacts:
path: ~/test-results
- persist_to_workspace:
root: /home/circleci
paths:
- test-results
The test runner command uses PCOV as the coverage driver and disables Xdebug as the former is faster. The ${CIRCLE_NODE_INDEX} prevents one runner from overwriting another runner’s file.
The --order-by=random --random-order-seed="${CIRCLE_BUILD_NUM}" flags allow us to deterministically (should we need to for debugging) rerun our tests in a predetermined random order when debugging failing erratic tests. The random order increases the likelihood of bringing erratic, test-order-dependent tests to light, so we can kill them off as soon as possible to keep our test suite healthy.
store_artifacts makes reports available for inspection after the job. A developer can open the HTML report or download the XML file without rerunning the build.
persist_to_workspace passes files to later jobs in the same workflow. The CRAP and mutation-testing jobs attach the workspace and use the merged coverage as input.
🗒️ — You’ll need to be running a version of PHPUnit that has my friend’s following contributions to the library in it to get CircleCI’s test splitting working with the above (the sed -n 's/^ - //p' part):
- https://github.com/sebastianbergmann/phpunit/pull/5462
- https://github.com/sebastianbergmann/phpunit/pull/5642
🗒️ - the --log-junit ~/test-results/junit-${CIRCLE_NODE_INDEX}.xml part is there so we can have a central test results file in the next step, should we need it in the future (however, CircleCI already merges the results in the Tests tab if you opt in to using that functionality).
After the parallel test job completes, the dedicated Merge coverage job processes the coverage directory into the different formats needed by different consumers:
- run:
name: Merge coverage
command: |
mkdir -p ~/merged-coverage/coverage-xml ~/merged-coverage/coverage-html
vendor/bin/phpcov merge \
--xml ~/merged-coverage/coverage-xml \
--crap4j ~/merged-coverage/crap4j.xml \
--html ~/merged-coverage/coverage-html \
~/test-results/coverage
- run:
name: 'Merge JUnit test results'
command: npx --yes --package=junit-report-merger@X.Y.Z jrm "$HOME/merged-coverage/junit.xml" "$HOME/test-results/junit-*.xml"
- store_artifacts:
path: ~/merged-coverage/coverage-html
destination: coverage-html
- store_artifacts:
path: ~/merged-coverage/coverage-xml
destination: coverage-xml
- store_artifacts:
path: ~/merged-coverage/crap4j.xml
destination: crap4j.xml
- persist_to_workspace:
root: /home/circleci
paths:
- merged-coverage
🗒️ - If you find your coverage files are too big to merge all at once, you can break the merging into chunks and go from there.
The below is how we run our CRAP analysis; the pipeline fails if the threshold is exceeded.
resource_class: small
steps:
- attach_workspace:
at: ~/
- run:
name: 'CRAP analysis'
command: vendor/bin/crap-check check ~/merged-coverage/crap4j.xml --threshold=14
We also have a linting step that runs tools like Pint, PHPStan, etc., but for brevity and because their setup is trivial, I’ve omitted them from this article.
Optional changes
We decided we wanted to lower the CRAP threshold value of the project to level 14 over the 30 we started with, as the methods it identified - to us - were in need of both simplifying and bolstering with coverage. It found every place our own refactoring instinct had been itching to tackle for years, and this gave us a good cause to do so - the process was quite therapeutic!
This lowering from 30 → 14 was part of one large refactor PR that the team accepted as a one-off, but you could stagger the lowering across multiple PRs if you wanted it to be more bite-sized. I highly recommend learning what refactoring truly is before attempting this, mind
The same happened for Rector. We started with no rules chosen, then enabled rule sets, scrutinised each automated change it made, and decided to either accept it or skip the rule if we found it broke any system behaviour.
🗒️ - We found it cheap to use an agent to backfill missing coverage by first reverting the unstaged changes Rector made, and then rerunning Rector to see if the tests still passed (after reviewing the generated coverage ofc!). This approach allowed us to churn through a lot of the changes very quickly.
🗒️ - Do not be disheartened if your codebase starts with a much higher score than ours. You’ll get there. Don’t let Best get in the way of Better!
Where it got complicated
Multi-threaded mutations
Over the span of the week, we initially experimented with just one thread and job runner for the entire mutation step to remove the infamous false positives and negatives and timeouts, etc etc Infection is known for.
We also settled on having it mutate only on the branch's diff from develop, and only once the branch was in PR (at least in DRAFT mode). This was to save on wasteful runs/re-runs.
After seeing how long it took for a single thread and job to run (9+ minutes added to our pipeline with minimal diff changes) and the causes of the false results, we experimented with enabling multi-threading (-j > 1).
This forced us to get creative and pre-create each thread’s dedicated test database (to prevent collisions) before running infection.
When multi-threading is enabled with Infection, it assigns each thread an incrementing by +1 integer under the TEST_TOKEN env variable that corresponds to its thread number.
Since we can control the -j value (how many threads we’ll have running), we know how many databases we need (one per thread).
I was able to make use of this deterministic TEST_TOKEN integer behaviour to create the databases for each thread before calling Infection, in a bash script similar to the following SQL in the mutation testing database:
# This is pseudocode as to not expose my employer's tech stack information!
THREADS=${1:-10}
GRANT ALL PRIVILEGES ON \`testing_test_%\`.* TO 'XYZ'@'%'; FLUSH PRIVILEGES;
for i in $(seq 1 $THREADS); do
CREATE DATABASE IF NOT EXISTS \`testing_test_$i\`;
done
Laravel’s database.php's 'database' value then just needs the following amendment to it ensure the corresponding database is used by each thread:
'database' => env('DB_DATABASE').(env('TEST_TOKEN') ? '_test_'.env('TEST_TOKEN') : ''),
I found 6 threads on the large.gen2 resource class in CircleCI was a nice sweet spot for us.
Chunking mutations
Mutli-threading improved performance, but we still had to wait for infection to run PER class’s line changes, so I experimented with adding a rudimentary chunking idea to split the infection run across multiple jobs (or “shards”), and have each one process each changed source file like so:
-
Planner -
git diffagainstorigin/developto list changed source files, weight them using CRAP4J complexity/coverage (and optional timings from previous runs), then write a manifest plus one shard file per runner. - Shard runners - each CircleCI node reads its shard file, ignores every other changed file for that run.
-
Collate - merge per-shard
summary.jsonfiles into one MSI for the PR.
Pipeline parameter (top of .circleci/config.yml):
parameters:
mutation-shard-count:
type: integer
default: 10
I was able to get an agent to create a script that another step calls to build a mutation testing “manifest” like so, and split it into these weighted “shard “ files each job could run based on their job number.
{
"shard_count": 3,
"requested_shard_count": 10,
"source_file_count": 7,
"shards": [
[
"app/Order/Services/OrderService.php",
"app/Order/Data/OrderData.php"
],
[
"app/Shared/Service/MailService.php",
"app/Shared/Data/OrderConfirmedData.php"
],
[
"app/Payment/Services/StripeService.php"
]
],
"weights": {
"app/Order/Services/OrderService.php": 48.5,
"app/Order/Data/OrderData.php": 12.0,
"app/Shared/Service/MailService.php": 36.2,
"app/Shared/Data/OrderConfirmedData.php": 8.1,
"app/Payment/Services/StripeService.php": 22.4
}
}
I started with 10 shard jobs. From the example above, there would be 3 shards running while the remaining 7 would end gracefully (20s total run time each) because they had nothing to run with.
This got what was coming out as a 9+ minutes exponential addition to our pipeline runtime down to < 2 minutes, but it would scale better without adding much more time to it.
The workflow ended up looking like the following:
Image showing the various steps of the pipeline
In order to get Infection to mutate just the changed lines in the given shard manifest file, I had to make Infection somehow exclude all the other file changes it would detect on the branch and run our intended shard file lines. Our early testing showed Infection would ignore the specific file it was given if --git-diff-lines was used.
Note: without --git-diff-lines, the entire file would be mutated. What the agent and I settled on was extracting a small script that, for each shard, reads every changed file from the manifest, filters out that shard’s assigned files, and merges the remainder into source.excludes as paths relative to source.directories — not mutators.global-ignore, which expects FQCNs. The generated config looks like
Example of all source files changed in this PR's diff:
app/Order/Services/OrderService.php,
app/Order/Data/OrderData.php
app/Shared/Service/MailService.php
app/Shared/Data/OrderConfirmedData.php
app/Payment/Services/StripeService.php
## Example files we want mutation testing running on for this given shard:
app/Order/Services/OrderService.php,
app/Order/Data/OrderData.php
Below is an example of what each shard could dynamically generate:
{
// Shard 0: mutate OrderService + OrderData; exclude the other changed files
"$schema": "vendor/infection/infection/resources/schema.json",
"logs": {
"html": "test-results"
},
"source": {
"directories": [
"app",
"routes"
],
"excludes": [
"*Test",
"Database",
"Shared/Service/MailService.php",
"Shared/Data/OrderConfirmedData.php",
"Payment/Services/StripeService.php"
]
},
"timeout": 1800,
"mutators": {
"@default": true
}
}
Said script (.circleci/scripts/build-infection-shard-config.php) strips each skip file‘s app/ or routes/ prefix (whatever is in source.directories) so Infection receives exclude paths such as Shared/Service/MailService.php rather than full repo paths.
🫶 -If anyone comes up with a better alternative to this approach to chunking, please say!
Below is a snippet of our mutation testing step with all these performance tweaks, using the shard-specific config file:
backend-mutation-tests-diff:
parameters:
shard-count:
type: integer
default: 6
resource_class: large.gen2
parallelism: << parameters.shard-count >>
steps:
- run:
name: 'Check PR Context'
command: |
if [ -z "${CIRCLE_PULL_REQUEST}" ]; then
echo "This pipeline is not associated with a Pull Request. Skipping mutation testing."
circleci-agent step halt
fi
- attach_workspace:
at: /tmp/workspace
- run:
name: 'Select weighted mutation shard'
command: |
manifest="/tmp/workspace/mutation-plan/manifest.json"
if [ ! -f "$manifest" ]; then
echo "Mutation manifest not found: $manifest"
exit 1
fi
active_shard_count=$(php -r '
$manifest = json_decode(file_get_contents($argv[1]), true, 512, JSON_THROW_ON_ERROR);
echo $manifest["shard_count"] ?? 0;
' "$manifest")
if [ "$CIRCLE_NODE_INDEX" -ge "$active_shard_count" ]; then
echo "No mutation shard assigned to node ${CIRCLE_NODE_INDEX}; skipping."
circleci-agent step halt
exit 0
fi
shard_file="/tmp/workspace/mutation-plan/shard-${CIRCLE_NODE_INDEX}.paths"
if [ ! -f "$shard_file" ]; then
echo "Mutation shard file not found: $shard_file"
exit 1
fi
if [ ! -s "$shard_file" ]; then
echo "No mutation paths assigned to node ${CIRCLE_NODE_INDEX}; skipping."
circleci-agent step halt
exit 0
fi
cp "$shard_file" /tmp/mutation-paths
- run:
name: 'Create mutation runner databases'
command: bash .circleci/scripts/02-create-mutation-testing-databases.sh << parameters.shard-count >>
- run:
name: 'Run PHP Infection for weighted mutation shard'
no_output_timeout: 30m
command: |
set -euo pipefail
report_dir="/home/circleci/build/reports/node-${CIRCLE_NODE_INDEX}"
mkdir -p "$report_dir"
# 1. Early exit if this shard is completely empty
if [ ! -s /tmp/mutation-paths ]; then
echo "No files assigned to this shard. Skipping Infection."
echo "{}" > "$report_dir/mutation-timings.json"
exit 0
fi
started_at=$(date +%s)
# 2. Exclude other shards' files via source.excludes (paths relative to source.directories)
php .circleci/scripts/build-infection-shard-config.php \
--infection-config="infection.json5" \
--manifest="/tmp/workspace/mutation-plan/manifest.json" \
--shard-paths="/tmp/mutation-paths" \
--output="infection-shard.json"
# 3. Run Infection exactly ONCE, applying git-diff-lines and the new config
XDEBUG_MODE=off php -d memory_limit=-1 -d pcov.enabled=1 vendor/bin/infection \
--configuration="infection-shard.json" \
--skip-initial-tests \
--git-diff-base="origin/develop" \
--git-diff-lines \
--coverage=/tmp/workspace/merged-coverage \
--logger-html="$report_dir/mutation-report.html" \
--logger-summary-json="$report_dir/summary.json" \
--logger-text="$report_dir/mutation-report.txt" \
--only-covering-test-cases \
--with-timeouts \
-j6 --min-msi=0 --min-covered-msi=0
duration=$(( $(date +%s) - started_at ))
# 4. Output the timing
timing_file="$(mktemp)"
printf 'node-%s\t%s\n' "$CIRCLE_NODE_INDEX" "$duration" >> "$timing_file"
php -r '
$timings = [];
foreach (file($argv[1], FILE_IGNORE_NEW_LINES | FILE_SKIP_EMPTY_LINES) as $line) {
[$path, $duration] = explode("\t", $line, 2);
$timings[$path] = (float) $duration;
}
file_put_contents(
$argv[2],
json_encode($timings, JSON_PRETTY_PRINT | JSON_THROW_ON_ERROR).PHP_EOL,
);
' "$timing_file" "$report_dir/mutation-timings.json"
- store_artifacts:
path: /home/circleci/build/reports
destination: mutation-reports
- persist_to_workspace:
root: /home/circleci
paths:
- build/reports
🗒️ - We found it useful to bump the “timeout“ value in infection.json5 to 15 minutes as the estimates Infection calculates is often pessimistic and not (usually) representative, thus more mutations are tested vs being skipped!
📚- See Infection’s documentation for further reading on the command’s flags: https://infection.github.io/
Results
In this particular repo, we had a suite of about 6,000 Laravel tests. With mutation testing, selecting only the necessary tests, chunking, and multithreading, the new steps added approximately 2 minutes-ish to the total runtime (depends on number of files changed ofc)!
🗒️- If given permission, I will link to a full, NDA compliant/sanitised, config.yml file you can use!
Mutation test report for just the PR’s diff, not the entire repo
CRAP analysis being run in < 1s with no issues found based on threshold
We enabled 1,074 rules and identified 15 that introduced breaking changes.
Prerequisites
Hopefully your applications have a CI/CD pipeline in place. If you don’t know what that is, or why it is important, God help you.
Just kidding, here’s a helpful video for those who are new to it: https://www.youtube.com/watch?v=i2DrLsnETk4
Most teams today (I hope) have some level of automated test coverage for their code, and it’s in their CI/CD pipeline. TDD was all the rage before the advent of commonplace AI, and it’s still unclear whether TDD makes sense in a world where AI writes most of the code (I was a big fan of it personally) as of the time of writing this.
Regardless, AI can (when guided with skill files written by experienced test writers) backfill coverage of existing code (known as Characterisation tests) and reviewed throughly. The book Working Effectively with Legacy Code by Michael Feathers is brilliant at explaining and guiding on how to achieve this, particularly in a corporate setting.
You don’t need to be in a PHP setting, though. Most languages have their own ports of these tools. Hopefully, the examples above give you the gist of how you could benefit from them in your own pipelines.
Next steps
Scaling: we have a much larger 10+ year old system I’m in the process of porting these CI pipeline changes to. Also, we need to see how scalability factors in. However, so far, the results have been the same: cheap to get running + not much of a hit to pipeline performance other than a few more runners here and there - stay tuned for part two!
Visibility: the generated reports are currently contained within CircleCI, but with the right permissions, our next goal is to create a final job that attaches these reports and their summaries to a comment that can be added to the PR itself for each commit run. This way the reports are immortalised on the PR itself (artefacts in CircleCI are ephemeral), where both humans and agents can ingest them.
The comment body can then be short and useful to both audiences:
## CI quality report
Commit: `<short-sha>`
Run: `<circleci-run-number>`
| Check | Status |
| --- | --- |
| Coverage | pass |
| CRAP | review needed |
| Mutation testing | review needed |
| Tests | pass |
Reports: [attached]
More tools: we’re experimenting with other tools that are too early to mention, but hopefully sharing this blog article will interest other developers, even if it’s just a talking point in your standup somewhere tomorrow. If you’ve got any ideas or feedback (constructive plz), please share it below! 🫶
🪴
This blog post is part of my digital garden: https://jaydenvicarey.substack.com/about. It’s not meant to be perfect, just what I’ve been up to.
Hope it was nice to read!











Top comments (0)