If you are evaluating how well a repository supports AI-assisted development, it is easy to mix two different jobs: measuring one repository quickly and adding a research observation to a maintained corpus. Those jobs need different levels of evidence.
harness-maturity-analysis gives you both paths. Its corpus workflow pins repositories to exact commits and records scanner versions. Its ad hoc command lets you inspect any local path or public repository without writing to the corpus.
This tutorial focuses on the second path. You will run a deterministic local score, read the dimension breakdown, identify the highest-value gaps, and understand what the result does not prove.
TL;DR
Clone the project, install its development dependencies, then run the documented score script with a repository URL or local path:
git clone https://github.com/paladini/harness-maturity-analysis
cd harness-maturity-analysis
npm install
npm run score -- https://github.com/owner/repo
The command prints a maturity level, a point total, dimension percentages, detected tooling, and several unmet checks. It does not add the target to corpus/manifest.json or create a generated report under corpus/reports/.
Prerequisites
You need Git, Node.js 18 or newer, and npm. The project declares the Node engine requirement in package.json. A network connection is needed when you pass a remote repository URL because the runner must obtain the source and the pinned scanner package.
The project is MIT licensed. Read the repository license before incorporating the scripts into another workflow.
The current public README documents harness-score@1.8.1 as the scanner pin used by the corpus. That version is part of the research evidence, not a promise that every future run will produce the same number if the target repository changes.
Run an ad hoc score
From the project directory, pass either a GitHub URL or a local directory:
npm run score -- https://github.com/owner/repo
npm run score -- C:\path\to\local\repository
The first command asks the project to obtain the remote source. The second analyzes files already on disk. In both cases, the ad hoc runner reads the target and sends it to the scanner; it does not install dependencies or execute code from the repository being inspected. That boundary matters when the target is unfamiliar.
For a first experiment, score the analysis project itself:
npm run score -- .
The score npm script delegates to corpus/score-adhoc.mjs. The script uses the study's configured scanner version, prints the score and dimension breakdown, and lists unmet checks that may offer the clearest improvement opportunities.
Understand the output
A result has several separate signals. The maturity level is a compact interpretation of the score. The point total explains how much of the available model the repository satisfies. Dimension percentages show whether the result is driven by context, skills, guardrails, feedback, CI, or hygiene.
The detected-tooling line is descriptive. If the scanner identifies editor or agent configuration, that tells you which signals were found in the inspected files. It does not certify that a team uses the tools effectively.
The unmet checks are a prioritization aid, not a backlog generated by an oracle. A missing type-checking check, for example, may be a reasonable next experiment for a JavaScript project, but it is not automatically a defect for every repository. Read the check definition and the relevant files before changing anything.
The project's methodology makes the same distinction for the larger study: the score describes a repository's harness, not the competence of the organization that owns it. A low score can be appropriate for a library that was never designed as an agent-first workspace.
Reproduce the observation
An ad hoc score is useful for exploration, but a research claim needs stronger identity. Record at least these values alongside the output:
- the repository URL;
- the commit SHA or tag inspected;
- the scanner version;
- the command and options used;
- the complete output;
- the date and environment.
Then rerun the same command at the same commit. If the score changes, investigate the scanner version, checkout contents, platform-specific files, and command arguments before interpreting the difference.
The corpus workflow formalizes this process. corpus/manifest.json stores exact repository commits and the scanner version. corpus/run.mjs clones pinned source into a cache, scans it, and writes generated reports. The generated artifacts are then rebuilt from those reports instead of being edited by hand.
That separation is the key design choice: use npm run score -- ... when you are asking, "How does this repository look right now?" Use the corpus workflow when you are ready to make a durable, reviewable data point.
Verify the local installation
Before trusting a local checkout, run the project's own checks:
npm test
npm run lint
The test suite covers parsing, manifest shape, selection logic, report generation, history handling, site rendering, and clone behavior. Linting checks the JavaScript and data files with Biome. These checks validate the analysis tool itself; they do not validate the quality of the repository you pass to npm run score.
After a score, check the working tree:
git status --short
For an ad hoc run, the expected result is no new corpus entry or generated report. If you intended to create study evidence and see unexpected files, stop and inspect the command and repository state before continuing.
Failure modes and security boundaries
The target is not reproducible
A branch name points to moving content. Prefer a commit SHA or release tag when comparing results over time. A score from main is a current snapshot, not a permanent fact.
The score looks surprisingly high or low
Read the file-level evidence and check whether the project type matches the model's assumptions. The study explicitly warns that a repository-local harness score should not be treated as an organizational ranking.
The remote checkout cannot be obtained
Check the URL, visibility, network access, and Git installation. Do not work around the failure by copying generated reports from another checkout. A missing checkout is missing evidence.
You are tempted to run target code
Do not. The documented runner is designed to inspect files without executing code from the scanned repository. Keep that boundary intact, especially for untrusted projects, and treat any workflow that installs target dependencies as a different security model.
The corpus changes unexpectedly
Do not hand-edit corpus/reports/, results/, or the Pages site. Those files are generated. Re-run the documented generator after changing an input, then review the diff.
FAQ
Does an ad hoc score change the public corpus?
No. It is intended for a one-off reading. Corpus additions require a separate pinned and reviewable workflow.
Can I compare two projects directly?
You can compare their outputs as an exploratory exercise, but keep repository type, commit identity, scanner version, and checkout conditions visible. The number alone is not a universal quality ranking.
Does L4 mean the repository is production-ready?
No. It means the inspected files satisfy the model's harness checks at that snapshot. It says nothing by itself about correctness, reliability, maintainability, licensing of dependencies, or operational safety.
Is this an AI judge?
The scanner is deterministic and filesystem-oriented. It does not replace human review, and the repository's research protocol treats blind human assessment as a separate question.
Takeaway
Use the ad hoc command to ask a focused question without polluting research history. Pin the target and scanner when the result needs to be reproduced. Most importantly, read the evidence behind the number before turning a maturity score into a technical decision.
What repository would you score first, and which scanner signal would you want to validate manually before trusting the result?
AI assistance disclosure: AI was used to help organize and edit this tutorial. Repository behavior, commands, versions, and limitations were checked against the project's public source and local verification output.
Top comments (0)