An AI coding agent can edit a repository faster than you can review the diff. Speed is not the same thing as a check. If the repository does not expose a command the agent can run, the agent still has to guess whether the edit works.
That surrounding system is the harness: the instructions, skills, hooks, tests, CI, and guardrails wrapped around the model. The model is rented. The harness is what you own. Harness Score is a small open-source scanner that measures that harness from the files in the repository. It does not call a model, and it does not phone home while it scans.
Version 1.8.1 is a useful excuse to talk about the project, because it fixes a specific kind of lie a score can tell. A repository already had tests. The scanner said it had no test runner.
The model is rented. The harness is yours.
Two repositories can use the same coding agent and get different results. One has an AGENTS.md that names the real commands, a test the agent can run after an edit, and a CI job that repeats that check. The other has a README and a hope that the next session will rediscover the conventions.
Harness Score exists to make that difference visible. Point it at a repository that uses Cursor, Claude Code, Devin, Windsurf, Cline, Continue, or another AI coding tool. You get a maturity level from L0 to L4, a score out of 108 points across six dimensions, and a ranked list of missing evidence. Each failed check says what was missing and where to look. The same commit produces the same score on a laptop and in CI, because the checks are filesystem facts: a file exists, a script is declared, a hook parses. They are not a judgment of whether the product is good.
The guide is at paladini.github.io/harness-score. The source is paladini/harness-score.
What the scanner actually checks
The six dimensions are the control system, not a style guide:
- Context and guides: can a new session find the commands and the boundaries?
- Skills and commands: are repeatable workflows packaged, or only described in prose?
- Hooks and guardrails: is anything enforced at edit or shell time, or is every rule optional?
- Sensors: is there a test runner, a linter, a type checker, a formatter, and at least one real test file?
- CI: does a pipeline exist, and does it run those sensors?
- Hygiene: are secrets, env files, and license signals handled in the tree the agent will read?
Sensors are the part an agent uses in the middle of a task. A test suite is not only a safety net for humans. It is how the agent checks its own edit before it writes a confident summary. The guide's practical bar is a fast suite and one obvious command. npm test is the familiar shape. It is not the only shape.
The version is the occasion, not the product
Harness Score 1.8.1 landed because of a public report, issue #88, from @lglucas. The repository had no package.json. It did have test files that load Node's built-in runner, and CI ran them with node --test. Node documents that command as needing no config file and no dependency (Running tests from the command line).
SNS-01, the "test runner configured" check, still failed. It knew how to see a package.json test script, Vitest, Jest, pytest, go test, and cargo test. It did not know how to see require('node:test'). The failure text said no runner was detected. That sent people looking for a missing framework. The runner was already there.
v1.8.1 passes that check when a test-like JavaScript or TypeScript file actually loads node:test, including import and require, and a subpath such as node:test/reporters. The evidence names the command an agent can run:
node:test built-in (node --test), e.g. scripts/test/a.test.js
A file that is only named a.test.js still fails. Jest, Vitest, and Node's runner share that filename, and a different check already scores "a test file exists." If test files are present and nothing wires a runner, the failure now says so:
Found 1 test file(s), e.g. a.test.js, but no test runner or standard entry point detected.
A CI line that only contains node --test does not, by itself, pass SNS-01. "CI runs the tests" is a separate check. The local signal is the load of node:test. A package.json script of "test": "node --test" already passed before this release, and it still does.
That is the product in one incident. The score is useful when it names the feedback the agent can run. It is harmful when it invents a missing tool. The rest of the scanner is the same idea applied to guides, skills, hooks, linters, and CI.
Try it on a repository you already have
You need Node.js 18 or newer. No API key. The scan does not install your dependencies and does not modify the tree.
Confirm the package version, then scan the current directory:
npx --yes harness-score@1.8.1 --version
npx --yes harness-score@1.8.1 .
--version should print 1.8.1. If it does not, the registry has not caught up with the GitHub release yet. Wait and pin the version again rather than trusting an unpinned npx harness-score, which may still resolve to 1.8.0.
Read the failed checks before you read the headline score. A missing test command, a missing linter, and a missing CI job are different repairs. The remediation line on each check is the next edit, and the guide links from the report go to the matching section.
If you want the Node case specifically, a minimal file is enough. This is the shape covered by the project's own tests, not a suggestion to delete your existing runner:
const { test } = require('node:test');
test('ok', () => {});
Put that in scripts/test/a.test.js in a repository with no package.json test script. On 1.8.1, SNS-01 should pass and mention node --test. A sibling file that never loads node:test should still fail, and the message should name the file.
You can also gate CI on a level once the repository deserves it:
- uses: paladini/harness-score@v1
The Action tag v1 moves only after a published release. Pin a scanner version when you need a stable comparison. Scores from 1.8.0 and 1.8.1 are not interchangeable for SNS-01: a repository that only has node:test can gain those points without any other change.
What a high score does not mean
A high score means the expected files and commands were found. It does not mean the tests cover the behavior you care about, that the rules in AGENTS.md are true, or that a human reviewed the diff. The scanner cannot tell a good test from a test that always passes.
It also does not require you to document the command in AGENTS.md. The remediation text asks for one obvious entry point and a note in the agent guide. The check itself looks for the runner. Writing the command down is still the thing an arriving agent reads first. node --test in the evidence is the scanner telling you the command. Your guide should say it too.
Do not add empty skills, unused hooks, or a package.json you do not need just to move the number. For a dependency-free Node repo, node --test is a real entry point. Adding npm only to satisfy an older scanner was the wrong fix. That is why this release exists.
Takeaway
Harness Score is a measuring tape for the system around an AI coding agent. v1.8.1 is one correction on that tape: Node's built-in test runner counts, and a repository that already has tests is no longer told that the runner is missing. The useful habit is the same either way. Give the agent one command that checks the edit, keep that command fast, and let CI run it again.
If an agent opened your repository today, what is the single command you would want it to run before it claims the task is done?
AI assistance disclosure: I used AI assistance to organize and edit this post. The scanner behavior, the v1.8.1 release notes, issue #88, and the Node.js test-runner documentation were checked against those primary sources.
Top comments (0)