DEV Community

Cover image for Linting Prose in CI
Steef-Jan Wiggers
Steef-Jan Wiggers

Posted on

Linting Prose in CI

A commenter on Watching the Contract Fail in Production put the conclusion of this whole series into one sentence: once agents consume tool descriptions as part of their decision-making, changing a sentence is no longer just documentation, it can be an interface change, so important wording changes should be treated like code changes. This post makes the build system agree. The companion repo now has a contract linter and a GitHub Action, and the first thing the linter did was fail its own repo.

Four rules, each one earned

The linter is about a hundred lines of dependency-free Python (measure/lint_contract.py), and every rule in it was paid for by a measurement earlier in the series:

Provenance. Every identifier parameter must say where its value comes from ("as returned by search_restaurants"). This is the sentence a JSON schema cannot express, and the difference between an agent that chains tools and one that invents ids; the confusion telemetry showed exactly that failure as the menu_unknown_restaurant sentinel.

When-to-call. Every tool description must contain a usage cue, not only a what-it-does sentence. Selection within a shortlist works because descriptions disambiguate between siblings, and when-to-call is where that disambiguation lives.

Token budget. The serialized tools array must stay under a budget, counted with the same cl100k tokenizer the discovery post used. The budget here is 350 tokens; every description edit now has a visible price.

Drift. The C# attribute descriptions and the APIM Bicep copies of the same tools must be character-for-character identical, because the contract post demonstrated that two copies of a contract fork within a week even with one author.

A fifth rule, the alias sentence that took paraphrase recall from 5 to 100 percent, is a warning rather than an error: proven valuable in the retrieval experiments, advisory at three tools.

The linter parses the contract straight out of the McpToolTrigger and McpToolProperty attributes, the authoritative copy, the same source the deployed description hash is computed from. One pleasant cross-check fell out of that: the parser's token count of the extracted contract was 277, exactly the number the discovery post measured from the live server's tools/list response. Two extraction paths, one contract, same count.

The first run: red, deservedly

I pushed the linter with the honest commit message "current contract fails its own lint," and the Action delivered:

GitHub Actions run failing the contract lint with exit code 1 on the commit that introduced the linter

Nothing here was staged. The drift errors are the exact fork the contract post documented by eye weeks ago, still sitting in the repo, now caught mechanically. The when-to-call errors are real too: get_menu's description said what the tool does and nothing about when to reach for it, and I wrote that description believing descriptions were the contract. The linter's first catch was its own author, which is the strongest argument I can offer for having one.

The fix: 25 tokens, priced

The fix commit added when-to-call clauses to two tools ("Call it after search_restaurants and before placing an order", "Use it only after get_menu has confirmed the exact item name") and aligned the three Bicep descriptions with the C#. The build went green, and the budget rule priced the edit: the contract grew from 277 to 302 tokens, 25 tokens for two sentences of disambiguation, well under budget and exactly the kind of spend the measurement posts said earns its keep.

Local run of the contract linter reporting zero errors and three alias warnings at 302 tokens

The commit pair is the point of the whole exercise. One push failed CI over sentences; the next push fixed sentences and passed. Wording changes went through the same gate code goes through: a diff, a red build, a fix, a green build. And because the deployed description hash is computed from these same attributes, this fix also changed the descHash, so the confusion telemetry from the previous post can compare agent behavior before and after the new wording. The lint gate and the telemetry loop are two ends of one pipeline: CI checks what can be checked statically, the dashboard watches what only production can reveal.

GitHub Actions run passing the contract lint on the commit that fixed the wording

What the gate can and cannot hold

Worth being precise about the boundary. These rules check the mechanical half of description quality: presence of provenance, presence of a usage cue, total cost, copy consistency. None of them can tell you whether the sentence is good, whether "Use it only after get_menu has confirmed the exact item name" actually improves selection, that is what the eval harness measures and what the telemetry observes. The lint is the cheap gate that runs in six seconds on every push; the eval is the expensive gate you run when the cheap one says the shape is right. Most teams have neither; the order to add them in is lint first, because it catches the failures that need no model at all.

The takeaway

The contract is part of the system behavior, so it gets the treatment behavior gets: versioned by hash, observed in telemetry, and now gated in CI. A one-word description edit can fail your build, and that is not overengineering, that is the build system catching up with where your interface actually lives. The linter, the workflow, and both commits, the red one and the green one, are in the companion repo; point measure/lint_contract.py at your own attributes and find out what your contract has been getting away with.

Top comments (0)