DEV Community

Cover image for Mentioned Is Not Binding: Which Solicitation Clauses Actually Bind the Contract
Raihan
Raihan

Posted on

Mentioned Is Not Binding: Which Solicitation Clauses Actually Bind the Contract

Rules plus a 0.46 MB classifier decide, for every FAR/DFARS clause number in a solicitation, whether it binds — evaluated in seven pre-registered rounds, now live in production.


fedproc-ledger results

Scope note, first paragraph as promised: the gold labels behind every number below come from a coding agent and two LLM judges, not human experts. The task is binding-set classification on real solicitations, and every interval is a cluster bootstrap over documents.

The problem with regex clause lists

VETR finds FAR clauses with regular expressions — 16 patterns, longest-prefix-first. That finds clause numbers well. But a number appearing in a solicitation does not mean the clause binds the contract. In real solicitations, FAR 52.212-5 contains a long checklist and the contracting officer checks only the clauses that apply; clauses get cited inside other clauses' text, in instructions, in tables of contents; and "52.212-5" in an amendment can mean the whole thing was deleted.

I measured the damage without any annotator: of 423,328 status-quo ledger entries across 6,472 documents, 58,294 (13.8%) are clauses whose every mention is a checklist item with an empty box. In checklist-heavy documents it's 20–31%. That is a lower bound — items whose box state was lost in extraction aren't counted.

The system: rules first, model second

The model never generates a clause number — numbers come from regex candidates validated against an eCFR registry (2,509 Title-48 sections with version history), and the model only classifies each mention's role: incorporated by reference, full text, selected, not selected, narrative, internal reference, and so on. Checkbox states are decided by rule (form widgets, glyphs, typed [X]/__ marks), never by the model.

The classifier itself is deliberately small: one-vs-rest logistic regression over structural features plus character TF-IDF, 0.46 MB of numpy weights, CPU-only, no LLM at inference. A number's score is the max over its mentions at q ≥ 0.5. Explicitly deleted clauses ("CLAUSE 52.247-59 IS DELETED") veto the number through the aggregation.

Seven pre-registered rounds, four consecutive wins

Each round fixed its frame, model, and numeric hypotheses in writing before any label, on documents no earlier round had seen — round 7 on a temporal hold-out of solicitations posted after the last acquisition day. Majority gold of the agent + two judges, binding set:

  • Round 4 (rules v1.2, 26 docs): F1 0.921 vs 0.780 (+0.142 [+0.085, +0.210])
  • Round 5 (24 docs): 0.926 vs 0.829 (+0.097 [+0.040, +0.158])
  • Round 6 (rules v1.3, 24 docs): 0.914 vs 0.807 (+0.107 [+0.057, +0.161])
  • Round 7 (rules v1.4, 27 docs): 0.890 vs 0.806 (+0.084 [+0.028, +0.143])

Specificity 70–79% vs 8–19% at equal recall. The first two rounds did not show this (one failed outright — the judges accepted unchecked boxes as binding), and those failures are in the paper, not hidden.

What bounds the score

Pooled annotator agreement is 75–82% (kappa 0.47–0.56). A score of 0.98 against these labels is not attainable without fitting annotator quirks — the ceiling is in the labels, not the model. The applicable-set reading (narrative requirements count) still favors the regexes. And the review-queue calibration I pre-registered never hit its target in four tries. All reported, all in the repo.

Live in production, open to all

The ledger now powers the clause analysis behind VETR compliance matrices (regexes kept as fallback), with a case study on vetrproposal.com. Everything is public: model (pickle-free numpy, Apache-2.0), dataset, code, paper draft (submitted to arXiv). Not legal advice; the model must not remove clauses silently.

Top comments (0)