Contract review is where "the model said so" is least acceptable. A reviewer wants the passage, not a paraphrase of it; "this contract has no termination-for-convenience clause" is a real answer, different from "I don't know"; and someone has to decide which answers a lawyer must look at. This post shows how solvi handles all three, with a real LLM answering on a small contract, and what the same method measured on 1,025 questions from real contracts. I wrote solvi; part 3 of 5 (overview).
Why this way
-
The answer type says what a valid answer is.
Maybe[Span[str]]: an exact piece of the text with its offsets, orUnknown— "not stated". solvi checks a span against the text; an answer that is not literally there is rejected, never passed on. - The model proposes, it does not certify. Its confidence is one input to a trust score, not the decision about who looks.
- The decision about who looks carries a promise calibrated on labelled questions of your own: below the threshold, a person gets the case.
A real LLM on a small contract
The contract in the script is a short master services agreement with seven sections. Two traps: Section 4 is called "Exclusivity" but is a non-compete, and Section 7 lets a party terminate for breach — which is not termination for convenience.
CONTRACT = """MASTER SERVICES AGREEMENT
1. Services. Northwind Analytics Ltd ("Provider") shall deliver the data services described in Exhibit A to
Bluefin Retail GmbH ("Client").
…
7. Breach. Either party may terminate this Agreement if the other party commits a material breach and fails to cure
it within thirty (30) days of written notice.
"""
KINDS = {
"governing_law": "Which passage says which state's or country's law governs the contract?",
"non_compete": "Which passage restricts a party from competing or serving competitors?",
"termination_for_convenience": "Which passage lets a party terminate without cause (for convenience), "
"regardless of breach?",
}
The model is any OpenAI-compatible server. The script reads the endpoint and key from the environment, so LLM_URL=https://openrouter.ai/api/v1 LLM_KEY=sk-... python s3_contracts.py runs it against OpenRouter, and an Ollama or vLLM URL runs it locally:
url = os.environ.get("LLM_URL", "http://127.0.0.1:8765/launch10/s3/v1")
model = llm(url, "openai/gpt-oss-120b", api_key=os.environ.get("LLM_KEY", "local"),
extra_body={"reasoning": {"effort": "low"}})
print("Part 1 — an LLM proposes, solvi keeps only passages that are in the contract")
for name, question in KINDS.items():
part = model.decision(name, question, "contract", Maybe[Span[str]])
d = part(contract=CONTRACT)
v = d.value
if v is Unknown:
print(f" {name:28s} not stated (confidence {float(d.conf):.2f})")
elif v is None or d.escalate:
print(f" {name:28s} escalated: {d.escalate}")
else:
same = CONTRACT[v.start:v.end] == str(v.value)
print(f" {name:28s} contract[{v.start}:{v.end}] = {str(v.value)[:70]!r}… literal: {same}, "
f"confidence {float(d.conf):.2f}")
Two details in that call. extra_body asks for reasoning, and with reasoning on solvi sends no forced JSON schema: on the task stand, forcing the schema on a reasoning model skipped its thinking and cost accuracy (a hallucination judge on RAGTruth eval: F1 0.733 with the schema enforced, 0.766 with the reply contract in the prompt). And an invalid or cut-off reply never becomes a value: the decision escalates.
What came back
The run below used gpt-oss-120b, three requests, about a hundredth of a cent:
Part 1 — an LLM proposes, solvi keeps only passages that are in the contract
governing_law contract[775:930] = '6. Governing Law. This Agreement is governed by the laws of the Republ'… literal: True, confidence 1.00
non_compete contract[450:637] = '4. Exclusivity. During the Term and for twelve (12) months after it, P'… literal: True, confidence 0.99
termination_for_convenience not stated (confidence 0.95)
The model saw through both traps: the non-compete is quoted from "Exclusivity", and the breach clause was not taken for termination for convenience — "not stated", with its own confidence. Each passage is contract[start:end], so a reviewer's UI can highlight it and an auditor can check it without trusting anyone.
What happens to a quote that is not quite in the text
Models type no-break spaces, curly quotes and odd dashes where the document has plain ones. Before 1.0, solvi rejected such an honest quote; since 1.0 a quote is matched on a normalized view (Unicode NFKC, dashes, quotes, whitespace), and what is stored is always the contract's own text at offsets into the original. A paraphrase is still rejected. Part 2 of the script shows it offline, with the same lookup solvi uses (find_quote):
for written in ["governed by the laws of the Republic of Ireland", # as in the text
"governed by the laws of the Republic\u00a0of\u00a0Ireland", # no-break spaces, as models type them
"governed by Irish law"]: # a paraphrase
q = find_quote(written, {"contract": CONTRACT})
print(f" {written!r:55s} -> " + (f"found at [{q.start}:{q.end}], stored as {q.value!r}" if q else "not in the text: rejected"))
Part 2 — what happens to a quote a model writes (offline)
'governed by the laws of the Republic of Ireland' -> found at [811:858], stored as 'governed by the laws of the Republic of Ireland'
'governed by the laws of the Republic\xa0of\xa0Ireland' -> found at [811:858], stored as 'governed by the laws of the Republic of Ireland'
'governed by Irish law' -> not in the text: rejected
A record whose quote needed the normalized view says so in the trace, and replay re-checks it. Catalog(quotes="literal") keeps the old, strict rule if you need it.
Who looks: a trust score under a promise
A seven-section contract fits in any context window; real ones do not, and real reviewers cannot read every answer. On solvi's task stand the same method ran on CUAD: 25 commercial contracts × 41 kinds of clause, 1,025 questions on the held-out split (measured with solvi 0.8.0 and gpt-oss-120b through OpenRouter). The solution (benchmarks/tasks/cuad/solution.py) adds three things to the code above:
-
Long documents:
long="retrieve"splits the contract into sections, BM25 picks the few that bear on the question (with the clause's own vocabulary inretrieve_query), and the model reads only those; quotes still point into the whole document. - A check: a second yes/no question about each quoted passage with a little context around it; a passage it refuses becomes "absent".
-
A trust fact computed in the catalog: how often this kind of answer was right for this kind of clause on dev × the LLM's confidence × the check's p(yes). Then
System.guarantee(signal="trust", max_risk=0.03): below the calibrated threshold the question goes to a person.
What it measured, against the same LLM asked directly:
| baseline (plain LLM call) | solvi solution | |
|---|---|---|
| accuracy | 0.882 | 0.899 |
| quotes literally in the contract | 260 of 398 | 263 of 263 |
| false claims where the clause is absent (of 721) | 60 | 27 |
| clauses found (of 304) | 243 | 227 |
| answered without a person; wrong among them | 100%; 11.8% | 67.6%; 4.0% |
The trust score is what made the last row possible: on dev it separated right from wrong answers with AUROC 0.83, where the LLM's own confidence managed 0.63 — at the same risk, the confidence alone would have let only 14% of dev be answered alone.
What it does not do
- It finds fewer clauses. 227 against 243 of 304 present: the check and the stricter quote rule cost recall. If missing a clause is worse than a false claim for you, that trade-off goes the wrong way.
- It does not make the model read better. Accuracy moved from 0.882 to 0.899; most of the gain is in which answers can be trusted, not in more right answers.
-
The promise needs labelled questions of your own contracts, and it covers contracts like those. A new contract type, a new language: calibrate again. Retrieval matches words, so put the document's own words (and its language) into
retrieve_query. - A quote in the text is not proof the quote answers the question. The check question helps; it is still a model.
Where next
- Part 1: routing tickets with a promise
- Part 2: product matching with a rule across answers
- Part 4: guarding a support agent's tool calls
- Part 5: an agent that learns the world it acts in
- Full code, outputs and a run on three real CUAD contracts: solvi-ai/recipes: 03-contract-clauses
- Docs: models in solvi, best practices for documents, the task stand
If you review contracts with an LLM today: what does a reviewer on your side need to see next to an answer before they stop re-reading the clause themselves? And which would you rather trade away — recall (some clauses missed) or the share answered without a lawyer?

Top comments (0)