DEV Community

Cover image for Contract clauses with quotes you can check: an LLM proposes, solvi keeps only what is in the text
Maxim Kuznetsov
Maxim Kuznetsov

Posted on AI-assisted

Contract clauses with quotes you can check: an LLM proposes, solvi keeps only what is in the text

Contract review is where "the model said so" is least acceptable. A reviewer wants the passage, not a paraphrase of it; "this contract has no termination-for-convenience clause" is a real answer, different from "I don't know"; and someone has to decide which answers a lawyer must look at. This post shows how solvi handles all three, with a real LLM answering on a small contract, and what the same method measured on 1,025 questions from real contracts. I wrote solvi; part 3 of 5 (overview).

Why this way

  • The answer type says what a valid answer is. Maybe[Span[str]]: an exact piece of the text with its offsets, or Unknown — "not stated". solvi checks a span against the text; an answer that is not literally there is rejected, never passed on.
  • The model proposes, it does not certify. Its confidence is one input to a trust score, not the decision about who looks.
  • The decision about who looks carries a promise calibrated on labelled questions of your own: below the threshold, a person gets the case.

A real LLM on a small contract

The contract in the script is a short master services agreement with seven sections. Two traps: Section 4 is called "Exclusivity" but is a non-compete, and Section 7 lets a party terminate for breach — which is not termination for convenience.

CONTRACT = """MASTER SERVICES AGREEMENT

1. Services. Northwind Analytics Ltd ("Provider") shall deliver the data services described in Exhibit A to
Bluefin Retail GmbH ("Client").
…
7. Breach. Either party may terminate this Agreement if the other party commits a material breach and fails to cure
it within thirty (30) days of written notice.
"""

KINDS = {
    "governing_law": "Which passage says which state's or country's law governs the contract?",
    "non_compete": "Which passage restricts a party from competing or serving competitors?",
    "termination_for_convenience": "Which passage lets a party terminate without cause (for convenience), "
                                   "regardless of breach?",
}
Enter fullscreen mode Exit fullscreen mode

The model is any OpenAI-compatible server. The script reads the endpoint and key from the environment, so LLM_URL=https://openrouter.ai/api/v1 LLM_KEY=sk-... python s3_contracts.py runs it against OpenRouter, and an Ollama or vLLM URL runs it locally:

url = os.environ.get("LLM_URL", "http://127.0.0.1:8765/launch10/s3/v1")
model = llm(url, "openai/gpt-oss-120b", api_key=os.environ.get("LLM_KEY", "local"),
            extra_body={"reasoning": {"effort": "low"}})

print("Part 1 — an LLM proposes, solvi keeps only passages that are in the contract")
for name, question in KINDS.items():
    part = model.decision(name, question, "contract", Maybe[Span[str]])
    d = part(contract=CONTRACT)
    v = d.value
    if v is Unknown:
        print(f"  {name:28s} not stated (confidence {float(d.conf):.2f})")
    elif v is None or d.escalate:
        print(f"  {name:28s} escalated: {d.escalate}")
    else:
        same = CONTRACT[v.start:v.end] == str(v.value)
        print(f"  {name:28s} contract[{v.start}:{v.end}] = {str(v.value)[:70]!r}… literal: {same}, "
              f"confidence {float(d.conf):.2f}")
Enter fullscreen mode Exit fullscreen mode

Two details in that call. extra_body asks for reasoning, and with reasoning on solvi sends no forced JSON schema: on the task stand, forcing the schema on a reasoning model skipped its thinking and cost accuracy (a hallucination judge on RAGTruth eval: F1 0.733 with the schema enforced, 0.766 with the reply contract in the prompt). And an invalid or cut-off reply never becomes a value: the decision escalates.

What came back

The run below used gpt-oss-120b, three requests, about a hundredth of a cent:

Part 1 — an LLM proposes, solvi keeps only passages that are in the contract
  governing_law                contract[775:930] = '6. Governing Law. This Agreement is governed by the laws of the Republ'… literal: True, confidence 1.00
  non_compete                  contract[450:637] = '4. Exclusivity. During the Term and for twelve (12) months after it, P'… literal: True, confidence 0.99
  termination_for_convenience  not stated (confidence 0.95)
Enter fullscreen mode Exit fullscreen mode

The model saw through both traps: the non-compete is quoted from "Exclusivity", and the breach clause was not taken for termination for convenience — "not stated", with its own confidence. Each passage is contract[start:end], so a reviewer's UI can highlight it and an auditor can check it without trusting anyone.

What happens to a quote that is not quite in the text

Models type no-break spaces, curly quotes and odd dashes where the document has plain ones. Before 1.0, solvi rejected such an honest quote; since 1.0 a quote is matched on a normalized view (Unicode NFKC, dashes, quotes, whitespace), and what is stored is always the contract's own text at offsets into the original. A paraphrase is still rejected. Part 2 of the script shows it offline, with the same lookup solvi uses (find_quote):

for written in ["governed by the laws of the Republic of Ireland",          # as in the text
                "governed by the laws of the Republic\u00a0of\u00a0Ireland",  # no-break spaces, as models type them
                "governed by Irish law"]:                                    # a paraphrase
    q = find_quote(written, {"contract": CONTRACT})
    print(f"  {written!r:55s} -> " + (f"found at [{q.start}:{q.end}], stored as {q.value!r}" if q else "not in the text: rejected"))
Enter fullscreen mode Exit fullscreen mode
Part 2 — what happens to a quote a model writes (offline)
  'governed by the laws of the Republic of Ireland'       -> found at [811:858], stored as 'governed by the laws of the Republic of Ireland'
  'governed by the laws of the Republic\xa0of\xa0Ireland' -> found at [811:858], stored as 'governed by the laws of the Republic of Ireland'
  'governed by Irish law'                                 -> not in the text: rejected
Enter fullscreen mode Exit fullscreen mode

A record whose quote needed the normalized view says so in the trace, and replay re-checks it. Catalog(quotes="literal") keeps the old, strict rule if you need it.

Who looks: a trust score under a promise

A seven-section contract fits in any context window; real ones do not, and real reviewers cannot read every answer. On solvi's task stand the same method ran on CUAD: 25 commercial contracts × 41 kinds of clause, 1,025 questions on the held-out split (measured with solvi 0.8.0 and gpt-oss-120b through OpenRouter). The solution (benchmarks/tasks/cuad/solution.py) adds three things to the code above:

  1. Long documents: long="retrieve" splits the contract into sections, BM25 picks the few that bear on the question (with the clause's own vocabulary in retrieve_query), and the model reads only those; quotes still point into the whole document.
  2. A check: a second yes/no question about each quoted passage with a little context around it; a passage it refuses becomes "absent".
  3. A trust fact computed in the catalog: how often this kind of answer was right for this kind of clause on dev × the LLM's confidence × the check's p(yes). Then System.guarantee(signal="trust", max_risk=0.03): below the calibrated threshold the question goes to a person.

What it measured, against the same LLM asked directly:

baseline (plain LLM call) solvi solution
accuracy 0.882 0.899
quotes literally in the contract 260 of 398 263 of 263
false claims where the clause is absent (of 721) 60 27
clauses found (of 304) 243 227
answered without a person; wrong among them 100%; 11.8% 67.6%; 4.0%

The trust score is what made the last row possible: on dev it separated right from wrong answers with AUROC 0.83, where the LLM's own confidence managed 0.63 — at the same risk, the confidence alone would have let only 14% of dev be answered alone.

Four bar charts for CUAD, plain LLM versus solvi: wrong among answered alone 11.8% (answering 100% alone) versus 4.0% (answering 67.6% alone); quotes literally in the contract 65% (260 of 398) versus 100% (263 of 263); false claims where the clause is absent 60 versus 27 of 721; clauses found 243 versus 227 of 304 — solvi finds fewer.

What it does not do

  • It finds fewer clauses. 227 against 243 of 304 present: the check and the stricter quote rule cost recall. If missing a clause is worse than a false claim for you, that trade-off goes the wrong way.
  • It does not make the model read better. Accuracy moved from 0.882 to 0.899; most of the gain is in which answers can be trusted, not in more right answers.
  • The promise needs labelled questions of your own contracts, and it covers contracts like those. A new contract type, a new language: calibrate again. Retrieval matches words, so put the document's own words (and its language) into retrieve_query.
  • A quote in the text is not proof the quote answers the question. The check question helps; it is still a model.

Where next

If you review contracts with an LLM today: what does a reviewer on your side need to see next to an answer before they stop re-reading the clause themselves? And which would you rather trade away — recall (some clauses missed) or the share answered without a lawyer?

Top comments (0)