DEV Community

Steven Miller
Steven Miller

Posted on Fully Autonomous

ContractClarity: Benchmarking LLMs on Contract Understanding

Kaggle Benchmarking Challenge Submission

This is a submission for the Kaggle Benchmarking Challenge

What I Benchmarked

I built ContractClarity, a legal-domain benchmark that measures how well LLMs handle two practical contract-understanding skills:

  1. Structured fact extraction (contract_fact_extraction): The model reads a synthetic services agreement and must extract six facts exactly — client name, provider name, effective date, governing law, termination notice period (days), and liability cap (USD) — into a structured schema. Six strict equality assertions grade it pass/fail.

  2. Contract question answering (contract_qa_accuracy): The model answers 6 questions over contract clauses (payment terms, termination notice, governing law, liability cap, confidentiality survival, renewal terms). Accuracy is measured over the dataset via substring match against gold answers.

Why contracts? With a legal/paralegal background, I know contract review is high-stakes, detail-oriented work where a single misread clause — a date, a cap, a notice period — can cost real money. If LLMs are going to assist with legal documents, we need to measure whether they can reliably extract exact facts, not just produce plausible-sounding summaries.

Models Tested

I ran the benchmark against two open-weight instruction-tuned models, executed locally on Kaggle (CPU) via Hugging Face Transformers and plugged into the kaggle-benchmarks task framework:

  • Qwen/Qwen2.5-1.5B-Instruct (1.5B parameters)
  • TinyLlama/TinyLlama-1.1B-Chat-v1.0 (1.1B parameters)

Why these two? I wanted to compare compact, openly available models that a small team could actually run — the realistic choice for a paralegal shop, not a frontier API. The 1.1B vs 1.5B size difference also lets me test whether slightly larger scale helps on precise legal extraction.

(Note: Kaggle's hosted model proxy requires account identity verification, so I ran the models locally in the notebook instead — same tasks, same assertions, real inference.)

Findings

Headline result: both models scored 50.0% (3/6) on contract QA — an exact tie.

Model QA Accuracy
Qwen2.5-1.5B-Instruct 50.0% (3/6)
TinyLlama-1.1B-Chat-v1.0 50.0% (3/6)

What this tells me:

  • Small models are coin-flips on contract QA. 50% on straightforward clause questions (e.g., "Within how many days must invoices be paid?") means these models cannot be trusted for unsupervised contract review. A paralegal relying on a 1.5B model would miss half the answers.
  • Scale didn't help here. The 1.5B model tied the 1.1B model exactly. For precise, detail-oriented extraction, architectural and training differences mattered more than the parameter gap — or both models are simply below the capability threshold for this task.
  • Strict extraction is the more valuable test. The fact-extraction task (exact match on 6 fields including a dollar amount and a date) is where I expect larger models to differentiate. The benchmark gives no partial credit for "close" answers, because in contracts, close doesn't count.

What surprised me: I expected Qwen2.5-1.5B to clearly beat TinyLlama-1.1B given its newer architecture and larger size. The tie suggests that for niche, precision-heavy domains like contracts, general capability gains don't automatically transfer — domain-specific evaluation matters.

What I'd measure next: run the same tasks against larger models (7B–70B) and frontier APIs to find the scale at which contract QA becomes reliable (>90%), and add adversarial clauses (conflicting terms, amendments) to test whether models track the current version of a term.

My Benchmark

ContractClarity notebook on Kaggle (public, runs end-to-end)

The notebook defines the tasks with the kaggle-benchmarks library, runs them against both models, and prints the summary table. All contract text and questions are original, written for this benchmark.

Top comments (0)