Give the agent its domain tools and a citation check before you spend a training run on a custom model. On CoCounsel Bench, Joel Hron said the prior CoCounsel answered about 25% of questions in one shot, and the rebuilt agent reached north of 70% almost overnight once it had the tools and basic instructions. Record every source that run opened, and reject any citation that is not on the list.
Joel Hron, Chief Technology Officer at Thomson Reuters, covered this on Chain of Thought in August 2026. He leads product engineering across legal, tax, audit, and compliance, including Westlaw, Practical Law, and CoCounsel. CoCounsel is the application. Thomson is the model.
What raised one-shot accuracy before the new model?
The tools, with short instructions for how to use them. Earlier CoCounsel versions were not coding-agent architectures, Hron said. The new CoCounsel decomposes Westlaw and Practical Law into tools the agent can call. CoCounsel Bench is their legal-task set: about 3,000 queries, with more added almost every day.
CoCounsel is built around the Claude Agent SDK and also uses models from OpenAI and Google. Thomson is the later step he described: train the model to use those tools, past skills and system prompts. The first CoCounsel job for Thomson is bulk document review, such as due diligence on large document sets. On tabular analysis, the high-volume review tied to that launch, he said Thomson was outperforming the prior model versions they were using by about five percentage points in accuracy. Frontier models stay in the stack for frontier intelligence work, he said. Thomson does not have to replace Claude, OpenAI, or Gemini for that path to proceed.
But we could answer about 25% of those questions in one shot on the prior version of CoCounsel. And when we built this new version of CoCounsel, almost overnight, we got north of 70% just by giving the agent the tools it needed to sort of do the work and some basic instructions on how to use those tools.
Joel Hron, Chief Technology Officer at Thomson Reuters, on Chain of Thought ep 69
How do you block a citation the agent never opened?
Keep a register of every case inspected and use it to ground the claims.
Hron described a citation ledger: a register of every case inspected during research, used to ground the claims the agent makes. Thomson Reuters has a patent pending on CoCounsel's approach to citation ledgers and citation verification. Westlaw's Litigation Document Analyzer applies a related check to a finished brief. It splits the document into claims, looks for case law that supports each claim, and says so when it cannot find a source.
Fabricated case names are a serious failure, Hron said. They still happen, and he sees them at a pretty low rate, less often than the news suggests. The failures he called harder to catch are misreadings of law the agent did open. Humans remain the primary check. He said there is no automated verification loop for these models like there is with code.
A minimal sketch to adapt, not a drop-in library.
class CitationLedger:
def __init__(self):
self.opened = set()
def record_open(self, source_id):
self.opened.add(source_id)
def cites_never_opened(self, answer_cites):
return [cite for cite in answer_cites if cite not in self.opened]
def accept_or_reject(answer, ledger):
blocked = ledger.cites_never_opened(answer.cites)
if blocked:
return {"status": "reject", "blocked": blocked}
return {"status": "ok", "cites": list(answer.cites)}
The bad hallucinations that are really difficult to catch are the things that are misinterpretations of the law. ... Or it interpreted this statement from the judge as fact versus opinion. These kinds of things might change the way that the agent operates.
Joel Hron, Chief Technology Officer at Thomson Reuters, on Chain of Thought ep 69
Where does continued pretraining sit relative to the tools?
Continuous pre-training is an earlier training stage, on a selected slice, before agentic reinforcement learning teaches tool use. Fresh case law stays a retrieval problem.
The pre-training text is drawn from Westlaw, Practical Law, Checkpoint (their tax research content), and the Reuters News Archive, after they select what matters most for the jobs. He said less than 10% of that corpus had been applied, and he expects that share to increase. They publish thousands, if not hundreds of thousands, of cases a day, so weights alone are not a practical way to stay current. The model is trained to use Westlaw, and the tax and news tools, for that.
And to this point, we started small there, but less than 10% of that corpus has been applied for continuous pre-training. So that's an element of the training process that we think will increase over time.
Joel Hron, Chief Technology Officer at Thomson Reuters, on Chain of Thought ep 69
FAQ
Did training Thomson produce the move from about 25% to north of 70%?
No. Hron tied that change to the rebuilt CoCounsel agent, the tools, and basic instructions for those tools. Training Thomson to use tools more directly is a later step. The first CoCounsel use he named for Thomson is bulk document review.
Does a citation ledger catch a misreading of a real case?
It grounds claims against cases the agent inspected. Fabricated names still occur at a low rate. Hron said the harder errors are misinterpretations of real law, such as reading a holding too broadly or treating a judge's statement as fact. Those still need a person.
How much of the corpus went into the first continued-pretraining pass?
Less than 10%, on text selected from Westlaw, Practical Law, Checkpoint, and the Reuters News Archive. Hron said that share should grow. He put their publishing pace at thousands, if not hundreds of thousands, of cases a day, so keep a retrieval tool beside the weights.
Takeaway
- Score one-shot accuracy on a fixed task set before and after the domain tools exist, as its own change, separate from a training run.
- Ship basic instructions with each tool. That is the pairing Hron tied to the one-shot move on CoCounsel Bench.
- Log every source the agent opens. Reject an answer that cites a source absent from the log.
- Have a reviewer check real cases for an overly broad holding and for a statement treated as fact rather than opinion.
- Run continued pretraining on a selected slice after the tools work, and keep retrieval for material that updates faster than a training cycle.
The full conversation, with the transcript, is on Chain of Thought.
Subscribe to the Chain of Thought newsletter for new episodes and write-ups like this one.
Drafted with AI assistance from the episode transcripts.
Top comments (0)