Rob Hill, Fortitude Omnis. Measured on 27 and 28 September 2026.
I sent a thousand Banking77 support messages through a model on my own RTX 3080 Ti first, and only passed the uncertain ones to Claude Opus 5.5. With a fine-tuned Laya decision model doing the local work, 73.0% of decisions stayed on the GPU, blended accuracy was 93.3% against 94.2% for Claude alone, and the estimated bill fell from £3,898 to £1,054 per million decisions. The caveats belong right here. It's one dataset (Banking77), one card, one night. The £ figures are estimates from published list prices, not invoices. And Claude's answers came from an interactive Claude Code session working through batched answer sheets, not from the API.
The twist is the model that did best. A plain fine-tuned MiniLM classifier kept 94.7% local at the same 94.2% blended accuracy as Claude alone. I'll come back to that. I also measured TypeSafe's hosted Jev on the same items, and it lands between the two on Banking77.
This post is the hands-on version. It shows how to run the server, how to calibrate a model's confidence, and how to get the same report for your own decision. The full write-up is on the Fortitude Omnis R&D page.
What is Tau?
Tau is two things, both Apache-2.0.
The Runtime is a .NET server that answers the /v1/systemone decision contract locally. It runs the open Laya and Von decision models through ONNX Runtime on CUDA, DirectML or a CPU. You send it some state and a set of questions, and it sends back an answer with a probability for each option.
The Workbench is a command-line tool called tau. It asks one question of any /v1/systemone endpoint: can I trust this model's confidence enough to gate on it, and what does gating save? It measures a hosted endpoint over the network the same way as a local one. Every stage writes its output to disk, and the last one writes a single HTML report with the misses left in.
How do I run a self-hosted /v1/systemone server?
The repo's README walks through fetching a model, exporting it to ONNX and fetching the ONNX Runtime natives. After that, starting the Runtime on an NVIDIA card is one line:
dotnet run --project src/Tau.Runtime -c Release -- --urls http://localhost:8088 --Tau:Provider=cuda
Leave out --Tau:Provider=cuda and it runs on the CPU. On Windows, --Tau:Provider=directml runs on any DirectX 12 GPU. It refuses to start if the provider you asked for won't load, which I prefer to a silent fallback to the CPU.
Then ask it something:
curl -s http://localhost:8088/v1/systemone \
-H "Content-Type: application/json" \
-d '{
"model": "jev-latest",
"state": "My card payment failed three times this morning and I need it sorted today.",
"questions": {
"category": {
"type": "choice",
"instructions": "Which support category fits this message?",
"criteria": {
"billing": "Billing or payments issue",
"technical": "App or website technical issue",
"account": "Account access or security issue"
}
},
"urgent": { "type": "noul", "instructions": "Does the customer need an answer today?" }
}
}'
jev-latest is an alias that lets the Runtime pick a model. English text goes to laya-en. You get back a choice with a probability for each option, plus a noul score for the yes/no question. GET /v1/models lists what's loaded.
From C#, the Tau.Client package wraps the same call and maps the answer onto an enum:
using Tau.Client;
using var client = new SystemOneClient(new Uri("http://localhost:8088/"));
var decision = await client.DecideAsync<Category>(
state: "My card payment failed three times this morning and I need it sorted today.",
instructions: "Which support category fits this message?");
Console.WriteLine($"{decision.Value} p={decision.Probabilities[decision.Value]:F2}");
enum Category { Billing, Technical, Account }
How fast is it on a gaming GPU?
On the 3080 Ti with CUDA, laya-en answers one question in 18.59 ms and ten questions about one state in 114.11 ms. On the i9-11900K's CPU, the single question takes 529.21 ms. Those are FP32 medians with no FP16 tricks. For comparison, TypeSafe's hosted Jev had a median of 223 ms, but that's measured over the network from my desk, so it isn't like for like.
How do I calibrate Laya's confidence scores?
This is the bit that matters, and the reason the Workbench exists. A decision model gives you an answer and a probability. You want to keep the answer when the probability is high and escalate when it's low. That only works if 0.9 means right nine times in ten.
Out of the box, it doesn't. On Banking77, laya-en was 37.2% accurate with an expected calibration error (ECE) of 0.502. Its stated confidence sat about 50 points away from its hit rate. A confidence score that looks certain and means nothing.
Calibration fixes the meaning of the number and leaves the model alone. Fitting a small calibrator on a separate 1,000-item split took laya-en's ECE from 0.502 to 0.056, and Von's from 0.185 to 0.021. Accuracy didn't move.
The Workbench does it in stages. You describe the decision in a decision.yaml (the question, the labelled data, the local models, the frontier model and a target error), then run:
tau label examples/banking77/decision.yaml # ingest cached frontier answers
tau measure examples/banking77/decision.yaml # raw accuracy and ECE, calibration and held-out splits
tau calibrate examples/banking77/decision.yaml # fit temperature and isotonic calibrators
# restart the Runtime with --Tau:CalibratorsDirectory=examples/banking77/calibrators, then:
tau measure examples/banking77/decision.yaml --phase calibrated
tau threshold examples/banking77/decision.yaml # pick the threshold for the target error
tau cascade examples/banking77/decision.yaml # simulate local-first and price it
tau report examples/banking77/decision.yaml # write report.json and report.html
tau run does the lot in order and skips stages whose outputs are current. It never calls a paid API unless your spec lists a hosted endpoint under external:, and then only under the budget you set there. If frontier answers are missing, tau label exports batches to be answered and exits with code 2.
One trap I nearly fell into. Laya's contract returns a confidence field for choice questions that's derived from the entropy of the whole distribution, so it isn't a calibrated probability. Calibrate and threshold on the probability of the chosen answer instead. That's why the C# example above prints Probabilities[decision.Value].
How do I decide how much can stay local?
A calibrated model is only useful if it's also right often enough. Raw laya-en wasn't: no threshold got its error below 5%. So I fine-tuned it on the 9,003-item training split, on the same card, in 1,518 seconds.
With a threshold of 0.95, picked on the calibration split and judged on the held-out split, the fine-tuned model kept 73.0% of decisions local and sent the rest to Claude, for an estimated £1,054 per million.
The money is arithmetic on list prices. Each Banking77 prompt is about 1,259 input tokens once you list all 77 intents, so Claude Opus 5.5 comes to £3,898 per million decisions, or £1,949 through the Batch API. Tokens are estimated from characters, not counted. I haven't costed the GPU, which I already owned for less serious reasons.
Why did a classic encoder win?
I put a fine-tuned all-MiniLM-L6-v2 through exactly the same measurement, calibration, threshold and cascade code. It scored 91.5% on the held-out set, with a calibrated ECE of 0.024, against 87.3% for the fine-tuned Laya. In the cascade it kept 94.7% local for an estimated £207 per million. It isn't served through Tau, so that figure leaves out local latency and energy.
On a fixed task with thousands of labelled examples, a small classifier is still the thing to beat. Decision models earn their keep on questions you haven't trained for, or many questions against one state. I'd rather the Workbench told me that than hid it.
The second dataset was worse. On synthetic support tickets, Claude agreed with the labels on 23.8% of items against a 41.0% majority baseline, MiniLM "learned" them to 55.6%, and scored against Claude instead, no local decision model kept more than 0.1%. On urgency, nothing local stands in for the frontier call.
How does a hosted model compare?
The Workbench takes hosted /v1/systemone endpoints as well as local ones. You list them under external: in the spec, with the name of the environment variable that holds the key. The key is read at run time and never written anywhere, and a budget guard stops the run before spend passes the limit you set. I used it on TypeSafe's hosted Jev (the responses say jev-1.13.0), on the same items as everything else. It cost an estimated $0.14 for Banking77 and $0.04 for the tickets, at the published $0.042 per million input tokens.
On Banking77 it's a solid generalist. It scored 79.2% held-out out of the box, ahead of laya-en's 37.2% and Von's 77.1%, with an ECE of 0.093 before calibration and 0.029 after. As the first stage of a cascade it kept 51.3% at 93.6% blended, for an estimated £1,952 per million against £3,898, with its own calls priced in. The fine-tuned Laya beat it on accuracy (87.3%) and on share kept local (73.0%), and MiniLM beat both.
The tickets are where it earns its place. Scored against Claude, it agreed on 52.5% of items out of the box, ahead of every local model. At the 20% target it kept 39.8% of decisions, with the served answers agreeing with Claude on 89.3%, for an estimated £614 per million against £997. Nothing local came close.
Two things to know before you gate on it. Jev returns probabilities rounded to 2 decimal places, and a hosted endpoint can't load your calibrator, so its calibrated figures are the Workbench applying one offline. You'd do the same on your side of the call.
What went wrong?
- I fitted the calibrators on rounded numbers. The Runtime rounds probabilities to 4 decimal places, the Workbench fitted its calibrators on that, and the Runtime applied them to the unrounded values. With 77 options most probabilities round to zero, so it mattered. My earlier figures (laya-en 0.065, Von 0.039) came from that. Refitted on full precision they're the 0.056 and 0.021 above. I'd also claimed the Workbench and the Runtime agree within 2.4e-4, having only checked inputs that didn't saturate. The decisions log has the fix.
- Three FP32 models on one 12 GB card slowed the last one down. The fine-tuned Laya, measured last, had a median of 15,219 ms per request in its raw phase, against 243 ms in a fresh Runtime. VRAM spill is the likely cause, not proven. Accuracy isn't affected, and the published latency comes from the clean phase.
- The classic encoder beat every decision model on Banking77, Tau's and Jev, on accuracy and on calibration.
- The ticket labels are close to noise. Claude agreed with them less often than a constant guess would.
- On urgency, no local model stands in for Claude. No local decision model kept more than 0.1% of decisions.
- Calibration barely helped the fine-tuned Laya: ECE went from 0.071 to 0.060. Fine-tuning had already fitted its temperature.
- laya-en missed the 50% calibration target on the tickets: ECE went from 0.273 to 0.155. The Workbench picks the calibrator with the lower calibration-split log loss, which chose isotonic even though temperature scaling had a calibration-split ECE of 0.016 against 0.216. I set that rule before seeing any held-out result, so I didn't change it.
- Von's tickets calibration moved ECE from 0.045 to 0.042. It was already close.
- My first MiniLM run scored 80.8% because I stopped it after five epochs with the loss still falling. That flatters whatever you compare it with, so I retrained it properly.
- Claude disagreed with the Banking77 labels on 5.8% of items. Some are Claude's mistakes and some are the dataset's.
How do I reproduce it?
The repo holds both worked examples with their reports, calibrators, cached frontier answers and per-item results. With the models fetched and the CUDA natives in place, one command reruns Banking77 end to end and rewrites its report:
./scripts/examples.ps1 -Example banking77
It checks the data and model packages first and prints the exact command for anything missing. Then read the Banking77 report and the tickets report. They were harder on me than I've been here.
The repo is at https://github.com/Fortitude-Group/tau.


Top comments (0)