DEV Community

FX-LgLL
FX-LgLL

Posted on

We open-sourced THX-01: a 322M decision model that matches Sonnet 5.5 on ticket classification at ~10 ms (Apache 2.0, pip install thx01)

Today we're releasing THX-01, a model for the boring decisions that
agents and backends make thousands of times a day: route this ticket, is this spam, which team, what's the invoice total,
which sentence supports that.

What it is

  • 322M parameters (mmBERT-base encoder + 15M decision head), Apache 2.0
  • Non-autoregressive: you send a state plus typed questions, and it returns a probability for every option in one forward pass. No generation, no JSON parsing, no retries.
  • Question types: choice, noul (yes/no), score, plus number (pulls a value stated in the document, or null), excerpt (verbatim span with offsets) and "cite": true on any question
  • ~10 ms per request on one GPU, ~3,000 decisions/s batched, and it runs on CPU
  • Speaks the /v1/systemone wire format, so it is a drop-in for existing TypeSafe-style service layers

Numbers (our support-ticket benchmark, 2,843 tickets in az/ru/en/tr, 15 categories)

model avg acc latency
THX-01 (322M) 98.4 ~10 ms
Claude Sonnet 5.5 98.5 1.5 s
Wahoo 1.5 97.8 145 ms
TypeSafe Jev 1.13 97.4 331 ms
Kev-4B 92.8 830 ms

Because it's trained with a strictly proper scoring-rule reward, the probabilities are calibrated (ECE 0.003). With a single
confidence threshold it handled ~75% of tickets fully automatically with zero errors on all four test sets.

pip install thx01

import thx01
agent = thx01.load("doofz/THX-01")
agent.decide("My card was charged twice for one order", {
    "team": {"type": "choice", "question": "Which team handles this?",
             "criteria": {"billing": "billing", "tech": "technical", "sales": "sales"}}})
Enter fullscreen mode Exit fullscreen mode

Model + docs: https://huggingface.co/doofz/THX-01

Top comments (0)