DEV Community

FX-LgLL
FX-LgLL

Posted on

We made our decision model's API free (and the weights are open)

Most "AI decisions" in production aren't open-ended generation. They're classification in disguise:

  • which team should handle this ticket?
  • is this e-mail phishing?
  • what's the total on this invoice?
  • which sentence in the contract supports that answer?

Teams usually send each of these to a large LLM: 1–2 seconds and a bill per call, a JSON answer to parse, and a "confidence" that means nothing.

We built THX-01 for exactly these decisions. Today we're making its hosted API free.

Try it now (no key, no sign-up)

curl https://api.hal-x.ai/v1/systemone \
  -H "Content-Type: application/json" \
  -d '{
    "state": "Hi, I was charged twice for one order and nobody answers. Refund me!",
    "questions": {
      "team":   {"type": "choice", "question": "Which team handles this?",
                 "criteria": {"billing": "billing", "tech": "technical", "sales": "sales"}},
      "refund": {"type": "noul", "question": "Does the customer want money back?"},
      "urgency":{"type": "score", "question": "How urgent?", "criteria": ["low", "medium", "high"]}
    }
  }'
Enter fullscreen mode Exit fullscreen mode

Every answer comes back with a probability for each option, computed in a single forward pass of about 10 ms.

The free tier allows 200 decisions per minute per IP. One question is one decision.

What it can answer

type returns
choice one option, with a probability for every option
noul P(yes)
score a level on an ordinal scale
number a value stated in the document, or null if it isn't there
excerpt a verbatim span with character offsets
"cite": true the passages that support any answer

How it compares

On our 2,843-ticket benchmark (Azerbaijani, Russian, English, Turkish; clean, corrupted, messy and independently written sets):

model avg accuracy latency
THX-01 (322M, open) 98.4 ~10 ms
TypeSafe Jev 1.13 97.4 331 ms
Kev-4B 92.8 830 ms

The model's confidence is calibrated, so a single threshold lets it handle about 75% of tickets automatically with zero errors on all four test sets.

Run it yourself

The weights are Apache 2.0:

pip install thx01
Enter fullscreen mode Exit fullscreen mode
import thx01
agent = thx01.load("doofz/THX-01")
agent.decide("My card was charged twice", {
    "team": {"type": "choice", "question": "Which team?",
             "criteria": {"billing": "billing", "tech": "technical"}}})
Enter fullscreen mode Exit fullscreen mode

It runs on a laptop CPU in a fraction of a second.

Top comments (0)