DEV Community

Yong Yu
Yong Yu

Posted on Originally published at yongboyu.hashnode.dev

Liquid AI's d1 Decision Models Went Open: Triage Support Tickets on a CPU With the 600M One (and Where It Fools You)

Attributed compile + one small real CPU run

Primary sources (Liquid AI, 2026-10-07): Open d1: Edge decision models for text, vision, and audio (Liquid AI blog), Multimodal open d1 decision models for the edge (Hugging Face blog), and the model cards for d1-3B and d1-omni-600M

Background: Introducing d1: The most capable decision model, now with vision (Liquid AI, 2026-10-05)

All benchmark and latency figures for d1-3B are Liquid AI's. The numbers I report for d1-omni-600M come from one run on a shared 8-vCPU Linux box with 12 tickets I wrote myself. That is a smoke test, not an evaluation.


This week Liquid AI put two of its d1 "decision models" on Hugging Face as open weights: d1-3B (text + images) and an experimental d1-omni-600M (text + images, or text + audio). It landed the same week OpenAI put its Decisions API into public beta, and the conversation on X quickly turned to running these decisions locally instead of paying per call.

The idea is simple and, for backend people, more interesting than another chat model. A decision model does not write text. You hand it a state (a string, JSON, an image) and a set of named questions, and it returns typed, calibrated answers in one forward pass, with zero output tokens. Three question types cover most pipeline glue:

type you give you get back
noul a yes/no question P(yes)
choice named options with descriptions choice, confidence, probabilities
score 2 to 10 ordered levels expected level, confidence, probabilities

That is exactly the shape of the ticket routers, intent classifiers, and moderation checks that a lot of us currently build with an LLM call, a prompt that says "answer only with JSON", and a parser that hopes for the best.

What Liquid AI is claiming

From the blog and model cards, in short:

  • d1-3B (3.12B params, built on LFM2.5-VL-3B) scores 48.57 on Decision Index v0.2.1, which Liquid says is the best under 10B and on par with a 35B-A3B decision model. Liquid reports 8 ms per decision on an RTX 4090, 30 ms on an Apple M5 Pro, and 50 ms on a Jetson Orin Nano.
  • d1-omni-600M is built on a 350M bidirectional encoder plus vision and audio encoders. Liquid calls it an early research release, scores it 15.95 on the same index, and deliberately publishes no speed numbers for it.
  • Both have day-one transformers support via trust_remote_code, and Liquid mentions llama.cpp support for the family.

The 3B model is the one Liquid is pushing. I ran the 600M one, for a boring reason: d1-3B wants about 12 GB in float32 on CPU, and the box I used had about 6 GB free. If you have a GPU or an M-series Mac, start with the 3B.

Run it yourself (CPU, about 10 minutes)

The 600M model card asks for transformers>=5.15 (the 3B card says >=5.14). so pin it explicitly:

python3 -m venv .venv && . .venv/bin/activate
pip install --index-url https://download.pytorch.org/whl/cpu torch torchvision
pip install "transformers>=5.15" pillow soundfile
Enter fullscreen mode Exit fullscreen mode

Then one call answers three questions over the same ticket:

import torch
from transformers import AutoModel

model = AutoModel.from_pretrained(
    "LiquidAI/d1-omni-600M", trust_remote_code=True, dtype=torch.float32
)

questions = {
    "refund": {"type": "noul", "instructions": "Is the customer asking for a refund?"},
    "team": {
        "type": "choice",
        "instructions": "Which team should handle this?",
        "criteria": {
            "billing": "Charges, refunds, invoices",
            "technical": "App or site faults",
            "fraud": "Suspected unauthorised use",
        },
    },
    "urgency": {
        "type": "score",
        "instructions": "How urgent is this?",
        "criteria": ["Can wait", "Today", "Blocking the customer now"],
    },
}

out = model.system_one(
    "I was charged twice for my October plan. Please refund the duplicate.", questions
)
print(out["answers"])
Enter fullscreen mode Exit fullscreen mode

On my run, that ticket came back as billing with confidence 0.994, P(refund) = 0.999, and an urgency score of 1.69 on the 0 to 2 scale, using 157 input tokens and 0 output tokens. Two notes from the card: use float32 on CPU (the model was trained in it), and avoid bfloat16 on GPU, which Liquid says flipped the top answer on a small share of rows. trust_remote_code also downloads Python files from the repo, so pin a revision before you ship this anywhere.

What happened on 12 tickets

I wrote 12 short support tickets and their expected team and refund labels before running anything: four billing, four technical, four fraud-ish, including a phishing email question and a hijacked account.

  • Team: 12 of 12 matched my labels. Refund (P >= 0.5): 12 of 12.
  • Latency, three questions in one pass, warm: median 208 ms, range 160 to 471 ms. One question alone: median 114 ms.
  • Batch: 48 single-question requests through system_one_batch took 3.63 s, about 13 tickets per second.

Those are CPU numbers on a shared Xeon box, so treat them as "fast enough for a background queue", not as a latency benchmark. Twelve easy tickets also say nothing about your data; the point is that the API shape works and the probabilities are usable.

The more useful part was where confidence dropped. The phishing question ("I got an email asking me to confirm my card number on a weird link. Is that you?") still went to fraud, but at 0.477. The hijacked-account ticket went to fraud at 0.633. Those are exactly the tickets I'd want a human to see, and the confidence field flags them without any prompt engineering.

Where it fools you: forced choices

A choice question always picks one of your options. I sent two inputs that are not tickets at all:

  • "What's your office address in Toronto?" went to billing at 0.496. The low confidence gives it away.
  • "asdf asdf test test" went to technical at 0.907. High confidence, and wrong.

So a confidence threshold alone does not catch junk. The obvious fix is a catch-all option, so I added "other": "None of the above, or not a real support request" and a separate noul gate, "Is this a genuine customer support request?". That fixed the junk: the gibberish moved to other (0.663) with P(request) = 0.004, and the address question moved to other (0.896).

It also moved two real tickets. With other available, the phishing question went to other (0.603) and "Do you have a nonprofit discount?" went to other (0.699) instead of billing. The clear tickets (double charge, crashing app) stayed put at 0.90 or higher. Adding a catch-all moves the decision boundary for every borderline input, not just the junk.

How I'd actually wire it

  1. Gate first. Ask a noul "is this a real request?" question and drop or queue anything with very low P(yes).
  2. Route with explicit options, no catch-all, and send anything under a confidence threshold (I'd start around 0.7 and tune it on real tickets) to a human or a bigger LLM.
  3. Write option descriptions like policy. "Suspected unauthorised use" did not obviously cover "someone asking if a phishing email is real". If phishing reports should go to fraud, say so in the description.
  4. Label 100 real tickets before trusting any of this. My 12 are a smoke test. Measure agreement, and check which tickets land between 0.4 and 0.7.
  5. Pin the revision of a trust_remote_code model, and try d1-3B if you have the memory. It's the one Liquid benchmarks.

Why this matters

For years the pattern has been to call an LLM, beg for JSON, and parse it. A small model that returns probabilities in one pass, runs on a CPU, and costs nothing per call is a real option for the boring 80% of classification work. It doesn't replace judgment on the hard 20%, but it tells you which tickets those are.

The scripts and raw output from this run (triage.py, other_option.py) need only the pip install above and no API key.


YongBo Yu is an AI engineer in Toronto building LLM workflows and agent systems. More at yongbo-yu.vercel.app and GitHub.

Top comments (0)