Result first: a 60-line zero-dependency Python script gets typed judgments out of jev-1.13 on AiHubMix — a routing label, a 0–2 severity score with its own legend, and a yes/no probability — in one round trip, with no generated text to parse.
The engineering problem this solves: classification via a chat model means prompting for JSON, then defending against the model not returning JSON. jev-1.13 removes that layer. You declare named questions, you get answers keyed by those names. The trade-off is that the endpoint is non-standard and your criteria design becomes the accuracy bottleneck.
Prerequisites
- Python 3 (standard library only — no
requestsneeded) - An AiHubMix API key, exported as
AIHUBMIX_API_KEY
The vendor sample uses requests. If it is not installed, urllib.request from the stdlib is a drop-in for a single POST and keeps the script dependency-free.
The endpoint is not OpenAI-compatible
POST https://aihubmix.com/v1/systemone
Authorization: Bearer $AIHUBMIX_API_KEY
Content-Type: application/json
This is a model-specific route, not /v1/chat/completions. No OpenAI-compatible SDK will work against it. Plan on raw HTTP.
Request contract
Two fields carry the work:
-
state— the text to be judged -
questions— an object where each key is a question name you choose
payload = {
"model": "jev-1.13",
"state": "Hi, I have been trying to connect my Stripe account for 3 days "
"and it keeps failing. I am losing sales. Please help ASAP.",
"questions": {
"department": {
"type": "choice",
"instructions": "Which team should handle this",
"criteria": {
"billing": "Payment or subscription issues",
"technical": "Bugs or integration problems",
"sales": "Pricing or account questions",
},
},
"frustration": {
"type": "score",
"instructions": "How frustrated the customer appears",
"criteria": [
"Calm, just stating facts",
"Frustrated but civil",
"Very angry, strong language",
],
},
"is_urgent": {
"type": "noul",
"instructions": "The message conveys urgency or time-sensitivity",
},
},
}
Note the asymmetry in criteria: choice takes a dict (label → definition), score takes an ordered list (index → definition). noul takes no criteria at all.
Response shapes
Each answer stores its value under a key matching its own type, so answer[answer["type"]] is the general accessor.
| type | value key | extra fields |
|---|---|---|
noul |
noul (0..1 probability of yes) |
none — no confidence
|
choice |
choice (your label) |
probabilities, confidence
|
score |
score (list index) |
legend, probabilities, confidence
|
Verified score envelope:
{
"type": "score",
"score": 1,
"legend": {"0": "Calm, just stating facts", "1": "Frustrated but civil", "2": "Very angry, strong language"},
"probabilities": {"0": 0, "1": 1, "2": 0},
"confidence": 1
}
legend echoes your own criteria back, indexed — so the integer is self-describing and you do not need a parallel lookup table client-side.
The top level also carries usage (input_tokens / output_tokens), id, provider (TypeSafe), and a resolved backend id: typesafe/jev-1.13-20260917.
The script
#!/usr/bin/env python3
import json, os, pathlib, sys, urllib.error, urllib.request
URL = "https://aihubmix.com/v1/systemone"
def get_key():
k = os.environ.get("AIHUBMIX_API_KEY")
if k:
return k.strip()
f = pathlib.Path(__file__).with_name(".aihubmix_key")
if f.exists():
return f.read_text().strip()
sys.exit("Missing API key: set AIHUBMIX_API_KEY or write .aihubmix_key")
def build_payload(state):
return {"model": "jev-1.13", "state": state, "questions": {...}} # as above
def ask(payload):
req = urllib.request.Request(
URL,
data=json.dumps(payload).encode(),
headers={"Authorization": "Bearer " + get_key(),
"Content-Type": "application/json"},
method="POST",
)
try:
with urllib.request.urlopen(req, timeout=60) as r:
return json.loads(r.read())
except urllib.error.HTTPError as e:
sys.exit(f"HTTP {e.code}: {e.read().decode(errors='replace')[:500]}")
if __name__ == "__main__":
data = ask(build_payload("..."))
for name, ans in data.get("answers", {}).items():
kind = ans["type"]
print(f"{name:12} {kind:7} {ans[kind]!r:28} confidence={ans.get('confidence')}")
print("usage:", data.get("usage"))
Keep the key out of source. Read it from the environment, or from a local file at mode 600 that your VCS ignores.
Observed output
department choice 'billing' confidence=0.51
frustration score 1 confidence=1
is_urgent noul 1
usage: {'input_tokens': 424, 'output_tokens': 73}
Token cost for the 3-question payload: 424 in / 73 out. Dropping to 2 questions: 355 in / 38 out. Question count moves both sides of the ledger, so prune questions you will not act on.
Failure modes
1. choice confidence drifts across identical requests. The same byte-identical payload returned confidence=0.51 on one call and 0.38 on the next. The selected label was stable (billing both times), but the stated certainty was not. Do not treat a single confidence reading as a stable property of an input.
2. Overlapping criteria cap your accuracy. billing ("Payment or subscription issues") and technical ("Bugs or integration problems") both describe a failing Stripe integration. The low confidence is the model correctly reporting an ambiguity in the schema, not a model defect. Write mutually exclusive definitions before blaming the model.
3. noul has no confidence. ans.get("confidence") yields None for every noul answer. If you branch on confidence, special-case noul — its probability is its confidence signal.
4. Module-level requests fire on import. Building and sending the request at module scope means import jev_min re-runs the whole call and spends tokens. Put the payload behind build_payload() and the send behind ask(), and guard the entry point with if __name__ == "__main__": — as above. The version I first wrote did not, and importing it to reuse the payload cost an extra billed call.
5. No SDK fallback. Because the route is non-standard, there is no OpenAI-compatible client to switch to when something breaks. Log the raw response body on non-2xx; the HTTPError branch above truncates at 500 characters, which has been enough.
Reproducibility limits
This is a three-call smoke test against one input. No latency measurement, no batching, no second test case, no repeat count high enough to characterize the confidence drift as a rate. Treat the drift as a reproduced observation that warrants your own repeat testing, not as a quantified error bar. Pricing and context-window figures on the model page are vendor-published — the page lists 64K in its header and 32K in the provider table, a discrepancy this test did not resolve.
Try it against your own data
The endpoint, the full question-type reference, and the published per-token pricing are all on the AiHubMix model page. jev-latest tracks the newest release — read the spec there, then pin the explicit version in code.
https://aihubmix.com/model/jev-latest
To reproduce this in about two minutes: grab an API key, paste the script above, and swap state for one of your own records. If the classification step in your pipeline is currently a chat completion plus a JSON parser plus a retry, this is the shape worth benchmarking against it.
Top comments (0)