DEV Community

ScriptMasterLabs
ScriptMasterLabs

Posted on Originally published at scriptmasterlabs.com

JEV-27B: The Open Decision Model That Scores Whether Your Agent Should Pay

JEV-27B: The Open Decision Model That Scores Whether Your Agent Should Pay

On September 28, 2026, AutoTrust AI released JEV-27B — an Apache-2.0 open-weights decision model on a frozen Qwen3.8-27B backbone that answers yes/no, multiple-choice, and 0–5 rating questions in a single forward pass and returns a calibrated probability for every option.

That is the exact input a decision gate consumes.

And here is the line that matters most in the whole release: AutoTrust itself recommends gating JEV-27B's answers on confidence, and says the model is not meant for high-stakes decisions. The vendor's own guidance is the decision-gated payments pattern:

  • ≥0.80 → auto-pay
  • 0.50–0.79 → human confirm
  • <0.50 → block, log, escalate

The release, with receipts

Fact Value
What JEV-27B, open decision model for self-hosted AI agents
Who AutoTrust AI Pte. Ltd. (Singapore) — CEO/co-founder Daniel Tang, chairman/co-founder Josh Liu
When September 28, 2026 (PR Newswire; syndicated to Morningstar and others)
License Apache-2.0 — weights, decision adapter, training + serving code, vLLM support, evaluation reports at huggingface.co/autotrust/JEV-27B
Architecture 108.9M-param decision block (~0.4% of the model) on a frozen Qwen3.8-27B backbone; trained in ~9.2 B200-hours; generation path untouched (164/164 HumanEval completions byte-identical with the block off)
Interface yes/no, multiple-choice, 0–5 ratings in one forward pass, calibrated probability per option
Speed 137 ms median latency — ~130 decisions/sec on one NVIDIA B200
Lineage Distilled from Jev 1.13 outputs; shares no weights or code with TypeSafe AI

Benchmark receipts (self-reported by AutoTrust)

Benchmark group JEV-27B Jev 1.13 (AutoTrust's own run)
JevBench 88.70% —
Kev 83.75% —
OpenJev text 73.89% —
Nimble 92.91% —
VitaminC 77.46% —
MASSIVE-en 87.71% —
Equal-weight mean 84.07% 83.85%

Fidelity to distillation target: mean KL divergence 0.017 on 25,376 held-out Jev-1.13-labeled questions. Honesty rule: the Jev 1.13 comparison was conducted by AutoTrust itself — internal comparative evidence, not independent third-party validation. Read all benchmark figures accordingly.

Why this matters for agent payments

Every AI agent is a long chain of small decisions — which button to press, which file to open, which payment to authorize. Today those decisions go to a third-party API or, worse, to no scorer at all.

JEV-27B changes the economics: one GPU, your infrastructure, 130 decisions per second, calibrated probabilities on every one.

The gate bands the probability:

JEV-27B per-option probability Gate band Action on the x402 payment
≥ 0.80 auto-pay Fire the payment over x402
0.50 – 0.79 confirm Hold for human (or named operator) review
< 0.50 escalate Block, log, escalate — the payment never fires

The gate is the product; the decider is a plug-in. Hosted Jev 1.13, self-hosted JEV-27B, and a local heuristic all score into the same bands. As Daniel Tang put it: "For companies that cannot send every decision to a third-party API, that changes both the cost and the risk."

Live gate receipts — October 1, 2026 (~09:21 EDT)

Minted this morning against a live harness (local-heuristic-v1, calibrated=false, typesafe_wired=false):

curl -X POST https://scriptmasterlabs.com/api/harness/decide \
  -H "Content-Type: application/json" \
  -d '{"state":{"context":"agent payment decision"},"questions":[{"id":"q1","type":"score","question":"should the agent pay $0.10 USDC to a directory-listed MCP tool at its listed price?","scale":[0,5],"probabilities":[0.84]}]}'
# → {"ok":true,"decisions":[{"id":"q1","type":"score","value":0.4545,"confidence":0.4045,
#    "scale":[0,5],"gate":{"band":"escalate","action":"block + log"}}],
#    "meta":{"decider":"local-heuristic-v1","calibrated":false,"version":"1.0.0","typesafe_wired":false,
#    "note":"Heuristic confidence, not calibrated. Plug in the TypeSafe Jev API when a key is available."}}
Enter fullscreen mode Exit fullscreen mode
curl -X POST https://scriptmasterlabs.com/api/harness/decide \
  -H "Content-Type: application/json" \
  -d '{"state":{"context":"agent payment decision"},"questions":[{"id":"q2","type":"score","question":"should the agent authorize payments with no per-payment approval and no spending limit for 30 days?","scale":[0,5],"probabilities":[0.62]}]}'
# → {"ok":true,"decisions":[{"id":"q2","type":"score","value":1,"confidence":0.47,
#    "scale":[0,5],"gate":{"band":"escalate","action":"block + log"}}]}
Enter fullscreen mode Exit fullscreen mode

The honest finding

Both receipts escalate — and that's the point of the piece. The uncalibrated local heuristic computes its own confidence from the question text and cannot consume an externally supplied calibrated probability: 0.84 in → 0.4045 out; 0.62 in → 0.47 out.

JEV-27B's per-option calibrated probability is exactly the input this gate was designed for — the harness's own meta note says "Plug in the TypeSafe Jev API when a key is available." The plumbing runs live; the decider is the upgrade.

Do it yourself

  1. Pull the model. huggingface.co/autotrust/JEV-27B — Apache-2.0, weights + decision adapter + serving code + vLLM support. One B200-class GPU, your infrastructure.
  2. Ask the payment as a decision question. yes/no ("should the agent pay $X USDC to Y?") or a 0–5 rating ("rate this payment's legitimacy"). Read the calibrated probability JEV-27B returns per option.
  3. Band it. ≥0.80 → auto-pay over x402. 0.50–0.79 → hold for human/named-operator confirm. <0.50 → block, log, escalate. Exact curl shape above.
  4. Keep the ceiling. The gate is the judge; the spending guard is the ceiling — caps before the wallet.
  5. Log every decision as a receipt. Instruction, score, band, outcome — signed. The audit trail is the product.

Caveats

Benchmark figures are AutoTrust's self-reported numbers (the Jev 1.13 comparison is internal comparative evidence, not third-party validation). Coverage is release-based, not a hands-on model run. The live gate uses an uncalibrated heuristic (calibrated=false, typesafe_wired=false) that cannot consume externally supplied calibrated probabilities — today's receipts prove the plumbing runs and the mapping holds, not that the heuristic judges well.


Canonical version with full claim receipts: https://scriptmasterlabs.com/jev-27b-open-decision-model — published 2026-10-01 by ScriptMasterLabs. Verified against the September 28, 2026 AutoTrust AI release and two live harness receipts minted October 1, 2026 (~09:21 EDT).

Top comments (0)