DEV Community

Cover image for Wiring a Third-Party Decision Model Into Claude Code, and What I Wouldn't Let It Do by Default
Kiell Tampubolon
Kiell Tampubolon

Posted on

Wiring a Third-Party Decision Model Into Claude Code, and What I Wouldn't Let It Do by Default

TL;DR

  • Downloaded a vendor's "paste this into your agent" setup guide for Jev, a fast/cheap decision model. Had Claude Code build the integration live instead of just reading about it.
  • Secured the real API key into gitignored storage before a single test call.
  • Caught the docs-fetching tool hallucinating the API endpoint (doubled path, 404), fixed it with a stricter re-fetch.
  • Built a reusable client, two on-demand skills, and a message-size router hook, pipe-tested through every branch, shipped off by default.
  • Closed the loop by having Jev itself pick this article's title. 0.99 confidence, under a second, a fraction of a cent.

Manual triage versus Jev-assisted triage: a person opening 100 emails one by one, versus Jev sorting the same pile in seconds with three hot leads flagged and Claude writing only those replies

I downloaded a setup guide for something called Jev: a model that promises "10x the power at 100x less cost" when you pair it with Claude. I didn't read it first. I opened Claude Code, told it to read the PDF, and asked for a plan I could run in that same session. What came back wasn't a plan to review later. It was a working integration, a real paid API call, a caught hallucination, and a question about how much of a vendor's own instructions I should let an agent follow without looking.

What Jev actually is

Jev (typesafe/jev-1.13, from TypeSafe AI) isn't a language model in the usual sense. It doesn't write anything. You send it a piece of text plus one or more typed questions, and it answers in one of three shapes: choice (pick one option), score (place it on a scale), or noul (the probability a yes/no statement is true), each with a confidence number attached.

A fictional sales email, three questions, one real call through OpenRouter (a service that proxies API access to many AI models behind one key): sales, a 2.97-3 lead score at 0.97-1.0 confidence, 0.75-0.78 probability it needed a same-day reply. Under a second. About two cents per thousand calls.

The guide's own benchmark backs this up at scale, run against 100 made-up emails for a fictional agency, recorded 21 September 2026:

Model Time Cost Lead score exactly right Hot leads found
Jev 3.6s $0.0032 89/100 15/15
Claude Haiku 4.5 (thinking off) 11.1s $0.061 84/100 15/15
Claude Fable 5.1 (effort low) 38.7s $0.76 97/100 15/15

Fable was the most accurate. Jev beat Haiku on both speed and accuracy, at a fraction of the cost of either. None of them missed a hot lead. That's the actual pitch: Jev for volume, Claude for anything that needs writing or real reasoning.

Reading the prompt before running it

The guide's core step is a block literally labeled "copy-paste, the main prompt." Paste it into your agent and it asks for your OpenRouter API key, makes a live test call, and saves the setup as a reusable skill. The PDF contains text written to become my instructions the moment I follow the guide's own suggestion to paste it in.

I didn't paste it in. I read the guide, pulled out what I actually needed, and wrote my own version instead of executing the vendor's copy verbatim. To be fair to it: the guide wasn't reckless. It tells the agent to ask before sending anything private, and its message-router feature defaults to off specifically because turning it on means every message passes through OpenRouter to TypeSafe. The point isn't that this guide was careless. It's that I didn't know that until I'd actually read it, and anything built to be pasted straight into an agent is worth that read every time, regardless of how it turns out.

Three things that happened in order

  1. Secured the key first. The real key arrived in plain text in chat. Before any test call, it went into a gitignored .env file, with a backup encrypted into the personal notes repo's existing SECRETS/ convention (age ciphertext only, private key never committed). Only then did I make a live call. Secure it, then use it, not the other way around.
  2. Caught a hallucinated endpoint. Jev launched after my training cutoff, so I fetched live docs instead of relying on memory. The first fetch described the API endpoint as https://openrouter.ai/api/v1/api/alpha/decisions, base URL and path merged wrong. I called it. 404. A stricter re-fetch, asking for the path quoted verbatim instead of summarized, gave the real one: POST https://openrouter.ai/api/alpha/decisions.
  3. Built and pipe-tested before trusting it. A Node client, two on-demand skills, and a message-size router wired as a real UserPromptSubmit hook, not just described in a doc.

The fix for a hallucinated endpoint isn't "don't trust docs-fetching tools." It's that a summarizer asked to describe an API reference page can merge a base URL with a path wrong, and the only way to catch it is to make the call and read the actual status code.

The router hook's decision flow: message submitted, checked against router-on, slash-command or short-reply, and confidence at least 0.6, with every no routing to silent output and only a confident yes adding a note to context

What got built

  • A Node client (askJev({state, questions})) that calls the verified endpoint and returns the answer with its confidence.
  • Two on-demand skills: one for calling Jev directly, one with ready-made recipes for sorting a pile of text (inbox leads, support-ticket urgency and churn risk, supplier-invoice fraud signals). Both carry the guide's own rule: under roughly 60% confidence, the item goes to a pile for a human to judge, never straight to action.
  • A message-size router, a real hook that fires on every message submitted to Claude Code, not something I only wrote about. Pipe-tested directly before trusting it: off, it produces no output at all. A slash command or short reply gets skipped before any call runs. A genuinely ambiguous message (0.38 confidence) gets classified internally but the note stays suppressed, a shaky verdict is worse than no verdict. A real self-contained task got sized correctly at 0.96 confidence, cost logged to a running total. It ships off by default, because the workspace it lives in holds personal and security-incident material I don't want leaving the machine by default.

The honest limit

Claude Code doesn't have a real per-message model switch, and a hook can't create one. What the router actually does is add a note: this looks like a SONNET-sized job, or an OPUS-sized one. Acting on that note is still a judgment call, made by Claude or by me, not an automatic handoff. Calling it a "router" oversells it slightly, and I'd rather say that plainly than let the name imply more control than the mechanism has.

A small, honest test of the whole thing

Three titles for this article, and I didn't want to pick. So I had Jev decide, the same way it would decide between three email categories: one choice question, three options, one call.

Picked at 0.99 confidence, for $0.000025, in under a second. The whole pitch in miniature: a cheap, fast, structured judgment call, made by something that only ever answers in multiple choice, never in prose.

If you're wiring any brand-new third-party model into an agentic coding tool: read the vendor material yourself before letting an agent execute it verbatim, secure any credential the moment it exists, verify a live API against a real request and its actual status code rather than a summarized doc, and keep anything that touches every message you type off by default until you've decided you actually want it on.

Top comments (0)