Claude Code makes a lot of small judgments that don't need a full LLM call: does this edit break a rule in my CLAUDE.md, which of my installed skills fits this prompt, does this diff need a careful review or a quick pass. I wanted to see whether a small, fast "decision model" could handle those instead. The result is jev-tools, an early Claude Code plugin. This post covers how it works, what broke when I ran it live, and what I actually measured, including the parts that did not work.
Repo: https://github.com/Rcidshacker/jev-tools (MIT, Python 3.10+, standard library only).
I'm the author. This is an independent project and is not affiliated with Codiv, OpenJev or TypeSafe AI.
The model in two minutes
OpenJev, served by Codiv, is what Codiv calls a "System One" model. It does not write text. You send a state (any string or JSON) plus typed questions, and it returns one answer per question in tens to hundreds of milliseconds:
| Question type | You send | You get |
|---|---|---|
| yes/no | instructions | a probability of yes |
| pick one | instructions + options | the pick and a probability per option |
| score | instructions + ordered levels | an expected level and per-level probabilities |
Because the answers are probabilities, decisions become thresholds in ordinary code. And because the model never writes free text, there is nothing to parse and nothing to hallucinate. (Codiv describes the probabilities as calibrated. I have not independently verified that.)
What the plugin does
-
Rule hook (
PreToolUseon Edit/Write): reads yourCLAUDE.md/AGENTS.mdrules, asks the model whether the pending edit breaks one, re-checks any hit with a stricter question, and inactivemode blocks the write and tells Claude which rule it broke. - Skill picker (opt-in): picks the one installed skill that fits your prompt.
-
review-precheck: seven yes/no policy questions on a git diff (secrets, new dependencies, auth, schema, weakened tests, swallowed errors, risky logic) decide "fast pass" versus "full review". -
rule-calibrate: replays your recent commits against your rules and tells you which rules are decisive, noisy, weak or quiet, before you enforce anything. -
find-filesandbrowser-nav: file discovery and a click-by-click browser navigator. -
status: shows whether the install is alive.
Design rules I set before writing code
- Thresholds live in code. The model answers typed questions; plain code decides what to do with the probability.
- Explicit failure policy per piece. The rule and skill hooks fail open, so an outage never blocks your edit or your prompt. The review pre-check fails safe: an outage routes to a full review, because silently skipping a review is the dangerous direction.
-
Shadow first. The plugin ships in
shadowmode. Hooks only log what they would have done. Nothing blocks or injects until you switch toactive. - Stdlib only, one shared client. Nothing to install, one file to audit.
What live testing broke
The first version passed all its offline tests. Running it against the real API and real Claude Code sessions found problems the mock never could:
| Finding | Fix |
|---|---|
| Every live call returned 403 | Codiv's edge rejects Python's default User-Agent, so the client now sends its own |
| A clean edit that reads its host from config scored 0.91 against "do not hardcode API hosts" | Added a strict second look on any hit; the same edit is now vetoed at the second question |
The browser navigator said blocked on a page where the goal was met |
The goal-met probability wobbled around the bar between identical calls (0.93, then below 0.8, same input). Added "nothing left to click and goal probably met means done" |
| Skill picker could inject on a shallow match | Added a stage-two confirmation on the pick's full description |
| File discovery scored 3 of 4, then 0 of 4 after a rewrite | Built a labelled evaluation of six variants instead of trusting one run |
What I measured
Everything below ran live against Codiv's hosted OpenJev in September 2026. The samples are small and run-to-run noise is real, so read these as smoke tests, not benchmarks.
| Component | Result |
|---|---|
| Rule enforcer | After the second look: 4 of 4 planted violations blocked, 4 of 4 clean edits allowed. Also blocked and allowed correctly inside real headless Claude Code sessions |
| Skill picker | 3 of 3 correct on my real skill roster (two matches, one correct "none") |
| Review pre-check | A rename-only diff routed fast. A diff with a hardcoded key, a swallowed exception and an emptied test file routed full with the right flags |
| Browser navigator | 5 of 5 steps on a synthetic login page. Never run against a real browser session |
| File discovery | No better than plain keyword counting: top-3 hits 4 to 6 of 8 versus 4 of 8, and identical reruns differ by up to 2 |
| Cost and speed | About 1 second per prompt or edit, 2 seconds when a violation is confirmed, about 5k input tokens per edit at 20 rules |
Two honest caveats. First, I wrote both the test edits and the rules, so the rule-enforcer numbers flatter the model. The calibration replay on a different project is the more honest signal: 20 rules against 24 real hunks, and nothing fired above 0.35. That is what a good rulebook on clean history should show, and it is also what a blind rulebook would show. Second, file discovery ships as a hint, not an oracle, because it did not beat a keyword counter. A one-week field report on a similar skill router found only about 5% of its suggestions were followed by the agent, which is why the skill hook is off by default and I recommend staying in shadow mode.
Privacy: what leaves your machine
This plugin sends text to a third-party API (api.codiv.ai). Depending on the piece, that is the file name and diff of a pending edit, your project rules, your prompt (only if you enable the skill hook), or a git diff. Shadow mode still sends, because the hook needs the answer to log it. Only JEV_MODE=off sends nothing.
A local redactor runs first and replaces private keys, vendor-style API keys, JWTs, bearer tokens, credentials in connection strings and secret-named assignments with [REDACTED]. Files named like secrets (.env*, *.pem, id_rsa* and similar) are never read into a request. It is a pattern match, not a guarantee. It does not catch names, emails or customer and employee data. Codiv's public docs say nothing about how request data is stored, so treat this like pasting into a public forum, and turn it off for any project with regulated or client data.
Known gaps
- File discovery does not beat keyword search on my evaluation.
- The browser navigator has only seen a synthetic page.
-
review-precheckandrule-calibratewere exercised by script, not through the plugin's skill loader. - The hooks call
python, so a system that only haspython3needs an alias. - All thresholds are defaults (flag at 0.80, confirm at 0.70), not tuned values.
Try it, and tell me what is wrong with it
You need Python 3.10+ on your PATH and a free OpenJev key from Codiv. In Claude Code:
/plugin marketplace add Rcidshacker/jev-tools
/plugin install jev-tools@jev-tools
Set the key as a user-level OPENJEV_API_KEY environment variable, never in a repo, then run the status skill to check the wiring. Leave it in shadow for a week and read the log before you turn on anything that blocks.
What I would like to hear:
- Would you run this in shadow mode on a real repo? What would stop you?
- Which other judgments would you hand to a cheap decision model?
- Are the default thresholds sensible, or how would you tune them?
- How painful was install, especially on Windows?
- What is missing from the privacy setup?
Comments here or issues on the repo both work.
Credits
The ideas came from projects that got there first, re-implemented small with no code copied: abide (rule calibration, second look), hermes-jev-skills (confidence floor, two-stage retrieval), jev-kit (shadow-first rollout), jevgate (confirming before acting) and jev-skill-router (plugin layout and an honest field report). Built with Claude Code.
Top comments (0)