If you are working out how to build agentic AI this year, the first decision is no longer which framework to adopt. It is which model handles the decisions your agent makes, and there are now two credible answers that arrived ten days apart: Needle 3 from Cactus Compute, an open on-device tool-calling model, and Jev from TypeSafe AI, a hosted decisions-only API.
Verdict: Needle 3 wins for builders who own their stack. It is open source, ships at 8-29 MB, runs CPU-only on a Raspberry Pi or a laptop, can be fine-tuned on your own data, and costs nothing per call (Cactus Compute). Jev wins for a narrower case: teams that need calibrated yes/no probabilities they can set thresholds against, sub-second latency, and zero operational work — accepting a closed hosted API with no fine-tuning and no offline mode (TypeSafe specs). If you are learning, prototyping on edge hardware, or handling data that cannot leave the device, start with Needle.
TL;DR
- Needle 3 (Cactus Compute): 8-29 MB open model that does tool calling, structured extraction, and embeddings. No chat, on purpose.
- Jev (TypeSafe AI): hosted "System One" API returning three primitives only — Choice, Score, and Noul (a calibrated yes/no probability). Early access opened 15 September 2026.
- The split that matters: writing layer (frontier cloud model) → decision layer (Needle locally or Jev hosted) → code that acts.
- Cost shape: Needle is free on your hardware; Jev is $0.042 per million input tokens with output free, since a pick is barely any output (TypeSafe specs via flowtivity.ai).
- Pick Needle if you want to own weights, run offline, and fine-tune. Pick Jev if you want probabilities with thresholds and no infrastructure.
- Last verified: 2026-09-25.
What is the difference between Needle 3 and Jev?
Needle 3 is described by Cactus Compute as an automation foundation model for tiny devices, and it deliberately gives up general conversation to do three narrow jobs well: picking and filling tool calls, pulling structured fields out of messy text, and producing embeddings for local search and routing (cactuscompute.com/needle). The lineage is worth knowing. Cactus first open-sourced Needle as a 26M-parameter model distilled from Gemini's tool-calling behaviour (Show HN thread), and the prior generation was a 14 MB function-calling model that turned plain English into local actions on a Raspberry Pi 5 with CPU only (Raspberry Pi news).
Jev takes the opposite route. TypeSafe AI opened early access on 15 September 2026 after a $40M seed round led by DCVC (launch coverage). It never generates text. You send it a state — a string, a JSON object, or an array — plus typed questions, and it scores all predefined answer options in one parallel pass. You get back a Choice from up to 255 options, a Score on a 2-10 rubric, or a Noul: a calibrated probability on a yes/no question. There is no string to parse and no off-schema answer to defend against (Jev engineering notes).
| Dimension | Needle 3 (Cactus) | Jev (TypeSafe) |
|---|---|---|
| What it is | 8-29 MB open on-device tool-calling model | Hosted decisions-only API |
| Returns | Tool calls, structured extraction, embeddings | Choice (255 options), Score (2-10), Noul (yes/no) |
| Runs on | Pi 5, Mac, phones, CPU only | TypeSafe cloud, POST /v1/systemone
|
| Price | Free; your hardware | $0.042 per M input tokens, output free |
| Latency | ~1,200 tokens/s decode on consumer CPUs | 70-500 ms reported |
| Fine-tuning | Yes — needle finetune plus LoRA |
No — shape via state and question instructions |
| Chat/text | None by design | None by design |
| Privacy | Data never leaves the device | State leaves your network |
| Best for | Builders who own the stack, offline, cost floor | Teams needing calibrated probabilities, zero ops |
Why does the split architecture beat picking a framework?
Because the layers now have different best-in-class answers. The writing and reading layer still belongs to frontier cloud models. In our own harness, across three trials each on an identical seven-constraint article-planning task, Gemini 3.8 Flash (High) and Claude Opus 4.6 (Thinking) both scored 17 of 17 on machine-checked constraint adherence; median wall time was 23 seconds for Gemini against 67 seconds for Opus (n=6, measured 2026-09-25, runs via the Antigravity CLI). Neither is the right tool for the binary routing calls an agent makes tens of thousands of times a day.
That is the gap Needle and Jev fill. Orchestration sits above both, and the framework question has its own answer — see our agentic AI frameworks comparison and the build-versus-buy view on orchestration. The distribution layer also settled recently: Agent Plugins 1.0 gave agent extensions a vendor-neutral packaging format, published by Vercel with AWS, Cursor, Microsoft and OpenAI, with Google joining as a maintainer the same day (Vercel).
How do you build with Needle 3 in three steps?
-
pip install cactus-needle. - Wrap a Python function as a tool with the
@needle.tooldecorator (MindStudio walkthrough). - Run the agent loop locally against those tools.
For hardware targets, build the engine and weights with needle build --platform macos-arm64, or squeeze an 8-layer subnetwork onto Pi-class hardware with needle build --platform linux-arm64 --layers 8 --out ./pi, then serve via needle --model needle3.cact --tools tools.json --serve (GitHub README). Fine-tuning is first-class rather than bolted on: needle finetune data.jsonl --epochs 10 --out adapter.safetensors, followed by needle build --lora adapter.safetensors --out tuned.cact. If your stack is Python-first, our Python for agentic AI guide covers the surrounding tooling.
How do you build with Jev in three steps?
- Get an early-access key; the current build is
jev-1.13.0behind thejev-latestalias. - Define your question spaces as Choice, Score or Noul, with the records and rules your agent should reason over placed in the state.
-
POSTstate plus questions to/v1/systemoneand branch on confidence thresholds. Official SDKs exist for Python (typesafe-sdk) and TypeScript (@typesafe-ai/sdk).
Reported results from the first week of early access give a sense of the cost shape: builders triaged 1,700 emails for 18 cents, and a site scan for low-quality AI text ran in 243 ms at $0.00015 per check (flowtivity.ai). TypeSafe's own comparisons claim up to 193x faster and 444x cheaper than Claude Fable 5.1 and GPT-6 Astra on System-One-shaped tasks (madewithjev.com). Treat vendor benchmarks as directional and measure on your workload.
What are the real tradeoffs?
Needle's cost is your engineering time. You manage builds per platform, you own the fine-tuning data, and quality on an 8-layer Pi subnetwork will not match the full model. Jev's cost is control. There is no fine-tuning — TypeSafe's position is that you shape behaviour by improving the state and the question instructions, and the same weights serve every account. Your state also leaves your network, which rules Jev out for some regulated or offline deployments. Neither model chats, so you still need a frontier model for drafting and summarising, and code for the acting layer. The reliability practices around all of this are covered in our engineering principles for agentic systems, and if you are still weighing whether an agent is warranted at all, see agentic AI versus traditional automation.
FAQ
Q: Can Needle 3 replace a frontier model like Gemini or Claude?
A: No. Needle 3 has no general chat ability by design; it handles tool calls, structured extraction and embeddings. You still need a frontier model for drafting, summarising and open-ended reasoning.
Q: Does Jev work offline or on-device?
A: Not officially. Jev is a hosted developer API with a single endpoint, and there is no consumer app. Community projects such as LocalJev and litjev imitate the request shape against local OpenAI-compatible models.
Q: What hardware do I need to run Needle 3?
A: A Raspberry Pi 5 class machine or better is enough. The previous generation ran CPU-only on a Pi 5 with 8 GB, and Needle 3 supports a reduced 8-layer build for constrained devices.
Q: What does a Noul actually return?
A: A calibrated probability on a yes/no question, which lets you set explicit thresholds — auto-approve above 0.9, route to a human between 0.6 and 0.9, reject below that.
Q: Which should I learn first if I am new to agentic AI?
A: Needle. It is free, local, open source, and installable with one pip command, so you can iterate on the tool-calling loop without a key, a bill, or a network round trip.
Q: Can I use both in one system?
A: Yes, and that is a reasonable design. Run Needle on-device for private, high-volume routing, and call Jev for the smaller set of decisions where you want calibrated probabilities you can audit.
Last verified: 2026-09-25. Vendor-reported figures are attributed to their source and have not been independently reproduced except where marked as our own harness runs. This article was produced with AI assistance under human editorial review; see how we work.
Top comments (0)