Making Jev-like structured inference simpler, openly.

At FuturixAI, we believe contributing to AI means more than building products. It means exploring different approaches, sharing what we learn, and making our work available for others to examine and improve.
Today, we're open-sourcing CARVE, an independent Jev-like model family built on Qwen3.8-27B and released under the Apache-2.0 license.
A different approach to structured outputs
Many AI applications require structured responses rather than free-form text. Whether it's classifying customer feedback, selecting an option, or assigning a score, the output needs to follow a predictable format that another system can use.
CARVE explores how to achieve this using Qwen's generative capabilities. Through few-shot prompting, the model is guided to generate free-form JSON in the required structure, which can then be parsed into typed Noul, Choice, or Score answers through a TypeSafe-compatible System One API.
This differs from Jev's approach, which directly scores candidate options rather than generating a structured response. CARVE investigates how far prompting and output constraints can take a conventional generative model towards structured decision-making.
What our evaluations show
On the public Decision 1.0 transfer-v9 benchmark, CARVE achieved 82.22% accuracy, correctly answering 860 out of 1,046 decisions.
For comparison, Jev's published score is 87.19%, while Kev-9B's is 79.25%. CARVE was measured locally, whereas the competitor scores come from externally published results. These figures compare accuracy on the same named protocol, not performance under controlled, identical hardware conditions.
On an NVIDIA RTX PRO 6000 Blackwell Server Edition, CARVE recorded 72.5 ms p50 latency and 91.0 ms p95 latency under warm serial inference. Under concurrent load, it sustained 20.23 requests per second, with zero observed HTTP or response-validation failures across 5,120 requests.
Our experiments also examine prompt strategies, few-shot formats, model weights, caching, adaptive demonstrations, and production behaviour. Results vary by task, reinforcing the importance of evaluating configurations rather than assuming one approach works best everywhere.
Built to be examined and improved
At FuturixAI, we believe open source gives ideas room to evolve beyond the teams that build them.
CARVE includes public protocols, frozen revisions, raw summaries, ablations, and negative results alongside the implementation. These resources allow developers and researchers to reproduce experiments, examine limitations, and investigate where the approach can be useful.
We're releasing CARVE under Apache-2.0 so the community can inspect the code, experiment with different prompting strategies, challenge our findings, and contribute improvements.
CARVE is an independent project. Its Jev-like designation describes its task category and interface; it is not affiliated with or endorsed by TypeSafe AI.
Let's build on it together
CARVE is FuturixAI's contribution to the open-source AI ecosystem. We're excited to see how developers and researchers evaluate it, where they find limitations, and what they build with it.
Explore the repository, reproduce the results, and share your feedback. Contributions are welcome.
Built on an open-source foundation. Released by FuturixAI for the community.
Top comments (0)