DEV Community

Galih Putro Aji
Galih Putro Aji

Posted on Originally published at Medium on

A $40 Million Startup Built an AI That Can’t Chat. Three Days Later, Its Free Version Appeared

Jev from TypeSafe and Laya from the community: what was promised, what was proven, and who actually came first.

On September 15, 2026, TypeSafe AI emerged from stealth with a $40 million seed round led by DCVC and a product deliberately designed to be unable to write sentences. Its name is Jev. Three days later, an engineer from Kerala named Nandakishor M released Laya, an open-source model that mimics Jev’s API, and its repository now has around 25,700 stars on GitHub. To Analytics India Magazine, he claimed that he had been working on the idea since April 2025, while researching ways to detect hallucinations before an LLM finished responding.

The story is easy to summarize: an expensive startup, copied by the community within days. The more honest version is messier. What was copied was the shape of the interface, not the model itself, and the Laya figures most frequently quoted come with footnotes that are rarely carried along with them.

Jev, an AI that only answers choices

Ordinary language models receive text and respond with text. Jev receives a “state” (an email, JSON document) along with a list of typed questions, then returns answers that can be read directly by a program: a single choice, a score, or a yes/no probability, all accompanied by a confidence level. Its founder, Diogo Almeida, describes Jev as a “frontier-intelligence function call.”

An example from Laya’s README, which uses the same format: a customer email complains about being charged twice and threatens to cancel the subscription unless refunded that day. The model answers “billing department,” plus a churn probability of around 89 percent. The decision threshold is in your code; above a certain number, the process runs automatically, while below it, it is routed to a human.

TypeSafe’s claims are bold. The price is $0.042 per million input tokens with free output, response time is 70 to 500 milliseconds, and their homepage states that it is 193.6x faster and 444.6x cheaper. Their launch blog itself acknowledges that those figures come from four workflows created by the internal team that may be biased, and that such a low price has not yet been proven to be free of subsidies. The transparency is commendable, but the benchmark is still their own.

The phrase “cannot hallucinate” needs to be read narrowly. What it means is that the answer cannot fall outside the schema you define, not that the answer is guaranteed to be correct. A wrong answer can still arrive in a perfect format.

Almeida previously worked on RLHF at OpenAI, and Forbes reports that its seed valuation was around $200 million. TypeSafe trained Jev entirely on synthetic data and has not disclosed its architecture; according to TechCrunch, outside observers suspect that there may be an open-weight LLM underneath.

Laya: the same interface, a different engine

Laya is a ModernBERT-large encoder with 421 million parameters and a decision head on top. All questions are answered in a single forward pass, taking around 33 to 40 milliseconds on a T4 GPU, licensed under Apache 2.0, and pip install laya is enough to try it. Its built-in server, laya-serve, exposes a POST /v1/systemone endpoint with the same request and response shape as Jev's API, according to its model card, so a Jev client only needs to change the base URL.

That is why the word “clone” is not quite accurate. Jev’s weights are closed and its architecture has not been published, while Laya is a standalone bidirectional encoder. What they share is only the concept and the API contract.

The community’s enthusiasm is real: around 25.7 thousand stars, 2.2 thousand forks, and a model card with 4.12 thousand likes. Ports for Apple silicon and Huawei Ascend NPUs have already appeared. One example documented in the README: fine-tuning Laya for a browser agent increased top-1 element selection accuracy from 0.10 to 0.66.

The numbers that need to be read twice


Jev and Laya accuracy according to the comparison table in Laya’s README. Jev’s figures come from a third party

The comparison table in the README places Laya above Jev on the typed-decisions benchmark: 0.766 versus 0.727. Read the footnote. The 0.766 figure belongs to a checkpoint that has already been fine-tuned on the benchmark’s own training split. Its base checkpoint only reaches 0.362, below the majority-class guess of 0.461, and the README itself states that Laya is a fast foundation for specialization, not a zero-shot decision engine. AG News is included in its training-data mixture. On Banking77, with dozens of labels, Jev leads by a wide margin: 0.870 versus 0.425.

The Jev figures in that table are not measured by Laya; its developer cites them from third-party testing. One of them, a 300-example pilot by AbdelStark comparing Jev with the GLiNER2.5 model, provides a mixed picture. Jev wins decisively on AG News and Banking77. On the emotion dataset, its accuracy is only 0.480, and on 16 percent of examples it assigned zero probability to the correct label. Jev’s latency of 236 to 256 milliseconds in that pilot was measured from France over the internet, while Laya’s 33 milliseconds runs locally on a GPU. Part of the difference is simply network distance.

Laya itself starts out with probabilities that are too overconfident. The average expected calibration error (ECE) of its default checkpoint is 0.466, and only falls to 0.081 after temperature fitting on held-out data. This means the confidence scores from the default version are not yet trustworthy before you tune them on your own data.

Who came first?

Nandakishor claims that he published the idea a year before Jev through two arXiv papers, and accuses TypeSafe of launching the same concept without a paper, open weights, or a dataset. His first paper, SalesRLAgent, predicts sales conversion probabilities using reinforcement learning and synthetic data from GPT-4o. His second paper routes LLM queries based on estimated confidence before the model answers, and its experiments use SmolLM2–360M with 72 training examples.

The two are aligned: a small model trained with RL to measure its own confidence. However, the text of the second paper does not contain typed questions based on choice, score, or boolean values, nor a reward based on a proper scoring rule, which is central to Laya. On Dev.to, Nandakishor describes the paper as an exact framework for schema-based decisions, and that description is broader than the paper itself.

TypeSafe’s side is equally fragile. The basic construction — an encoder that scores labels provided only at inference time — has existed for years through NLI-based zero-shot classification and the GLiNER family of models since 2023. An analysis on Substack by akmaier concludes that both sides claim more originality than is supported by the public record. TypeSafe, according to coverage through the end of September, has not responded publicly.

The real test

The idea belongs to no one. TypeSafe packaged it into a ready-to-use service without fine-tuning; Nandakishor showed how quickly such an API contract could be reproduced, and that half of its value only appears after the model is trained on your own data. Both can be true at the same time.

If your ticketing or content-moderation workflow is already running on top of an LLM, do not stop at the README. Collect two hundred to three hundred real examples that you have already labeled, run both models, then calculate not just accuracy, but what percentage of decisions can be automated at a five-percent error threshold. That is the metric used by the AbdelStark pilot, and it determines how many tickets truly no longer need to be handled by a human. That number does not appear in any README.

The Bottom Line

Ultimately, the standoff between Jev and Laya is less about who holds the deed to discriminative decision architectures and more about a practical shift away from generic generative text toward strict schema reliability. While TypeSafe packages an instant, out-of-the-box pipeline for teams that want immediate results without fine-tuning, Laya demonstrates the agile power of open-source models ready to be specialized and calibrated directly on internal domain data. Instead of relying on headline-grabbing benchmarks or marketing claims, the real decision rests in your own operational pipeline: run a few hundred real examples through both, measure your risk thresholds, and see which one actually reduces human workload without breaking production


Top comments (0)