DEV Community

aimodels-fyi
aimodels-fyi

Posted on Originally published at aimodels.fyi

A beginner's guide to the Laya model by Convaiinnovations on Huggingface

This is a simplified guide to an AI model called Laya maintained by Convaiinnovations. If you like these kinds of analysis, you should join AImodels.fyi or follow us on Twitter.

Overview

laya is an open-source, non-autoregressive System 1 decision model from convaiinnovations. It accepts a text state—such as an email, support ticket, conversation state, or JSON document—and typed questions, then returns structured answers with probabilities and confidence scores. The most important point is that it is a decision model, not a text generator: it does not produce free-form text, which removes generation parsing errors and text hallucinations, but it also cannot explain its decisions in natural language. The model uses a fully fine-tuned, bidirectional ModernBERT-large backbone with 395 million parameters and a decision head trained from scratch. The head contains two transformer layers, an option-marker scorer, and an act/escalate head, bringing the total to 421 million parameters. Each question has a 512-token budget covering the question, options, and state; longer states are truncated. The model runs through the transformers ecosystem and the provided laya Python package, with loading through laya.load("convaiinnovations/laya"). Training used 100% human-annotated real-world datasets, 7,313 updates, one epoch, and about 1.96 hours of training. The checkpoint includes email triage, phishing detection, department routing, conversation trajectory modeling with TD(lambda = 1.0), and per-option-count temperature calibration.

Best use cases

Support-ticket and email routing. Use choice questions to route messages to billing, technical support, sales, or other departments. The checkpoint reports 99.1% accuracy and ECE of 0.009 for intent and customer routing, making this its strongest reported task family. Its typed option probabilities also support routing policies that send uncertain cases to human agents.

Email triage and phishing screening. The model can classify spam or phishing risk, identify department ownership, estimate urgency, and detect churn signals in one batched call. The README example combines four questions over one customer email. Email triage and phishing performance is lower than routing performance—73.2% accuracy—but its ECE of 0.017 indicates useful calibration on the reported benchmark. Treat the probability as a gating signal, not as proof that an email is safe.

Moderation and content-safety decisions. Use choice or noul questions for policy labels, safety boundaries, or escalation decisions. The checkpoint reports 96.7% accuracy and ECE of 0.061 for moderation and safety. The model returns probabilities rather than generated moderation prose, which fits systems that need deterministic downstream actions.

Ordinal prioritization and urgency scoring. Use score questions with criteria such as “not urgent,” “soon,” and “critical deadline or blocking issue.” The output includes an expected ordinal level, a distribution, and confidence. RLCD training includes ranked probability score rewards for ordinal questions, which targets calibrated distributions rather than only the top label.

Selective automation and human escalation. Use confidence thresholds to automate high-confidence cases and escalate the rest. The benchmark reports 92.2% accuracy at 50% coverage with confidence-gated automation at confidence ≥ 0.85. This makes the model a fit for queues where abstention or escalation matters more than maximum coverage.

Limitations

The model supports text only and English only. It does not accept images, audio, or other modalities. Each question has a 512-token input budget covering the question, options, and state, and longer states are truncated. If a document contains important information near the end, truncation can change the decision.

The model does not generate explanations, summaries, rewritten text, or other natural-language responses. Its claim of eliminating hallucinations applies to text generation: it still can make incorrect classifications or produce poorly calibrated probabilities outside its evaluation distribution. In-task macro accuracy is 0.838 with macro ECE of 0.060, but zero-shot performance on held-out task families falls to 0.651 accuracy with macro ECE of 0.207. That gap is the clearest warning against deploying the checkpoint without local validation.

Email triage is a weaker area than routing, moderation, emotion and tone, or fact checking. Arithmetic, counting, date comparisons, and multi-hop index lookups should remain in deterministic code. The model should not replace a rules engine for exact calculations or structured database retrieval.

The README provides no VRAM requirement, model file format, quantization configuration, or fixed CPU latency. It reports about 33–38 ms for multi-question evaluation on GPU, 38.4 ms P50 latency for one question with a 42.1 ms P95, 156.0 ms for 10 questions with a 158.4 ms P95, and 721.4 ms for 50 questions. Hardware claims cover commodity GPUs, Mac MPS, and CPU, but the supplied material does not specify hardware models or memory requirements.

The Apache 2.0 license permits commercial use subject to the license terms. Convai Innovations offers commercial support, enterprise integration, and custom fine-tuning. The supplied material does not describe demographic bias testing, safety red-team results, or a maintenance schedule, so assess those areas before production use.

How it compares

llava-llama-2-13b-chat-lightning-preview

Choose laya for typed classification, calibrated probabilities, low-latency batching, and local decision workflows over text. Choose llava-llama-2-13b-chat-lightning-preview for an autoregressive chatbot that needs free-form responses or multimodal interaction. The tradeoff is specialization versus generality: laya has 421 million parameters, a 512-token question budget, and reported 38.4 ms single-question P50 latency, while LLaVA is a 13-billion-parameter autoregressive model trained for multimodal instruction following. LLaVA uses the Llama 2 Community License; laya uses Apache 2.0.

llava-v1.5-13b

Choose laya when the output must be a typed choice, ordinal score, or boolean probability and the system needs confidence-aware escalation. Choose llava-v1.5-13b when a 13-billion-parameter autoregressive multimodal assistant is more suitable for open-ended instruction following. The available data does not provide a directly comparable latency or accuracy benchmark for LLaVA, so do not treat laya’s reported 83.8% in-task macro accuracy as a head-to-head quality result. LLaVA has a Llama 2 Community License, while laya is Apache 2.0.

llava-v1.5-7b

Choose laya for English text decisions, structured outputs, calibration, and batch evaluation of many questions. Choose llava-v1.5-7b when you need a smaller LLaVA chatbot for multimodal, autoregressive instruction-following tasks. The 7-billion-parameter LLaVA variant is larger than laya but may offer broader generative capability; the supplied information does not provide VRAM, speed, or task-accuracy comparisons. laya offers Apache 2.0 licensing and self-hosted inference with no recurring token charge, while LLaVA uses the Llama 2 Community License.

llava-v1.5-7b

Choose laya when you need a decision endpoint rather than a conversational model, especially for routing, moderation, phishing screening, and selective automation. Choose llava-v1.5-7b when you need multimodal input or generated answers. The supplied links contain a duplicate model name and URL, so there is no separate technical specification to compare. The central tradeoff remains typed, calibrated classification versus autoregressive multimodal conversation.

llava-v1.5-13b

Choose laya for local, air-gapped decision systems with explicit confidence thresholds and no data egress. Choose llava-v1.5-13b for broader multimodal instruction-following and free-form generation. The supplied material does not establish which model has higher general quality across shared tasks; it does establish that laya reports task-specific calibration, latency, and selective-automation results, while LLaVA is described as an autoregressive chatbot trained on GPT-generated multimodal instruction-following data.

Technical specifications

laya uses a non-autoregressive, bidirectional architecture:

  • Backbone: ModernBERT-large, 395 million parameters, fully fine-tuned.
  • Decision head: two transformer layers, an option-marker scorer, and an act/escalate head.
  • Total parameters: 421 million.
  • Option scoring: each option is scored at its own [MASK] marker token; a softmax is applied across the options for that question.
  • Input budget: 512 tokens per question, including question, options, and state.
  • Batching: all questions can run in one forward pass.
  • Reported GPU multi-question latency: about 33–38 ms.
  • Training method: RLCD, or Reinforcement Learning for Calibrated Decisions.
  • RLCD objective: log score and spherical score; ranked probability score is added for ordinal score questions.
  • Exploration: zero-mean Gaussian noise added to logits.
  • Dialogue training: Temporal Difference learning with Monte Carlo targets, TD(lambda = 1.0), over prefix slices to prevent outcome leakage.
  • Data: 100% human-annotated real-world datasets; no synthetic shortcuts.
  • Fine-tuning: 7,313 updates, one epoch, about 1.96 hours.
  • Calibration temperatures: [1.637, 1.251, 1.983], with per-option-count scaling.
  • Library: transformers; Python package: laya.
  • License: Apache 2.0.
  • Deployment claims: commodity GPUs, Mac MPS, CPU, local, air-gapped, and on-device operation.
  • Cost claim: $0.00 per 1 million input tokens when self-hosted. The comparison table lists $0.042 per 1 million input tokens for TypeSafe Jev.
  • No quantization options, VRAM requirement, model file format, or exact CPU benchmark is specified.

Reported evaluation results include 0.838 in-task macro accuracy and 0.060 macro ECE, versus 0.651 zero-shot macro accuracy and 0.207 macro ECE. Task-family results are intent and routing at 0.991 accuracy and 0.009 ECE; moderation and safety at 0.967 and 0.061; emotion and tone at 0.906 and 0.018; email triage and phishing at 0.732 and 0.017; and inference and fact checking at 0.883 and 0.054. Detailed results, reliability diagrams, and risk-coverage curves are in the repository’s eval/ directory.

Model inputs and outputs

Inputs

  • State: Text, email, ticket, conversation state, or JSON document.
  • Question collection: A mapping of question names to typed question definitions.
  • choice question: An instruction plus named options and descriptions, represented in the example by a criteria dictionary.
  • score question: An instruction plus an ordered list of rubric levels, represented in the example by a criteria list.
  • noul question: An instruction that defines a boolean proposition.
  • Language: English.
  • Length: 512 tokens per question across the question, options, and state; longer states are truncated.
  • Modality: Text only.

Outputs

  • choice: Selected option, probability for each option, and calibrated confidence.
  • score: Expected ordinal level, distribution across rubric levels, and confidence.
  • noul: Calibrated probability P(true) from 0.0 to 1.0.
  • Result container: The example reads answers from result["answers"].
  • Post-processing: No generated-text parsing is required. Applications should apply their own confidence thresholds, escalation rules, deterministic arithmetic, date logic, and database lookups.

Getting started

Install the package and load the checkpoint:

pip install laya
Enter fullscreen mode Exit fullscreen mode
import laya

agent = laya.load("convaiinnovations/laya")

state = {
    "from": "customer@acme.com",
    "subject": "Duplicate billing on March invoice #4411",
    "body": (
        "Hi team, we were billed twice for March. Please refund the duplicate "
        "before Friday or we will cancel our plan."
    ),
}

questions = {
    "department": {
        "type": "choice",
        "instructions": "Which department should handle this email?",
        "criteria": {
            "billing": "invoices, payments, refunds",
            "technical": "bugs, outages, integrations",
            "sales": "pricing, contracts, demos",
            "other": "everything else",
        },
    },
    "urgency": {
        "type": "score",
        "instructions": "How urgent is this request?",
        "criteria": [
            "not urgent",
            "soon",
            "critical deadline or blocking issue",
        ],
    },
    "churn_risk": {
        "type": "noul",
        "instructions": (
            "Does the user threaten to cancel or switch to a competitor?"
        ),
    },
    "is_phishing": {
        "type": "noul",
        "instructions": "Is this email a phishing or scam attempt?",
    },
}

result = agent.predict(state, questions)
answers = result["answers"]

print(
    "Department:",
    answers["department"]["choice"],
    f"(confidence: {answers['department']['confidence']:.2f})",
)
print("Urgency:", f"{answers['urgency']['score']:.2f} / 2.0")
print("Churn risk:", f"{answers['churn_risk']['noul']:.1%}")
print("Phishing:", f"{answers['is_phishing']['noul']:.1%}")
Enter fullscreen mode Exit fullscreen mode

Frequently asked questions

Q: Can I use laya commercially?

A: The checkpoint is released under the Apache 2.0 License, which supports commercial use subject to the license terms. Convai Innovations also offers commercial support, enterprise integration, and custom fine-tuning.

Q: What hardware or VRAM do I need?

A: The README states that it runs on commodity GPUs, Mac MPS, or CPU. It does not specify a VRAM minimum, supported GPU models, or CPU memory requirement, so measure resource use on the target deployment.

Q: How fast is inference?

A: The reported one-question latency is 38.4 ms P50 with 42.1 ms P95 on GPU. Ten questions take 156.0 ms with 158.4 ms P95, and 50 questions take 721.4 ms; the architecture section reports about 33–38 ms for multi-question evaluation in one forward pass.

Q: What input format does laya expect?

A: Pass a state containing text, an email, a ticket, a conversation state, or a JSON document, plus a mapping of typed questions. Each question uses choice, score, or noul and includes an instruction with options or rubric criteria where required.

Q: Is calibration reliable for every task?

A: No. Calibration is measured on benchmark datasets and must be tested on your distribution. In-task macro ECE is 0.060, while zero-shot macro ECE rises to 0.207.

Q: What are the main failure modes?

A: Longer inputs are truncated at 512 tokens per question, and the model is not intended for arithmetic, counting, date comparisons, or multi-hop index lookups. Email triage and phishing performance is 73.2% accuracy, so those decisions need threshold testing and human escalation.

Q: Can I use it for free-form answers or explanations?

A: No. It returns typed decisions, probabilities, distributions, and confidence scores. It does not generate explanatory text, summaries, or conversational responses.

Q: Can I fine-tune the checkpoint?

A: The maintainer offers custom fine-tuning, and the checkpoint is distributed under Apache 2.0. The supplied documentation does not provide a fine-tuning command, dataset schema, or training recipe beyond the reported RLCD procedure.

Click here to read the full guide to Laya

Top comments (0)