Mastering Jev and Laya: A Practical Guide to System One Decision Models

A large share of the "AI" in production systems today is a very expensive way of answering very small questions. Is this ticket about billing or a bug? Is this email a scam? Should this request go to the cheap model or the expensive one? Is the user trying to jailbreak the assistant?

We send those questions to a generative model, ask it to "respond only with JSON", parse whatever comes back, and write retry logic for the times it doesn't. It works, mostly. It is also slow, billed per output token, and prone to inventing a fourth option when you gave it three.

I've been building routing and triage logic into enterprise systems for most of my 29 years in this industry: CRMs, clinic management, law-firm case intake. Before LLMs it was rule engines and hand-built classifiers. After LLMs it became prompts. What arrived this month, Jev from TypeSafe AI and Laya from Convai Innovations, is the first thing that feels designed for this job rather than borrowed for it.

This guide explains what they are, how to use them well, where they fit, and where they don't. I ran every code sample below on my own workstation (one NVIDIA RTX 3090) with laya 0.3.21, and the outputs shown are the real outputs. Where I cite numbers I did not measure, I say so and link the source.

What a "System One" model is

The name comes from Daniel Kahneman's split between System 1 (fast, intuitive judgment) and System 2 (slow, deliberate reasoning). An LLM writing an essay is doing something like System 2. Deciding "this is a billing question" should be System 1: one glance, one answer.

A System One model does not generate text at all. You give it two things:

  • State: the thing being judged. A ticket, an email, a chat message, a JSON record.
  • Questions: typed, bounded questions about that state, with the possible answers defined up front.

It returns an answer for every question in a single forward pass, each with a probability distribution. There are three question types, and both Jev and Laya use the same three:

TypeYou defineYou get backTypical use
choiceA set of labelled optionsThe chosen option, a probability per option, confidenceRouting, intent, category
scoreAn ordered list of levelsAn expected score, a probability per level, confidenceUrgency, severity, difficulty
noulA yes/no statementThe probability that it is true (0 to 1)Flags, guardrails, signals

Because the model can only pick from options you defined, it cannot return an option that doesn't exist. That's not the same as being right, a point I'll come back to more than once, but it removes a whole category of parsing bugs.

Jev and Laya: same idea, opposite trade-offs

Jev launched on 15 September 2026 as a managed API. Laya followed three days later as open weights under the Apache 2.0 license, built by Nandakishor M at Convai Innovations. They share the same question format almost exactly, which makes it easy to prototype on one and move to the other.

Jev (TypeSafe AI)Laya (Convai Innovations)
How you run itHosted API onlyYour hardware: GPU, Apple Silicon, or CPU
WeightsClosedOpen, Apache 2.0
Price$0.042 per million input tokens, output freeFree; you pay for compute
Latency (published)70–500 ms end to end per TypeSafe; independent p50 of 236–276 ms~33–40 ms on a T4 per the project README
Options per choiceUp to 255Accuracy drops noticeably past ~20
Context64k tokens per request512 tokens (English), 1,024 (others), configurable up to 8,192
LanguagesEnglish first100+ via a multilingual checkpoint and automatic routing
Fine-tuningNot availableYes, with a free Kaggle notebook (2×T4, about 4–5 hours)
Zero-shot qualityStrong out of the boxWeak on general checkpoints; best after fine-tuning

The short version: Jev buys you quality without effort. Laya buys you control, speed and privacy, in exchange for effort. Neither is strictly better, and knowing when each one wins is most of the skill.

The rest of this guide is hands-on with Laya, because I could run it locally and verify every line. The design lessons (small questions, confidence gating, measuring before trusting) apply to Jev just as much.

Getting set up

Laya needs Python 3.10 or newer:

pip install laya

The first call downloads the checkpoint from Hugging Face. On a GPU it runs out of the box. It also runs on CPU, but slowly: one independent test on a 4-vCPU VPS saw a median of around 49 seconds per call. Plan on a GPU for anything real.

All the scripts in this guide, with pinned versions, are available to download as a zip or browse individually. Clone, install, and run them yourself; your numbers will differ with your hardware, and that's the point.

Lesson 1: your first decision, and how to read it

Here is a support ticket answered with all three question types in one call:

from laya import Router

router = Router()  # downloads the checkpoint on first use

state = (
    "Hi, we were billed twice for March. Please refund the duplicate "
    "today or we will cancel our plan."
)

questions = {
    "department": {
        "type": "choice",
        "instructions": "Which department should handle this ticket?",
        "criteria": {
            "billing": "invoices, payments, refunds",
            "technical": "bugs, outages, system errors",
            "other": "everything else",
        },
    },
    "urgency": {
        "type": "score",
        "instructions": "How urgent is this ticket?",
        "criteria": ["not urgent", "soon", "blocking"],
    },
    "churn_risk": {
        "type": "noul",
        "instructions": "Does the customer threaten to cancel or leave?",
    },
}

result = router.predict(state, questions)
answers = result["answers"]

print(answers["department"]["choice"], answers["department"]["probabilities"])
print(answers["urgency"]["score"], answers["urgency"]["probabilities"])
print(answers["churn_risk"]["noul"])
print(result["routing"]["model"], result["usage"])

Output:

billing {'billing': 0.987, 'technical': 0.0073, 'other': 0.0056}
1.7946 {'0': 0.0345, '1': 0.1365, '2': 0.8291}
0.8389
english {'input_tokens': 166, 'output_tokens': 0}

A few things to notice:

  • One pass, three answers. All the questions share one forward pass over the state. Asking three questions costs barely more than asking one (I measure this in Lesson 7).
  • score is an expected value, not a label. 1.79 on a 0–2 scale means "mostly blocking, some chance of soon". Round it if you need a level, or read the probabilities directly.
  • The Router chose a checkpoint for you. It detected English and used the English model. Remember that; it matters in Lesson 3.
  • output_tokens is zero. Nothing is generated, which is why output is free on Jev as well.

Each answer also carries two confidence numbers, and they are easy to mix up. answer_confidence is the probability of the answer being reported. That is the number calibration applies to, so it's the one to threshold on. confidence is an entropy-based measure of how peaked the whole distribution is, on a different scale. The library's own docstring warns against comparing the two against the same threshold. Gate your decisions on answer_confidence.

Lesson 2: ask small questions, then combine them in code

The most important habit, and the one TypeSafe's documentation stresses too, is to make each question something a knowledgeable person could answer in a few seconds. If a judgment needs several considerations, ask for each one separately and combine them in your own code.

Here are four emails, each asked about in two ways: one broad "is this phishing?" question, and four narrow signals that code combines into a weighted risk score.

from laya import Router

router = Router()

emails = {
    "invoice_scam": (
        "From: accounts@paypa1-billing.co\n"
        "Subject: Final notice - account suspended\n"
        "Your account will be closed in 24 hours. Verify your password and card "
        "number at http://paypa1-billing.co/verify to avoid permanent suspension."
    ),
    "real_invoice": (
        "From: billing@atlassian.com\n"
        "Subject: Your Jira invoice for September\n"
        "Hi Angsuman, your invoice #48213 for $240.00 is attached. It will be "
        "charged to the card on file on October 1. No action is needed."
    ),
    "ceo_fraud": (
        "From: ceo.office@gmail.com\n"
        "Subject: Urgent and confidential\n"
        "I'm in a board meeting and can't talk. I need you to buy five $200 gift "
        "cards right now and send me the codes. Don't tell anyone yet."
    ),
    "team_update": (
        "From: priya@ourcompany.com\n"
        "Subject: Sprint review moved\n"
        "Heads up, the sprint review is moving to Thursday 3pm because of the "
        "holiday. Same room, same agenda."
    ),
}

# One big, vague question...
monolithic = {
    "is_phishing": {
        "type": "noul",
        "instructions": "Is this email a phishing or fraud attempt?",
    }
}

# ...versus several small, concrete ones that code combines.
signals = {
    "urgency_pressure": {
        "type": "noul",
        "instructions": "Does the email pressure the reader to act immediately or face a penalty?",
    },
    "asks_for_secrets": {
        "type": "noul",
        "instructions": "Does the email ask for a password, card number, gift card codes or other secrets?",
    },
    "odd_sender": {
        "type": "noul",
        "instructions": "Is the sender address a free webmail account or a misspelled look-alike domain?",
    },
    "asks_for_secrecy": {
        "type": "noul",
        "instructions": "Does the email ask the reader to keep the request secret or skip normal process?",
    },
}

WEIGHTS = {"urgency_pressure": 1.0, "asks_for_secrets": 2.0, "odd_sender": 1.0, "asks_for_secrecy": 1.0}

CHECKPOINT = "multilingual"  # see "Pick the checkpoint by measurement" below

for name, email in emails.items():
    mono = router.predict(email, monolithic, model=CHECKPOINT)["answers"]["is_phishing"]["noul"]
    answers = router.predict(email, signals, model=CHECKPOINT)["answers"]
    risk = sum(WEIGHTS[k] * answers[k]["noul"] for k in signals) / sum(WEIGHTS.values())
    detail = "  ".join(f"{k}={answers[k]['noul']:.2f}" for k in signals)
    print(f"{name:13} monolithic={mono:.2f}  combined_risk={risk:.2f}\n    {detail}")

Output:

invoice_scam  monolithic=1.00  combined_risk=1.00
    urgency_pressure=0.99  asks_for_secrets=0.99  odd_sender=1.00  asks_for_secrecy=1.00
real_invoice  monolithic=0.00  combined_risk=0.00
    urgency_pressure=0.00  asks_for_secrets=0.00  odd_sender=0.00  asks_for_secrecy=0.01
ceo_fraud     monolithic=1.00  combined_risk=0.79
    urgency_pressure=0.28  asks_for_secrets=0.97  odd_sender=0.76  asks_for_secrecy=0.98
team_update   monolithic=0.00  combined_risk=0.01
    urgency_pressure=0.00  asks_for_secrets=0.00  odd_sender=0.01  asks_for_secrecy=0.02

Both approaches separate the scams from the legitimate mail here. So why bother decomposing?

  1. You can explain the decision. "Flagged because it asks for gift card codes and asks you to keep it secret" is something a compliance team or an end user can act on. "The model said 1.00" is not.
  2. You can tune without retraining. The weights are yours. If your business cares more about credential theft than urgency, change one number.
  3. You can see exactly where the model is weak. Look at ceo_fraud: "buy five gift cards right now" scored only 0.28 on urgency. With one monolithic question you would never have seen that. With signals, you know which question to reword or which examples to add when you fine-tune.

This isn't only a Laya trick. Runware reported that asking Jev "is this phishing?" directly scored 62.6% on their test, while splitting the task into five specific signals combined by code reached 95%.

Lesson 3: pick the checkpoint by measurement, not by assumption

In the code above, CHECKPOINT = "multilingual". That wasn't my first choice. The same script on the default English checkpoint, which the Router picks for English text, gave this:

invoice_scam  monolithic=0.66  combined_risk=0.51
    urgency_pressure=0.85  asks_for_secrets=0.25  odd_sender=0.95  asks_for_secrecy=0.26
real_invoice  monolithic=0.00  combined_risk=0.10
    urgency_pressure=0.22  asks_for_secrets=0.03  odd_sender=0.00  asks_for_secrecy=0.22
ceo_fraud     monolithic=0.65  combined_risk=0.78
    urgency_pressure=0.42  asks_for_secrets=0.76  odd_sender=0.97  asks_for_secrecy=0.98
team_update   monolithic=0.00  combined_risk=0.07
    urgency_pressure=0.04  asks_for_secrets=0.05  odd_sender=0.00  asks_for_secrecy=0.20

The English checkpoint rated an email asking for a "password and card number" at only 0.25 on asks_for_secrets. The multilingual checkpoint, on the same English email, gave 0.99.

Then I ran the opposite experiment: 24 hand-labelled support tickets across three departments (the code is in Lesson 5). This time the English checkpoint won. Here is everything I measured across Laya's three checkpoints:

Task (my tests)englishmultilingualtyped-decisions
Phishing signals, 4 emails (Lesson 2)Missed obvious signalsClean separation—
Ticket routing, 24 tickets, 3 departments0.960.880.96
34 banking intents, 6 queries, two-stage (Lesson 6)2/65/66/6

These test sets are tiny, so don't read the numbers as benchmarks. The pattern is still clear: the Router picks a checkpoint by language, not by how good it is at your task. It's a sensible default, not an optimization. Once you know your task, run your own labelled examples through each checkpoint and pin the winner with model=.

Lesson 4: typed schemas and confidence gating

Hand-writing question dictionaries gets tedious. Laya can build them from a Pydantic model: Literal fields become choice questions, bool fields become noul, and the field description becomes the instruction. That keeps your decision schema in one typed place that the rest of your code can import.

Pair that with the second key habit: a threshold for every action, set by the stakes, kept in code. Checking a balance at 50% confidence is fine. Closing an account is not.

from typing import Literal

from pydantic import BaseModel, Field

import laya
from laya import Router


class BankingIntent(BaseModel):
    action: Literal["check_balance", "transfer_money", "close_account", "other"] = Field(
        description="What does the customer want the assistant to do?"
    )
    frustrated: bool = Field(description="Is the customer frustrated or angry?")


# Thresholds live in code, next to the action they protect, not in a prompt.
STAKES = {
    "check_balance": 0.50,
    "transfer_money": 0.85,
    "close_account": 0.95,
    "other": 1.01,  # never auto-execute "other"
}

router = Router()

messages = [
    "what's my current balance on the savings account?",
    "send 500 dollars to my brother's account, same as last month",
    "I'm done with this bank. Close everything today.",
    "can you move some money around for me?",
]

for text in messages:
    result = laya.decide(router, text, schema=BankingIntent, return_details=True)
    action = result.values["action"]
    confidence = result.confidence["action"]
    if confidence >= STAKES[action]:
        route = f"execute {action}"
    else:
        route = f"confirm with user ({action}?)"
    print(f"{text[:48]:48} -> {route:32} conf={confidence:.2f} "
          f"frustrated={result.values['frustrated']}")

Output:

what's my current balance on the savings account -> execute check_balance            conf=0.80 frustrated=False
send 500 dollars to my brother's account, same a -> execute transfer_money           conf=0.95 frustrated=False
I'm done with this bank. Close everything today. -> execute close_account            conf=0.98 frustrated=True
can you move some money around for me?           -> confirm with user (transfer_money?) conf=0.41 frustrated=False

The vague request, "move some money around", came back as transfer_money with only 0.41 confidence, so it goes back to the user for confirmation instead of executing. This is the pattern that makes System One models safe in agents: the model proposes, and your code decides whether the proposal is strong enough to act on. If you'd rather not branch yourself, predict(..., min_confidence=0.85) marks weak answers with low_confidence: True.

One caution from independent testing of Jev: calibration can be uneven. A Towards Data Science analysis found answers reported at 0.7–0.9 confidence were right only 53% of the time, while answers at 1.00 were right 97% of the time. Don't assume a confidence number means what it says. Check it on your own data, which is the next lesson.

Lesson 5: evaluate and calibrate on your own data

If you take one thing from this guide, make it this: build a labelled evaluation set before you put a decision model in production. Twenty examples is enough to start; a few hundred is where it gets trustworthy. This script measures accuracy and Expected Calibration Error (ECE: how far stated confidence drifts from actual accuracy), then fits a temperature, a single number that rescales the probabilities so confidence better matches reality.

import numpy as np

import laya
from laya import Router

QUESTION = {
    "department": {
        "type": "choice",
        "instructions": "Which department should handle this ticket?",
        "criteria": {
            "billing": "invoices, payments, refunds, pricing",
            "technical": "bugs, outages, errors, integrations",
            "account": "login, password, users, permissions",
        },
    }
}

LABELED = [
    ("I was charged twice this month, please refund one.", "billing"),
    ("Can I get an invoice with our VAT number on it?", "billing"),
    ("Why did my plan price go up from $20 to $25?", "billing"),
    ("Our card expired, how do I update payment details?", "billing"),
    ("Do you offer a discount for annual billing?", "billing"),
    ("The receipt shows the wrong company name.", "billing"),
    ("I cancelled last week but was still billed.", "billing"),
    ("Please send me all invoices for 2025.", "billing"),
    ("The dashboard shows a 500 error since this morning.", "technical"),
    ("Webhooks stopped firing after your last release.", "technical"),
    ("CSV export produces an empty file.", "technical"),
    ("The mobile app crashes when I open settings.", "technical"),
    ("Your API returns 429 even at low traffic.", "technical"),
    ("Slack integration posts duplicate messages.", "technical"),
    ("Page takes 30 seconds to load reports.", "technical"),
    ("Search returns no results for existing records.", "technical"),
    ("I can't log in, it says my password is wrong.", "account"),
    ("How do I add a new teammate to our workspace?", "account"),
    ("Please remove John's access, he left the company.", "account"),
    ("I never received the password reset email.", "account"),
    ("Can I change the owner of our organization?", "account"),
    ("Two-factor codes from my authenticator are rejected.", "account"),
    ("How do I make Priya an admin?", "account"),
    ("My account got locked after too many attempts.", "account"),
]
LABELS = list(QUESTION["department"]["criteria"])

router = Router()


def evaluate(checkpoint):
    probs, gold = [], []
    for text, label in LABELED:
        answer = router.predict(text, QUESTION, model=checkpoint)["answers"]["department"]
        probs.append([answer["probabilities"][k] for k in LABELS])
        gold.append(LABELS.index(label))
    return np.array(probs), np.array(gold)


def report(name, probs, gold):
    correct = probs.argmax(1) == gold
    ece = laya.ece_score(probs.max(1), correct, bins=5)
    print(f"{name:28} accuracy={correct.mean():.2f}  ECE={ece:.3f}")


def rescale(probs, temperature):
    logits = np.log(np.clip(probs, 1e-9, 1)) / temperature
    e = np.exp(logits - logits.max(1, keepdims=True))
    return e / e.sum(1, keepdims=True)


for checkpoint in ["english", "multilingual", "typed-decisions"]:
    probs, gold = evaluate(checkpoint)
    report(checkpoint, probs, gold)

# Fit one temperature on even rows, check it on odd rows.
probs, gold = evaluate("multilingual")
fit, hold = slice(0, None, 2), slice(1, None, 2)


def nll(t):
    p = rescale(probs[fit], t)
    return -np.log(p[np.arange(len(p)), gold[fit]] + 1e-12).mean()


best_t = min(np.linspace(0.25, 5, 96), key=nll)
print(f"\nfitted temperature: {best_t:.2f}")
report("multilingual, held-out raw", probs[hold], gold[hold])
report("multilingual, held-out fit", rescale(probs[hold], best_t), gold[hold])

Output:

english                      accuracy=0.96  ECE=0.214
multilingual                 accuracy=0.88  ECE=0.051
typed-decisions              accuracy=0.96  ECE=0.288

fitted temperature: 1.75
multilingual, held-out raw   accuracy=0.83  ECE=0.141
multilingual, held-out fit   accuracy=0.83  ECE=0.064

Three takeaways:

  • Accuracy and calibration are different questions. The English checkpoint was more accurate here (0.96), but its confidence was further from reality (ECE 0.214). Multilingual was less accurate but more honest about it.
  • Temperature fitting helps, cheaply. On the held-out half, ECE fell from 0.141 to 0.064 after fitting one number. Laya's README reports the same effect at scale: ECE on the English checkpoint drops from 0.466 to 0.081 after temperature fitting. Fit one temperature per question type and option count, as the README advises.
  • Twelve held-out examples is far too few. I'm showing the method, not claiming these exact numbers. Use hundreds of real examples from your own traffic.

When zero-shot accuracy isn't enough, the next step is fine-tuning. Laya ships a notebook (notebooks/laya_finetune_typed_decisions_2xT4_kaggle.ipynb) that fine-tunes on Kaggle's free T4 GPUs in about four to five hours for ~30k questions, then fits calibration temperatures and pushes to Hugging Face. On their typed-decisions benchmark that takes accuracy from 0.362 zero-shot to 0.766. I covered the broader economics of fine-tuning in The Hidden Costs of Fine-Tuning Open Source LLMs. A 421M-parameter classifier is a far cheaper thing to fine-tune than an LLM, and it doesn't suffer the same knowledge-freshness problems, because it is judging content you hand it rather than recalling facts.

Lesson 6: many options? Build a tree

Laya's documented weak spot is a choice with many options. All options share a fixed token budget, so with dozens of labels each one gets only a few tokens to be distinguished. The published Banking77 result (77 intents) is 0.425 for Laya against 0.870 for Jev.

I tested this with 34 banking intents and six realistic customer messages. A flat 34-option question got 2 or 3 right out of 6, depending on the checkpoint. Shortlisting the options by embedding similarity first (laya.predict_shortlist) didn't help: it kept 2 of 6 and sometimes dropped the correct label before Laya ever saw it.

What worked was the oldest trick in classification: a two-level tree. First choose among five broad areas, then among a handful of intents within that area.

from laya import Router

TREE = {
    "card": {
        "description": "physical or virtual card: delivery, activation, PIN, loss, not working",
        "intents": {
            "card_arrival": "a new card has not arrived yet",
            "activate_card": "how to activate a card",
            "card_not_working": "card is declined or does not work",
            "lost_or_stolen_card": "card was lost or stolen",
            "change_pin": "change or unblock the PIN",
            "virtual_card_not_working": "virtual card problems",
        },
    },
    "money_movement": {
        "description": "sending, receiving or topping up money, and ATM withdrawals",
        "intents": {
            "transfer_pending": "a transfer has not arrived yet",
            "transfer_failed": "a transfer failed or was returned",
            "cancel_transfer": "cancel a transfer",
            "top_up_failed": "adding money to the account failed",
            "atm_wrong_amount": "ATM gave out the wrong amount of cash",
            "receiving_money": "how to receive money",
        },
    },
    "fees_and_rates": {
        "description": "fees, charges, and currency exchange rates",
        "intents": {
            "card_payment_fee": "extra charge on a card payment",
            "exchange_rate": "which exchange rate is used",
            "exchange_fee": "fee for exchanging currency",
            "transfer_fee": "fee for sending a transfer",
            "cash_withdrawal_fee": "fee for taking out cash",
        },
    },
    "refunds": {
        "description": "getting money back for purchases",
        "intents": {
            "request_refund": "ask for a refund from a merchant",
            "refund_not_showing": "a refund has not appeared yet",
        },
    },
    "account_and_wallets": {
        "description": "identity, personal details, closing the account, Apple Pay or Google Pay",
        "intents": {
            "verify_identity": "how to verify identity",
            "edit_personal_details": "change name, address or details",
            "terminate_account": "close the account",
            "apple_pay_or_google_pay": "use the card with Apple Pay or Google Pay",
        },
    },
}

router = Router()
CHECKPOINT = "typed-decisions"  # measured best for this task, see the table above


def classify(text):
    stage1 = {"category": {
        "type": "choice",
        "instructions": "Which area is the customer asking about?",
        "criteria": {name: node["description"] for name, node in TREE.items()},
    }}
    category = router.predict(text, stage1, model=CHECKPOINT)["answers"]["category"]["choice"]
    stage2 = {"intent": {
        "type": "choice",
        "instructions": "What exactly does the customer need?",
        "criteria": TREE[category]["intents"],
    }}
    return category, router.predict(text, stage2, model=CHECKPOINT)["answers"]["intent"]["choice"]


queries = {
    "I took 100 euros out at the ATM but the machine only gave me 60": "atm_wrong_amount",
    "My new card still hasn't shown up after two weeks": "card_arrival",
    "I sent money to my landlord yesterday and it's still not there": "transfer_pending",
    "Why was I charged extra when I paid in dollars abroad?": "card_payment_fee",
    "How do I get my money back for an order that never came?": "request_refund",
    "Can I use my card with my iPhone wallet?": "apple_pay_or_google_pay",
}
hits = 0
for query, gold in queries.items():
    category, intent = classify(query)
    hits += intent == gold
    print(f"{gold:24} -> {category:20} / {intent}")
print(f"\ncorrect: {hits}/{len(queries)}")

Output:

atm_wrong_amount         -> money_movement       / atm_wrong_amount
card_arrival             -> card                 / card_arrival
transfer_pending         -> money_movement       / transfer_pending
card_payment_fee         -> fees_and_rates       / card_payment_fee
request_refund           -> refunds              / request_refund
apple_pay_or_google_pay  -> account_and_wallets  / apple_pay_or_google_pay

correct: 6/6

Six out of six, from a checkpoint that got 2–3 of 6 on the flat version. Two things made the difference: each question now has at most six options, and the options carry real descriptions rather than bare labels. The cost is a second forward pass, around 30 ms on my machine.

If your label set is wide, flat, and you have no labelled data, this is exactly where Jev earns its fee: it handles up to 255 options without restructuring.

Lesson 7: batch everything, and measure

Speed is one of Laya's biggest selling points, so I measured it on my own machine rather than quoting the README:

import statistics
import time

import laya
from laya import Router

router = Router(preload=True)
ticket = "Hi, we were billed twice for March. Please refund the duplicate today."

one_q = {"refund": {"type": "noul", "instructions": "Is the customer asking for a refund?"}}
ten_q = {
    f"q{i}": {"type": "noul", "instructions": text}
    for i, text in enumerate([
        "Is the customer asking for a refund?",
        "Is the customer angry?",
        "Does the customer mention a specific month?",
        "Does the customer threaten to cancel?",
        "Is this about a bug in the product?",
        "Does the customer mention a competitor?",
        "Is this a sales inquiry?",
        "Does the customer ask for a phone call?",
        "Is the message written in English?",
        "Does the customer mention a duplicate charge?",
    ])
}


def p50_ms(fn, runs=50):
    for _ in range(5):  # warm-up
        fn()
    times = []
    for _ in range(runs):
        t = time.perf_counter()
        fn()
        times.append((time.perf_counter() - t) * 1000)
    return statistics.median(times)


print(f"1 question,   1 call : {p50_ms(lambda: router.predict(ticket, one_q)):6.1f} ms p50")
print(f"10 questions, 1 call : {p50_ms(lambda: router.predict(ticket, ten_q)):6.1f} ms p50")

batch = [{"state": f"{ticket} Order #{i}.", "questions": one_q} for i in range(512)]
router.predict_batch(batch[:32], batch_size=64)  # warm-up
t = time.perf_counter()
router.predict_batch(batch, batch_size=64)
elapsed = time.perf_counter() - t
print(f"512 tickets batched  : {elapsed:6.2f} s  ({512 / elapsed:,.0f} decisions/s)")

Output on an RTX 3090:

1 question,   1 call :   31.6 ms p50
10 questions, 1 call :   38.9 ms p50
512 tickets batched  :   0.72 s  (707 decisions/s)

Two numbers matter here:

  • Ten questions cost about 7 ms more than one. State is encoded once and the questions ride along. TypeSafe describes the same effect for Jev as speculative fan-out: ask every question you might need in one call, because the marginal cost is close to zero and answers never influence each other.
  • Over 700 decisions per second on one consumer GPU. At that rate a million decisions takes about 24 minutes of GPU time. (Latency varied by a few milliseconds between runs; across my runs single-question p50 sat between 27 and 32 ms.) On a rented T4, the Laya maintainers' reference card, it will be slower, but the per-decision cost is still a fraction of a cent.

For comparison, Jev's published end-to-end latency is 70–500 ms per call, with independent measurements around 240–280 ms p50 from the US West Coast. That's fine for back-office pipelines but noticeable inside an interactive request path.

Where these models shine: real use cases

With the mechanics covered, here is where I'd reach for a System One model, roughly in order of how much value I'd expect.

1. Routing requests to the right LLM. Most LLM traffic doesn't need your most expensive model. A decision model can grade every incoming request in a few milliseconds and send the easy ones to a small model. Laya ships a preset for exactly this:

import laya
from laya import Router

router = Router()
questions = laya.router_questions()  # difficulty, domain, needs_tools, is_sensitive


def pick_model(request: str) -> str:
    a = router.predict({"request": request}, questions)["answers"]
    difficulty = round(a["difficulty"]["score"])  # 0 trivial ... 3 hard
    if a["is_sensitive"]["noul"] > 0.5 or difficulty >= 3:
        return "claude-opus-5-5"
    if difficulty >= 2 or a["domain"]["choice"] in ("code", "math_or_logic"):
        return "claude-sonnet-5-5"
    return "claude-haiku-4-5"


requests = [
    "hi there!",
    "What's the capital of Australia?",
    "Refactor this 400-line Django view into a service layer and add tests.",
    "My landlord is withholding my deposit. What are my legal options in Karnataka?",
]
for r in requests:
    print(f"{pick_model(r):18} <- {r}")

Output:

claude-haiku-4-5   <- hi there!
claude-haiku-4-5   <- What's the capital of Australia?
claude-sonnet-5-5  <- Refactor this 400-line Django view into a service layer and add tests.
claude-opus-5-5    <- My landlord is withholding my deposit. What are my legal options in Karnataka?

Note what the code does and the model doesn't: the policy (sensitive or hard goes to the strongest model, code goes to the mid-tier) lives in plain Python where you can review it, test it and change it. A legal question about a deposit went to the strongest model because is_sensitive fired, not because anyone wrote a rule about Karnataka tenancy law.

2. Guardrails in front of an LLM. Checking for jailbreaks, prompt injection and sensitive data in every prompt is a perfect noul workload: high volume, latency-critical, and the questions never change. laya.guard_questions() gives you a starting set. Flowtivity's benchmark write-up reports Laya at 0.980 on phishing detection and 0.993 on Enron spam, though guardrail detection itself is more modest (about 0.76). Treat it as a fast first filter, not your only line of defence.

3. Ticket and message triage. Department, urgency, churn risk, sentiment, language: the classic help-desk problem, and the example both projects lead with. In CRM and clinic systems I've built, triage rules were historically brittle keyword lists that someone had to maintain forever. A calibrated decision model with a human-review queue for low-confidence cases is simply a better version of that system.

4. Multilingual intake. This is where Laya's Router earns its keep, and for anyone building for India it matters a lot:

from laya import Router

router = Router()
questions = {
    "department": {
        "type": "choice",
        "instructions": "Which department should handle this ticket?",
        "criteria": {
            "billing": "invoices, payments, refunds",
            "technical": "bugs, outages, system errors",
            "other": "everything else",
        },
    }
}

tickets = {
    "Hindi": "मेरे खाते से दो बार पैसे कट गए हैं, कृपया रिफंड करें।",
    "Bengali": "অ্যাপটি খুললেই ক্র্যাশ করছে, লগইন করতে পারছি না।",
    "German": "Die Rechnung für September ist doppelt abgebucht worden.",
    "Tamil": "செயலி திறக்கவே இல்லை, பிழை செய்தி வருகிறது.",
}

for lang, text in tickets.items():
    res = router.predict(text, questions)
    ans = res["answers"]["department"]
    print(f"{lang:8} checkpoint={res['routing']['model']:12} "
          f"-> {ans['choice']:9} ({ans['answer_confidence']:.2f})")

Output:

Hindi    checkpoint=multilingual -> billing   (1.00)
Bengali  checkpoint=multilingual -> technical (0.99)
German   checkpoint=multilingual -> billing   (0.70)
Tamil    checkpoint=multilingual -> technical (0.93)

Hindi, Bengali and Tamil tickets were routed correctly, with no translation step and no language-specific rules. The Router detected the script and switched checkpoints automatically. Laya's README reports 45 of 51 tested languages as usable on the multilingual checkpoint. Jev is currently English-first.

5. Agent control points. Autonomous agents make lots of small decisions: which tool to call, whether a result looks right, whether to ask the human. Each one is a natural System One question with a confidence gate. A recent arXiv paper on autonomous penetration testing uses exactly this structure: a System One model adjudicates findings, recalibrates severity, prunes which agents to run, and confirms results before they're reported, which cuts false positives and wasted compute.

6. Event-driven pipelines. If your architecture is built on events, a decision model is just another consumer: read an event, answer a few typed questions, emit a new event carrying the answers. At hundreds of decisions a second per GPU, it keeps up with real streams. I argued in How Event-Driven Architecture is Shaping the Next Generation of AI Agents that the event bus is the right backbone for AI systems, and typed decisions are one of the cleanest things to hang off it.

7. Regulated and air-gapped environments. Healthcare, legal, finance and government often can't send data to a third-party API at all. Laya runs entirely on your own hardware, so the data never leaves. For clinic and law-firm software this alone can decide the question.

8. Cheap evaluation at scale. "Is this answer grounded in the source?" or "does this reply follow our tone guide?", asked over thousands of LLM outputs, is an evaluation job that used to require an LLM-as-judge. LangChain has already added Jev to LangSmith Evals for this.

Where they don't belong

Knowing the limits is just as important. Don't use a System One model for:

  • Anything that must produce text. Drafting a reply, explaining a decision, summarizing a document. That's what LLMs are for. The winning pattern is usually the decision model routing and gating, with the LLM writing.
  • Arithmetic, counting and date comparison. TypeSafe documents these as known weak spots in Jev 1.13, and they're weak spots for classifiers generally. If the question is "is this invoice more than 30 days overdue?", compute the date difference in code and ask the model only what code can't answer.
  • Multi-step reasoning. If a person couldn't answer it in a few seconds, split it into questions that they could, or hand it to an LLM.
  • Long documents on Laya's defaults. The English checkpoint reads 512 tokens by default. Raise max_len (up to 8,192 on the multilingual checkpoint), split the document, or use Jev's 64k context.
  • Wide, flat label sets on zero-shot Laya. See Lesson 6.

And one thing no constrained output fixes: a confident wrong answer from your list is still wrong. Constraining outputs removes parsing failures, not judgment failures. Your confidence gates and review queues are there for that.

How much benefit can you actually expect?

Honest numbers, sourced:

  • Latency. Measured on my machine: about 30 ms for one question, under 40 ms for ten. A typical LLM call returning JSON takes hundreds of milliseconds to seconds, mostly spent generating tokens a decision model never produces. For anything in a user-facing request path, this is the difference between "noticeable" and "invisible".
  • Cost. Jev bills $0.042 per million input tokens with free output. At a generous 1,000 tokens per decision, a million decisions costs about $42. Laya is free to run; my measured 707 decisions per second means a million decisions in roughly 24 GPU-minutes. In one worked example, Towards Data Science found a single combined Jev call cost about $0.0005, versus about $0.006 for 13 separate LLM calls answering the same questions, a 12× difference. TypeSafe's own documentation reports 12.2× cheaper and 10× faster for batched questions.
  • Accuracy. Mixed, and task-specific. Jev beat an 80B-parameter open LLM by 4.7 points on banking classification (81.1% vs 76.4%) in the TDS test. Laya fine-tuned beats Jev on its own typed-decisions benchmark (0.766 vs 0.727) but loses badly on wide label sets without restructuring (0.425 vs 0.870 on Banking77). Zero-shot Laya on general checkpoints is near chance on hard tasks.
  • Engineering. The underrated benefit. No prompt-parsing code, no retry-on-malformed-JSON, and no model inventing categories. Typed answers with probabilities slot straight into ordinary control flow and ordinary unit tests.

If I had to put one number on it: for a system making high volumes of bounded decisions with an LLM today, moving those decisions to a System One model typically cuts their cost and latency by an order of magnitude, while accuracy holds or improves if you invest in question design and evaluation. If you skip that investment, you'll get fast, cheap, confidently wrong answers.

Choosing between them

Start with Jev if:

  • You have no labelled data and want good answers today.
  • Your documents are long (contracts, email threads) or your label sets are wide.
  • Your data can leave your environment and the latency budget allows a few hundred milliseconds.

Choose Laya if:

  • Data residency or privacy rules keep you on your own hardware.
  • You need decisions inside an interactive request path, in tens of milliseconds.
  • You're running high volumes where per-token pricing adds up.
  • You serve non-English users, especially across many languages.
  • You're willing to fine-tune, and want to own and inspect the model you depend on.

In practice my recommendation is both, behind one thin interface. The question format is nearly identical, so prototype on Jev for strong zero-shot answers, collect the real traffic and outcomes, and use that labelled data to fine-tune Laya for the high-volume paths. Keep your thresholds, question definitions and evaluation set in your own code. As Flowtivity put it in their benchmark write-up, a model lead is not a moat. What lasts is your data, your thresholds, your audit trail, and the workflow built around them.

A mastery checklist

  • Keep each question small enough that a knowledgeable person could answer it in seconds. Combine the answers in code.
  • Always include an "other" option when the input might not fit your categories.
  • Batch every question you might need into one call.
  • Gate actions on answer_confidence, with thresholds set by the stakes and kept in code.
  • Build a labelled evaluation set before production. Measure accuracy and calibration.
  • Compare checkpoints on your own task. Pin the winner.
  • Fit calibration temperatures before trusting confidence numbers.
  • Past ~20 options on Laya, build a tree.
  • Do arithmetic, dates and lookups in code, never in the model.
  • Let the LLM write and the decision model decide.

The bigger picture

For the last few years we've been using generative models as universal hammers, and a lot of production AI has looked like nails being hit with sledgehammers. System One models are one of the first clear signs that the field is specializing again: small, fast, calibrated components doing narrow jobs, orchestrated by ordinary, testable code.

That's a return to good software architecture more than a revolution. The systems that age well have always been the ones where each part does one job, the boundaries are typed, and the decisions are auditable. Jev and Laya make it much easier to build AI systems that way. The discipline to do it well — small questions, measured thresholds, honest evaluation — is still up to us.


Sources and further reading: Laya on GitHub · Laya on Hugging Face · TypeSafe AI documentation · Jev vs Laya comparison (Wilson Wu) · Laya benchmarked (Flowtivity) · Jev, Laya and decision models (Runware) · Jev vs LLMs (Towards Data Science) · A coding guide to Jev (MarkTechPost) · System One models for pentest agents (arXiv)