dhruvkumar patel/ data scientist
Jev vs LAYA: which to choose?
Jev vs LAYA: which to choose?

Most production systems contain an LLM call whose whole output is one label: which queue gets this ticket, whether a prompt is an injection attempt, whether a retrieved passage is relevant enough to pass to the answering model. The usual implementation asks a chat model for JSON with a label and a confidence field, parses it, retries on malformed output, and routes anything above 0.8. The 0.8 was never checked against outcomes.

nibzard's decision-model benchmark (v2 report) puts numbers on what that costs. On a 77-way banking-intent task, thinking-mode models took up to 5.6 s p50 per decision (glm-5.3), and eight LLMs cost between $0.19 and $2.48 per thousand decisions. gpt-5.4-nano answered "none", which was outside the schema, on 6.2% of one suite. The result that matters most for thresholds came from shuffling the option order on otherwise identical items: gpt-5.4-mini changed its answer on 37% of them, and the Anthropic models on 30–34%. A confidence score on an answer that moves when you reorder the labels is not something you can put a threshold on.

Loading image: asset
Image
Two products launched within three days of each other in September 2026 aimed at exactly this call. Jev, from TypeSafe AI, is a hosted "System One model" that returns probabilities, never text, at $0.042 per million input tokens with output free. LAYA, from Nandakishor M, is an Apache-2.0 set of 322M–421M-parameter encoder checkpoints that speaks Jev's wire protocol, and its README opens with a chart showing it beating Jev. For any given one-word call, the question is whether to keep it on the LLM, move it to Jev, self-host LAYA, or train a classifier. This post goes further than the video version into LAYA's code and the benchmark details.

What a decision model returns

Both systems take POST /v1/systemone with a state (a string, a JSON object or an array of text) and a dictionary of named questions. A choice question returns the chosen option, a probability per option and a confidence. A score question places the state on an ordered rubric. A noul, TypeSafe's name for a yes/no question, returns one probability that a statement is true. The response can only contain options you declared, so an invented label or a parse failure can't happen: nibzard sent 4,125 requests to Jev and got zero schema violations. All 225 failures were rejections at the option cap, which returns 400 Too many choices. Must have at most 255 choices.

Loading image: asset
Image
LAYA's laya.serve implements the same endpoint and the same response schema, down to the usage block, so a Jev client only needs its base URL changed. That is the most practical thing LAYA ships. You can put a hosted and a self-hosted decision model behind one flag and A/B them without touching the caller.

There are three ways to build something that returns this contract. The cheapest keeps a generative model and stops sampling from it: put the options in the prompt and read the next-token logits for each option. The openjev project did this with a frozen Qwen3.5-4B on one RTX 3090, and 21 yes/no criteria over one state took 1.023 s through the logits against 5.332 s when the model generated the same answers as a JSON array (111 output tokens). That fixes speed and schema validity, but not calibration, because the logits still come from a model post-trained to produce preferred text. The second way is the older NLP answer: a bidirectional encoder that reads the input and each candidate label together and scores the label. NLI zero-shot classifiers and cross-encoder rerankers work like this, and so does LAYA. The third way is whatever Jev is. TypeSafe hasn't published the architecture, and its docs describe only the behavior: the state is ingested once, and "every question is evaluated in parallel and in isolation against the same state in one go".

The first reaction many engineers had on launch day was that this is just classification, and it is mostly right. TypeSafe describes its post-training, which it calls RLCD, as optimizing for probabilities that match outcome frequencies. In statistics, that is optimizing a strictly proper scoring rule: a reward whose expected value is maximized only by reporting your true belief. Accuracy is not proper. If the right answer is A with probability 0.7, the expected accuracy reward is maximized by reporting A with probability 1.0. The log score is maximized at exactly 0.7, and the expected log score is negative cross-entropy. Every softmax classifier trained with cross-entropy is already trained on a strictly proper scoring rule. What a decision model has to show is that it delivers a runtime label space, probabilities you can threshold, and accuracy close to the LLM it replaces, all in one call.

How LAYA scores options, and where its budgets bite

LAYA's design is readable in laya/common.py. Each question becomes one sequence:

text
[CLS] <type> instructions [SEP] [MASK] option0 [MASK] option1 ... [SEP] state [SEP]

The sequence runs through ModernBERT (421M parameters in the English laya checkpoint, with a 512-token context), plus a learned question-type embedding and two more encoder layers. The model reads one scalar off each option's [MASK] position, divides by a temperature chosen by question type and option count, and takes a softmax across the options. The label space lives in the input, so it can change on every request.

Loading image: asset
Image

Agent._encode_state builds one such sequence per question, each with the full state appended. Five questions about one email means the email is encoded five times. LAYA's own T4 latency table shows the cost: the English checkpoint takes 39.5 ms for one question, 84.5 ms for five, 158.6 ms for ten and 771 ms for fifty. Jev prices and behaves like the other design, where the state is encoded once and questions are cheap. TypeSafe's parallel-questions cookbook measured 13 questions over the GDPR Wikipedia article at $0.000497 and 0.27 s as one call, against $0.006090 and 2.71 s as 13 separate calls (the cookbook was written against jev-1.12). Retrieval engineers know this trade from rerankers. A cross-encoder lets every option token attend to every state token in every layer, and pays for it by recomputing the state per question. How Jev keeps question-to-state interaction deep while encoding the state once is not public.

The budgets matter more than the latency for most calls. build_sequence reserves a head budget for the instruction and options (head_max_len, 192 tokens on laya and 256 on the other two checkpoints), and when the options overflow it, every option is cut to the same length:

python
    opt_budget = head_max_len - sum(len(o) for o in opt_ids)
    if opt_budget < 16:
        per = max(4, (head_max_len - 16) // max(1, len(opt_ids)))
        opt_ids = [o[:per] for o in opt_ids]

That length includes the option's [MASK] token. At 77 options (Banking77) each option keeps three tokens of label text, and "card arrival" and "card delivery estimate" become hard to tell apart. Both base checkpoints score exactly 0.425 on Banking77. Two different encoders landing on the same number points to the budget as the ceiling, and the README says as much. Its advice is to keep choice questions under about 20 options, raise head_max_len, or shortlist with predict_shortlist, which keeps the top k labels by embedding similarity before the forward pass. Issue #102 reports a top-20 shortlist moving one Banking77 run from 54.3% to 60.8%; the repo hasn't remeasured it.

The state side is quieter and matters more. On the English checkpoint the state gets about 320 tokens, and anything past that is cut from the right with no error or warning (lists are cut from the left, so the newest turn of a conversation survives). A long support email is decided on its first ~320 tokens. Jev's documented state budget is 32k tokens.

The confidence field is computed, and the formulas differ

Neither system has a separate confidence head. TypeSafe's docs call confidence "a statistic computed from the probability distribution", and the interactive example on the docs page uses (K × p_max − 1)/(K − 1), labeled as an approximation. LAYA computes one minus normalized entropy, 1 − H(p)/log K, for choice and score questions, and max(p, 1 − p) for nouls. For three options at probabilities (0.6, 0.3, 0.1), Jev's approximation gives 0.40 and LAYA's formula gives about 0.18. A threshold tuned on one model means something else on the other, and on LAYA it also shifts with the number of options. Tune thresholds per model, per question type and per option count, on your own data.

Jev's nouls have their own trap. Its jaggedness page for jev-1.13 gives the example itself: for "I was charged twice for the same order", the questions "Is the customer asking for a refund?" and "Is the customer asking for something other than a refund?" return 0.72 and 0.47, which sum to 1.19. The page's advice is to enforce identities in code. Two nouls are not a probability distribution, even when they read like one.

What calibration buys

AbdelStark's preregistered pilot measured the thing a calibrated model is for: coverage at a fixed error budget, the share of items you can auto-accept while keeping the error rate on those items under a target. With 100 held-out examples per condition and jev-1.13.0, at 5% error or less, Jev could auto-accept 83% of AG News items and 86% of Banking77 items (the 72-label BTZSC variant). GLiNER2.5, run locally, managed 24% and 27%. The accuracy gaps between the two models were 21 and 26 points; the coverage gap was about 60. That multiplier is the argument for this whole product category. When the probabilities track outcomes, a threshold turns directly into automation you can budget.

Loading image: asset
Image

The same pilot shows the limit. On DAIR Emotion, Jev scored 0.480 accuracy, coverage at 5% error was zero, and Jev put exactly zero probability on the true label for 16% of examples (NLL 5.588). Whether that is rounding in the API output or a real collapse hasn't been established. With 100 examples per condition, treat all of these as a pilot.

What LAYA's launch chart compares

Most of the launch-week comparisons put two numbers side by side that were measured differently:

ClaimLAYA sideJev sideWhat differs
Accuracy 0.766 vs 0.727 (typed-decisions)fine-tuned on that dataset's train splitzero-shottraining
ECE 0.081 vs 0.246after a temperature refit on held-out in-domain dataraw, on nibzard's no-correct-answer suitedata, and the definition of confidence
Latency 32.8 ms vs 236–276 msforward pass on a local T4HTTPS round trip from France and the USnetwork
AG News 0.950 vs 0.910AG News is in LAYA's training mix100-example zero-shot pilottraining data

Where both are measured raw on typed-decisions, Jev has the better ECE (0.144 against 0.213) and the better soft accuracy against the teacher distributions (0.580 against 0.471). The base LAYA checkpoints score 0.362 and 0.342 there, below the 0.461 per-question majority baseline. The fine-tuned 0.766 sits above the dataset's 0.735 teacher self-agreement ceiling, which Luni's independent evaluation reads as fitting labeler noise. I couldn't find where Jev's 0.727 on this dataset was first published, so treat it as the chart's own number. The ECE row needs one more correction: 0.246 comes from nibzard's suite S5, whose items have no correct option. On S1, 77-way banking where answers exist, nibzard measured Jev's ECE at 0.083, about the same as LAYA's refit 0.081.

Loading image: asset
Image

LAYA's own README is more candid than its chart: "Laya is a fast base to specialise, not a zero-shot decision engine."

Where calibration breaks

Calibration is a property of a model on a distribution. Each of the following is a case where the distribution changed and the probabilities didn't notice.

Items with no right answer

nibzard's suite S5 asks questions where no option is correct and counts how often each model reports confidence at or below 0.5. Seven of the eight LLMs did so on 97.3–100% of items; gpt-5.4-mini did on 64.7%. Jev did on 49.7%, with a mean confidence of 0.54. (The video version of this post said every LLM; gpt-5.4-mini is the exception.) The benchmark's README, rewritten on 26 September, warns that LLM confidence is a prompted probability while Jev's is a provider-defined score, so treat the comparison as rough. The practical fix doesn't depend on it: give Jev an explicit "none of these" option, because it will otherwise pick the least wrong answer with moderate confidence.

Inputs unlike the training data

On MASSIVE intent (20 options, so random is 0.05), LAYA's English checkpoint scores 0.000 on Khmer at a mean confidence of 0.952. Across 51 languages its mean confidence never drops below 0.885, whatever the accuracy. No threshold catches that, because the model is wrong for a reason it can't see. LAYA's answer is to decide before the forward pass: laya/lang.py checks the Unicode script in under 0.5 ms, applies a function-word heuristic to Latin text, and routes anything not English to the multilingual checkpoint. Routed, 45 of 51 languages clear three times random, against 23 of 51 for the English checkpoint alone. The lesson reaches past languages. Any confidence-gated cascade needs an explicit out-of-distribution check in front of the gate.

Loading image: asset
Image

A calibration layer fitted on the wrong data

This comment in laya/common.py is the clearest single example:

python
# A fitted temperature below 1 sharpens the logits instead of softening them. The shipped
# `choice:11+` bucket is 0.1006, which multiplies them ~10x: a 0.24 top probability is published as
# 0.99, so a caller gating on confidence is told a coin flip is a certainty. No honest calibration
# needs to sharpen this hard, so refuse to apply one that does.
TEMP_MIN = 0.5
TEMP_MAX = 5.0

The 0.3.10 release notes say calibration used to be fitted on training items (issue #186). A temperature fitted on items the model has memorized gets pushed toward sharpening, and 0.1006 sits at the fitter's lower bound of 0.1; that this is how the shipped value arose is my inference, not something the repo states. After the clamp, re-running the 51-language sweep moved macro ECE from 0.733 to 0.571, with accuracy unchanged, since temperature never changes the argmax.

A threshold in the wrong place

Luni ran raw LAYA on 2,000 balanced emails from the PhishNChips set. It scored 0.505 accuracy with 1.2% recall, calling almost everything "not phishing", yet its AUROC was 0.678, against 0.689 for Jev as Luni reports it (the page doesn't name where Jev's figures come from). The ranking was about as good; the threshold was wrong. Temperature scaling can't fix that, because it has no bias term and can never move a prediction across 0.5. Platt scaling (a fitted slope and bias), trained on 1,000 of the emails and scored on the other 1,000, lifted accuracy to 0.611, against 0.626 for Jev. Claude Haiku 4.5 scored 0.813 on the same set, about 20 points ahead of both decision models. Anyone who has shipped a CTR model will recognize this: ranking and calibration are separate problems, and a calibration layer with a bias term is standard infrastructure.

What LAYA's training code shows

LAYA's author uses TypeSafe's name, RLCD, for his own method. TypeSafe has described RLCD only by its goal, so there is no basis for calling them the same algorithm. LAYA's reward, proper_reward, combines strictly proper rules: the log score (floored at log 1e−4), a weighted spherical score, and a ranked probability score on ordinal questions, which is the one part plain cross-entropy lacks, since it counts a guess one rubric level off as better than one three levels off. The optimizer adds Gaussian noise to the logits several times per question, scores each perturbed distribution, and weights each perturbation by its advantage within the group. That is a Monte Carlo estimate of the gradient of a smoothed proper score, a higher-variance route to roughly where minimizing the negative score directly would go.

The Dev.to launch post describes base training as "pure policy gradient (zero supervised cross-entropy loss)". The base training script isn't in the repo, so that can't be checked. The fine-tuning notebook that produced the typed-decisions checkpoint behind the "beats Jev" row is in the repo, and its loss line is:

python
loss = (loss_rl + 1.0 * loss_ce) / GRAD_ACCUM + 0.0 * act.sum()

loss_ce is soft cross-entropy against the teacher LLMs' probability distributions, at full weight. The checkpoint is a distillation of the teachers with a noisy regularizer on top. The 0.0 * act.sum() term means the act head, the component meant to decide between acting and escalating, gets no gradient in fine-tuning. The README says it "carries no usable signal yet": its raw logits rank correctness at AUROC 0.30 on 396 labeled decisions, worse than chance, against 0.77 for the plain confidence field. The calibration in the shipped models comes from the post-hoc fit. The base checkpoints ship with mean ECE 0.466 (laya) and 0.314 (laya-multilingual), and a temperature fit on held-out data takes them to 0.081 and 0.106. That fit minimizes negative log-likelihood, which is cross-entropy again. Any classifier can have the same layer.

The option the chart left out

dylantom2012 ran 10,000 decisions over four public datasets on an M5 Pro CPU. Hosted Jev reached 79.3% macro accuracy at 381 ms p50, almost all of it network. A ModernBERT-base encoder with a linear head, trained on about 2,000 labels per task, reached 82.6% at 15.8 ms. A potion-base-8M static-embedding model with a linear head reached 78.2% at 0.1 ms in 82 MB of RAM. With no labels at all, a zero-shot ModernBERT-base cross-encoder (149M parameters) reached 78.7%, 0.6 points behind Jev. The change that moved accuracy most in that run, by 12–24 points, was architectural: scoring the text and the option jointly instead of embedding them separately. It's a single seed on one machine, and Jev's Banking77 accuracy varies across sources (0.870 in AbdelStark's 72-label pilot, 76.3% in nibzard's full 77-way suite, 77.8% in dylantom's run), so read the gap as "comparable" more than "better". But a fixed-label task with a couple of thousand examples doesn't need either product.

Loading image: asset
Image
Which one I'd pick

Before changing the model, split the decision into small questions. TypeSafe's workflow evals score agreement with a reference built from two frontier models, on workflows TypeSafe wrote, so they aren't neutral; but the decomposition result holds for every model in them. Claude Haiku 4.5 goes from 18.1% as one big prompt to 53.6% as workflow questions, Luna from 51.9% to 66.8%, and Opus 5 from 64.8% to 73.1%. Atomic questions with the weights and logic in code is good design whichever model answers them.

Loading image: asset
Image

Keep the LLM when the decision needs several hops of reasoning, arithmetic or date comparison, all of which TypeSafe's jaggedness page lists as known weaknesses, or when a wrong answer is expensive and volume is low. Invoice processing is the workflow where Jev trails Opus 5 most (61.8% against 78.4%), and on phishing Haiku 4.5 was about 20 points ahead of both decision models.

Train a small classifier when the labels are fixed and you can get about 2,000 examples per task. It was the most accurate option in the one benchmark that included it, and it runs in milliseconds on a CPU.

Use Jev when the labels change per request or you have none, the text is English, there are 255 options or fewer, and volume or latency is the problem. Against a well-chosen fast LLM the gains are real but modest: 2.7–4.6 times cheaper per decision than the cheapest LLMs nibzard measured, 1.2 times faster than gpt-oss-120b on Cerebras (331 ms p50), and the most stable answers under option reordering (13% flips). The 40–200× speedups in the launch post hold against thinking-mode models. Add a "none of these" option, keep nouls apart, and pin jev-1.13.0, because the jev-latest alias moves on each release and a new version shifts the probability scale your thresholds were fitted on.

Self-host LAYA when the data can't leave your infrastructure, you need latency in the tens of milliseconds, and you're willing to fine-tune. Plan for 4–5 hours of fine-tuning on Kaggle's free 2×T4 for 4 epochs over about 30,000 questions; Platt or isotonic calibration on held-out data; fewer than 20 options, or a shortlist; states under ~320 tokens on the English checkpoint, or the 1,024-token checkpoints; and the router in front. On CPU, thread settings dominate: one contributed laptop run (Ryzen 9 6900HX under WSL2) went from 9,396 ms p50 to 783 ms after torch.set_num_threads(8) and torch.set_num_interop_threads(1). The repo's browser-agent example shows the ceiling after fine-tuning: element top-1 accuracy among about 45 candidates went from 0.10 to 0.66, and real-task success from 0% to 62%, at 17–23 ms per step on one 16 GB GPU.

The model is the cheap part of all four. At list price a Jev decision over a 500-token state costs 500 × $0.042 / 10⁶ = $0.000021, so a million decisions cost about $21 in input tokens; nibzard's measured $0.07 per thousand banking decisions implies about 1,670 input tokens each, most of them the 77 options. LAYA's multilingual checkpoint serves 332 questions per second batched on one T4, about 0.84 T4-hours per million single-question decisions. What costs money is the work around the model: labels, a calibration fit on held-out data, a "none of these" option, and a check for inputs the model never saw.

What is still open

Nobody has published what Jev is: no architecture, parameter count, base model or training data. Flat latency across 2 to 255 options, the state-once pricing and the hard cap at 255 narrow the design space without settling it. Whether LAYA's policy-gradient term adds anything beyond the full-weight cross-entropy term needs an ablation, and none is published. The DAIR Emotion zero-probability result needs the raw API values to tell rounding from collapse.

The comparison that would settle the practical question also doesn't exist yet: one realistic LLM-as-classifier task run through all four paths (the current LLM, Jev, LAYA with Platt calibration, and a ModernBERT head on the available labels), reporting accuracy, coverage at 5% error, p50 latency and cost per thousand decisions. I haven't run it. Every number in this post comes from someone else's benchmark or from the projects' own code and docs.

Sources

TopicsLLM ClassificationCalibrationTypeSafe JEVLAYAModernBERTZero Short ClassificationText Classification
Citation

Cite this essay

If you reference or build upon this analysis in research, technical reports, or blog posts, please cite this work:

Standard (APA)

Patel, D. (2026). Jev vs LAYA: Which Should Replace Your LLM Call?. Dhruvkumar Patel's Engineering & Research Blog. https://www.stackdhruv.com/blog/jev-vs-laya-llm-classifier

BibTeX
@article{patel2026jevvslayallmclas,
  author    = {Dhruvkumar Patel},
  title     = {Jev vs LAYA: Which Should Replace Your LLM Call?},
  journal   = {Dhruvkumar Patel's Engineering & Research Blog},
  year      = {2026},
  url       = {https://www.stackdhruv.com/blog/jev-vs-laya-llm-classifier}
}