AI Routing

Decision Models: Ask for a Verdict, Not an Essay

Routing, triage and guardrail calls rarely need a paragraph, only a label. Decision models return that label as a typed value with a probability. Set two thresholds on your own labels, let the verdict route work, and never let it approve anything irreversible.

Cloud X Ops TeamDevOps & AI Integration
September 26, 2026
12 min read

TypeSafe opened early access to Jev, a model that answers with a label and a probability instead of prose, on 15 September. Ten days on you can run open weights yourself, Convai Innovations' Laya or Jared Palmer's Kev, and the Ollaya runtime serves them in Jev's request shape. The first critiques are in too: Alex Molas on calibration, and a Check Point study in which forged documents flipped Jev's verdict in 59% of attacks.

Plenty of production LLM calls only pick a label: which team gets a ticket, whether a reply leaks a card number. The probability that comes back with the label is the useful part. Code can act on it, if it means what it says.

What a decision model returns

You send a state (a ticket, a document, a tool call) and typed questions about it. TypeSafe's API has three: a Choice picks one of up to 255 options, a Score rates against a rubric, and a Noul answers true or false as a number from 0 to 1. What you can run today:

  • Jev (hosted). $0.042 per million input tokens, output free, 70 to 500 ms end to end by TypeSafe's figures, early access, no published weights. TypeSafe says its speed figures come from evals "generally run from our laptops on the West Coast", where the service is based. Of the price it says: "We can't prove it isn't subsidized".
  • Laya (open). Apache 2.0 encoders on ModernBERT-large and mmBERT-base, quoted at 32.8 ms a question on a T4 GPU for the multilingual checkpoint, 39.5 ms for English, and 193 to 464 ms on CPU. The English checkpoint leaves about 320 of its 512 tokens for the text, and the card calls the base checkpoints "a fast base to specialise, not a zero-shot decision engine".
  • Kev (open). LoRA and a pointer head on Qwen3.5 up to 9B and Qwen3.8 at 27B. Its README puts Kev-27B at 0.848 accuracy to Jev's 0.857 on unseen sources, adding "this isn't a controlled comparison". The 27B needs an 80 GB GPU.
  • Ollaya (runtime). Beta, Apache 2.0, serving /v1/systemone in TypeSafe's request shape, so TypeSafe's Python SDK works unchanged. It carries Laya and the smaller Kev models, not the 27B. Every model runs on CPU; its own latency figures are from an RTX 4090.

Jev and Ollaya share one API, so the code below runs on either. The fast figures are GPU numbers, so measure yours on the hardware you will run.

Use one wherever the answer is already a label

If your code string-matches a model's reply (if "refund" in reply), that call wanted a type. The usual places:

  • Routing. Intent to queue, ticket to team, task to model size.
  • Triage. Priority as a Score, "needs a human now" as a Noul.
  • Gates. Personal data in an outbound reply, a policy breach in user input.
  • Tool call screening. A Noul on "this call is destructive", fed to policy code as one signal.
  • Eval judges. Pass or fail against a rubric, read as a number instead of parsed from prose.

It cannot write: drafting, summarising and multi-step reasoning stay on an LLM, and every option is defined up front. Quality is still open. As Mo Bitar put it in a video The Register quoted, "I know it's fast. I know it's cheap, but is it good?"

Watch a score become a decision

A thousand synthetic tickets, built so a verdict at 0.9 is right nine times in ten, fall into bins by top probability as the panel below comes into view. Press Route tickets and two handles drop in at the thresholds in route.py further down: 0.92 and up goes to a team queue, 0.60 to 0.92 to a person, lower to the old LLM route. Drag them and watch automation, shipped errors and review load move.

Threshold Bands · 1,000 Synthetic Tickets
laya/v1/systemone
0scored
LLM route
0
model abstains
p < 0.60
Person
0
0 wrong, caught
0.60 <= p < 0.92
Queue
0
0.0% automated
p >= 0.92
Unreviewed
0wrong
0.0% of automated
5% error budget
$ awaiting tickets
Synthetic tickets, calibrated by construction. The band between the thresholds is where people buy accuracy back.

Two thresholds, set on your own labels

Molas points out that calibration is a property of your traffic as well as the model. He also found posts where Jev put 0.92 on a fair coin landing heads, enough to clear the 0.92 AUTO threshold below, and calls that worse than drift, since the true probability was in the prompt. A few hundred of your own labels, he says, can be enough to fit Platt scaling on top: a small logistic fit that maps the model's probability to how often it was right on your data.

Kev's README adds a second test, how well a model ranks right answers above wrong ones: at a 5% error budget, it says, Kev can automate 45 to 57% of decisions and Jev 70%. Both sets of numbers come from someone else's data, so shadow first. Call both routes, act on the old one, log label and probability against the outcome, then set two numbers.

  • AUTO: the lowest probability where errors above it stay inside your budget; those verdicts act alone.
  • REVIEW: the floor of the band a person checks. Below it, the old route decides.
route.py
# Intent routing on a decision model. Ollaya answers TypeSafe's API, so the
# same body goes to api.typesafe.ai with a Bearer key and model "jev-latest".
import requests

URL = "http://localhost:11435/v1/systemone"
AUTO, REVIEW = 0.92, 0.60    # examples: fit yours on shadow run labels

def route(ticket: str) -> tuple[str, str]:
    body = {"model": "laya", "state": ticket, "questions": {"intent": {
        "type": "choice",
        "instructions": "What does the customer want?",
        "criteria": {"invoice": "Needs an invoice or receipt",
                     "refund": "Wants money back",
                     "other": "Anything else"}}}}
    try:
        r = requests.post(URL, json=body, timeout=2)
        r.raise_for_status()
        ans = r.json()["answers"]["intent"]
        label = ans["choice"]
        p = float(ans["probabilities"][label])
    except (requests.RequestException, KeyError, TypeError, ValueError):
        return "other", "llm"            # FAIL BACK: the old LLM route decides
    # CALIBRATE: p = platt(p), fitted on shadow run labels
    # LOG: label, p and outcome, every call
    if p >= AUTO:
        return label, "queue"            # ACT: straight to that team's queue
    if p >= REVIEW:
        return label, "human"            # REVIEW BAND: a person confirms
    return "other", "llm"                # ABSTAIN: too unsure to route

Two traps. The response also carries a confidence field, "derived from the answer's probability distribution" per TypeSafe's docs; Ollaya's example shows 0.9547 against a top probability of 0.9698. Threshold on the one you calibrated. And mind length: Kev trained on states of up to 384 tokens, Laya's English checkpoint reads about 320, and the smaller Kev models lose accuracy on long documents, though the 27B holds up much better. Once live, alert on score drift.

A verdict is not a security boundary

Typed output feels safe because the model cannot answer off the list. Check Point published a test of that on 24 September, titled “A Decision Model Breaks Like Any Other Language Model”, run against a due diligence assistant that returns a risk level and an investment verdict. The attacks that worked never told the model what to output. They appended a plausible addendum instead: a clean audit opinion, a regulatory file number, a revised risk table. Of the attacks on Jev, 59% succeeded, after 4.6 turns on average, at $0.54 per break. The obvious defences did not help. Marking the document untrusted still let 16 of 27 runs break, the same count as pasting it inline with the instructions, and an instruction to ignore embedded instructions changed the count by one run in 27, which Check Point called noise.

Check Point ran the same attacks on two unnamed low-cost chat models. With reasoning off they broke more often than Jev (100% and 67%), and reasoning, the best defence it measured, is a setting Jev lacks. Check Point calls its work a directional study, not a benchmark, with an attacker who could see the model's probabilities, and TypeSafe's own docs list adversarial content among Jev's known weak spots. As the study puts it, "The output format constrains the shape of the answer. It does nothing about the content of the input."

Forged context still moves the verdict

Put a rule in code behind every verdict that can cost money. A refund above the order total, a payment to a new payee or a deploy outside the change window should fail there whatever the label says.

Let a decision model say which queue a ticket belongs in. Never let it be the reason money moves.

Takeaways

  • If your code string-matches a model's reply, that call wanted a type: routing, triage, gates and eval judges are the usual places.
  • Jev is hosted and in early access; Laya and the smaller Kev models run on your own hardware behind the same API through Ollaya, so measure latency there rather than trust the published GPU figures.
  • Fit the calibration on a few hundred of your own labels from a shadow run before you trust a probability.
  • Set AUTO at the lowest probability whose errors stay inside your budget, send the band down to REVIEW to a person, and let the old route decide below it.
  • Check Point's attackers flipped Jev's verdict in 59% of attempts with forged addenda, so a verdict can route work but should never approve an irreversible action alone.

Paying a chat model to pick a label?

We find the calls in your pipeline that only need a verdict, shadow a decision model against them, calibrate the thresholds on your own data and keep every irreversible action behind code.

Review your routing
route.sh
SECURE
cloudxops@router:~$ ./shadow.sh intent-routing
# Comparing verdicts with the current LLM route...
[OK] calibration fitted on your labelled rows
[INFO] 0.60 to 0.92 goes to a person
[READY] cut over on a green shadow run
$ ▋
Routing confidence