Skip to main content

Jev: a decision model you put in front of your LLMs to route traffic

· 7 min read
Rafael Fernandes
NLP Engineer & Tech Writer at WiLine
Share:
Routing · AI News

Jev decides, your LLMs answer

A request comes inThe right model, in 127 ms

TypeSafe shipped its first model on 15 September, and the interesting thing about Jev is what it refuses to do. It doesn't write you a paragraph. You give it a request and it hands back one structured value — a label, a class, a decision — in, they say, 70 to 500 ms. Founder Diogo Almeida's framing is the clearest line in the post: "Think of Jev as a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out."

Most models are built to talk to people. Jev is built to be called by code — and the first job that shape fits is routing.

A model built to decide, not to talk​

The name tells the story twice. "System One" is Kahneman's term for fast, automatic judgement — the snap decision — against "System Two," the slow deliberate reasoning we reach for a big LLM to do. And "Jev" is for William Stanley Jevons, the economist of the efficiency paradox: make something cheaper and people use far more of it. TypeSafe is betting cheap machine decisions get used everywhere.

A chat model generates tokens one at a time until a paragraph exists. Jev does the opposite — TypeSafe says it "gives up string generation," samples all outputs in a single parallel query, and returns one value from a set you define in advance. Two properties fall out, and both matter for routing:

  • The output is typed and structured, from a known set (SIMPLE, MEDIUM, COMPLEX, REASONING), not prose you have to parse. TypeSafe claims it "never makes type errors" — worth being precise here, because they are: that 0% is "not empirical. Schema matching is guaranteed, thus we can confidently add 0%." It's a property of constraining the output to a schema, not a measured result.
  • Each decision carries a calibrated probability — "higher confidence means higher accuracy," trained via a method they call Reinforcement Learning for Calibrated Decisions (RLCD). If that holds, code can act on the number: take the label when confident, escalate when not.

Their headline numbers ("40x-200x faster," output "too cheap to meter," a workflow eval at "193.6x faster, 444.6x cheaper") are vendor claims on their own evals, and to their credit they flag the bias — the evals were "made by individuals on our model capabilities team." Treat them as claims until someone independent re-runs them.

How routing works, and why the decider is the problem​

Model routing is a cost play. Not every request needs your biggest model, so you put a cheap model and an expensive one behind one endpoint and something in front decides which answers. Easy turns take the cheap tier; hard ones take the expensive tier; the bill drops. That "something in front" is a classifier:

Build that classifier the usual way — ask another LLM "is this simple or complex?" — and every request pays for an LLM call before it pays for the answer. That call is slow and costs tokens, on every request, whether it landed on the cheap tier or not. The decision meant to save money is quietly spending it.

Jev in the decider slot​

Jev fits that slot. You don't ask it to answer — you ask which tier the request belongs to, and the gateway routes on the label:

1Request arrivesat the gateway
2Jev labels it~127 ms, one tier
3Gateway routesto the matching model
4Model answerssmall or large
The classifier is step 2. Make it cheap and the arrangement pays off; make it an LLM and it taxes every call.

In LiteLLM's Auto Router that is a config change, not new code — you name the tiers, map each to a model, and set classifier_type: jev:

- model_name: jev-router
litellm_params:
model: auto_router/complexity_router
complexity_router_config:
tiers:
SIMPLE: {{openai_small}}
MEDIUM: {{openai_large}}
COMPLEX: {{anthropic}}
REASONING: {{anthropic_large}}
classifier_type: jev

What the swap is worth — and where the LLM falls down​

On 20 September LiteLLM's Moe Khalil published a benchmark titled "JEV Classifier: 5.43x as Fast as Haiku, 96% Lower Cost." He ran 80 cases three times each — 240 classification calls — with Jev (jev-1.13.0) against Claude Haiku 4.5 as the "classify with an LLM" baseline:

Classifier latency (p50)5.43× faster
Haiku 4.5 classifier688.40 ms
Jev classifier126.81 ms
Matched the expected tier+21 pts
Haiku 4.5 classifier73.75% (177/240)
Jev classifier95.00% (228/240)
Cost for 240 calls~96% cheaper
Haiku 4.5 classifier$0.1985
Jev classifier$0.0077

The average hides the real finding, which is in the per-tier numbers. On SIMPLE both were perfect (100%). But on MEDIUM Haiku collapsed to 38.33% where Jev held 85%, and on REASONING Haiku managed 76.67% against Jev's 100%. In other words, the LLM classifier is fine at telling trivial from non-trivial, and unreliable exactly in the middle band where routing decisions actually save or cost you money.

Two caveats, both LiteLLM's own, stated plainly: the expected tiers "were authored with the synthetic prompts, without independent review," and the benchmark measured the routing decision, not the quality of the final answer — "downstream answer quality was not measured." To their credit they published a frozen, hash-verified reproduction archive so the run can be re-analysed. Read the result as "Jev picked the tier fast, cheap, and the way they expected," not "your answers get 96% cheaper end to end."

Where this fits on WEC​

We build this exact setup in the tutorial on routing by complexity — a gateway in front of the WEC Inference API that scores each request and sends it to the right tier, Qwen3.5:9B for the simple turns and Qwen3.5:122B for the hard ones. And we have shown how a classifier gets it wrong on a bare "yes", sending work to the wrong model.

In both, the classifier is the part nobody prices. Put an LLM in that slot and you add two-thirds of a second and a token charge to every call before the real model is even chosen. A model like Jev is a bet that the decision in front of your models can be near-free — so routing pays off instead of taxing itself.

It is early, it is hosted, and the numbers still need someone independent to re-run them. For a WEC stack that keeps inference in-house, "hosted" is the real question mark — you would be sending every request's routing decision to a third party. But the shape is the part that travels: a model that decides in milliseconds, sitting in front of the models that answer.

Sources​

Comments & questions

Hit an error, spotted a typo, or have a question? Leave a note below.