Skip to main content

2 posts tagged with "routing"

View all tags

Jev: a decision model you put in front of your LLMs to route traffic

· 7 min read
Rafael Fernandes
NLP Engineer & Tech Writer at WiLine
Share:
Routing · AI News

Jev decides, your LLMs answer

A request comes inThe right model, in 127 ms

TypeSafe shipped its first model on 15 September, and the interesting thing about Jev is what it refuses to do. It doesn't write you a paragraph. You give it a request and it hands back one structured value — a label, a class, a decision — in, they say, 70 to 500 ms. Founder Diogo Almeida's framing is the clearest line in the post: "Think of Jev as a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out."

Most models are built to talk to people. Jev is built to be called by code — and the first job that shape fits is routing.

A router that can't see the conversation can't classify "yes"

· 7 min read
Rafael Fernandes
NLP Engineer & Tech Writer at WiLine
Share:
Routing · AI News

Classifying the word "yes"

Score this messageScore what it approves

A model router's job is to read a request and decide which model should answer it. Cheap questions go to a small model, hard ones to a large one, and the bill comes down. The whole arrangement rests on being able to tell the difference.

Then a user types "yes".

Or "continue". Or "do it". Nothing in those two or three characters says whether the work being approved is a spelling fix or a database migration. A router scoring the current message in isolation sees a very short string with no technical vocabulary, and does the obvious thing: cheapest model.

Which means if you route requests to save money, your cheapest tier is probably absorbing work it should never have seen — and your savings figure is partly fake. On 4 August LiteLLM published a benchmark that measures both halves of that: how wrong the routing gets, and what it costs to fix.