Skip to main content

3 posts tagged with "evals"

View all tags

A router that can't see the conversation can't classify "yes"

· 7 min read
Rafael Fernandes
NLP Engineer & Tech Writer at WiLine
Share:
Routing · AI News

Classifying the word "yes"

Score this messageScore what it approves

A model router's job is to read a request and decide which model should answer it. Cheap questions go to a small model, hard ones to a large one, and the bill comes down. The whole arrangement rests on being able to tell the difference.

Then a user types "yes".

Or "continue". Or "do it". Nothing in those two or three characters says whether the work being approved is a spelling fix or a database migration. A router scoring the current message in isolation sees a very short string with no technical vocabulary, and does the obvious thing: cheapest model.

Which means if you route requests to save money, your cheapest tier is probably absorbing work it should never have seen — and your savings figure is partly fake. On 4 August LiteLLM published a benchmark that measures both halves of that: how wrong the routing gets, and what it costs to fix.

Spec-Driven Development: is it the solution to Vibe Coding?

· 13 min read
Rafael Fernandes
NLP Engineer & Tech Writer at WiLine
Share:
Engineering practice · AI News

Spec-driven development, tested

Write the promptWrite the contract

Someone posted spec-driven development on LinkedIn this week as the answer to vibe coding — to prompting an agent, half-understanding what you're building, and ending up with code you can't vouch for. The linked toolkit has 127,000 stars and comes from GitHub itself. The pitch lands.

So I installed it and pointed it at a deliberately trivial task. One of the three principles it wrote for me was a dependency policy I never asked for — hold that thought.

Twenty minutes isn't a verdict, though. Two engineers have tested this properly, on real problems, long enough for the seams to show. They used different tools, on different continents, seven months apart — and both reached for the same comparison, unprompted: the last time our industry tried to generate working code from documents. On the one question that decides whether any of this survives contact with AI features, they flatly contradict each other. Neither has a measurement.

GPT-5.6: OpenAI's new pitch is cheaper per task, not just smarter — verify it on your workload before you switch

· 10 min read
Rafael Fernandes
NLP Engineer & Tech Writer at WiLine
Share:
Models · AI News

The frontier race just changed lanes: from smarter to cheaper per task

Benchmark pointsToken economics

OpenAI shipped GPT-5.6 on July 9 — a family of three models (Luna, Terra, Sol) — and the headline claim isn't a leaderboard score. It's an efficiency number: frontier coding performance on less than half the output tokens. If you build agents, that's a claim about your bill, not about bragging rights. It's also exactly the kind of claim you should measure yourself.