Skip to main content

14 posts tagged with "ai-news"

View all tags

Jev: a decision model you put in front of your LLMs to route traffic

· 7 min read
Rafael Fernandes
NLP Engineer & Tech Writer at WiLine
Share:
Routing · AI News

Jev decides, your LLMs answer

A request comes inThe right model, in 127 ms

TypeSafe shipped its first model on 15 September, and the interesting thing about Jev is what it refuses to do. It doesn't write you a paragraph. You give it a request and it hands back one structured value — a label, a class, a decision — in, they say, 70 to 500 ms. Founder Diogo Almeida's framing is the clearest line in the post: "Think of Jev as a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out."

Most models are built to talk to people. Jev is built to be called by code — and the first job that shape fits is routing.

The bug that halved a gateway's throughput without a single error

· 6 min read
Rafael Fernandes
NLP Engineer & Tech Writer at WiLine
Share:
Gateways · AI NewsA gateway losing half its throughput with no errors reported

Pfizer runs an AI gateway — the single service every internal tool talks to when it wants a model. During a routine upgrade check, its throughput fell by half. No errors. No failed requests. Nothing wrong on any dashboard.

They published what they found together with the LiteLLM team, and the cause is small enough to fit in a paragraph: someone wrote a setting that said don't encrypt this connection, and the system encrypted it anyway.

MCP's Fix for Bloated Tool Catalogues Has a Bill Attached — And It's in the Docs, Not the Roadmap

· 8 min read
Rafael Fernandes
NLP Engineer & Tech Writer at WiLine
Share:
Protocols · AI NewsProgressive tool discovery and the provider prompt cache

The MCP maintainers' current roadmap names a problem most people building agents have felt without measuring: a server's tool catalogue is charged to the model before anyone asks a question. Under Improved primitives, the post is blunt about it — "Connecting to a server with a hundred tools means the model pays for that entire surface before the user has asked a single question, and tool selection tends to get worse as the list grows."

What makes this worth reading is not the roadmap. It's that the fix is already written down, in the client documentation, alongside a warning that it can cost more than the problem it solves — and that warning gets far less attention than the fix it qualifies.

LangChain Just Made MCP First-Class — We Ran It the Same Day, and Three Things Don't Work Yet

· 10 min read
Rafael Fernandes
NLP Engineer & Tech Writer at WiLine
Share:
Agent Frameworks · AI NewsMCP support lands in the main LangChain package

In July the MCP specification ripped out sessions. We wrote at the time that the change was infrastructure, not changelog — that a stateless core would let servers sit behind ordinary load balancers and survive redeploys, and that clients would need to catch up.

Today LangChain caught up. MCP support moved out of the separate langchain-mcp-adapters package and into langchain itself, rebuilt on FastMCP, with two features the old spec made impossible: elicitation as a LangGraph interrupt, and a cacheable tool catalog.

We installed it the same afternoon and pointed it at a server we wrote against the new spec. The announcement is accurate. The ecosystem around it is not ready — and one of the gaps quietly converts a human-approval gate into a sentence the model invents.

An agent can narrow its own web search, but never widen it

· 4 min read
Rafael Fernandes
NLP Engineer & Tech Writer at WiLine
Share:
Agents · AI News

Narrow only, never wider

Please use trusted sourcesTrusted sources are all there are

Your agent can search the web. You write in the prompt: only use sec.gov and the big financial wires. It usually listens. The times it does not are the times you learn that a line in a prompt is a request, not a rule.

On 19 August AWS added site filters to the web search tool in Bedrock AgentCore. Small feature. But the way it handles a disagreement is worth knowing, because that is what makes it a guardrail instead of a note.

A router that can't see the conversation can't classify "yes"

· 7 min read
Rafael Fernandes
NLP Engineer & Tech Writer at WiLine
Share:
Routing · AI News

Classifying the word "yes"

Score this messageScore what it approves

A model router's job is to read a request and decide which model should answer it. Cheap questions go to a small model, hard ones to a large one, and the bill comes down. The whole arrangement rests on being able to tell the difference.

Then a user types "yes".

Or "continue". Or "do it". Nothing in those two or three characters says whether the work being approved is a spelling fix or a database migration. A router scoring the current message in isolation sees a very short string with no technical vocabulary, and does the obvious thing: cheapest model.

Which means if you route requests to save money, your cheapest tier is probably absorbing work it should never have seen — and your savings figure is partly fake. On 4 August LiteLLM published a benchmark that measures both halves of that: how wrong the routing gets, and what it costs to fix.

Mask Your Logs, Not Your Prompts

· 7 min read
Rafael Fernandes
NLP Engineer & Tech Writer at WiLine
Share:
Privacy · AI News

Mask your logs, not your prompts

Redact before the modelRedact before the logs

Almost every "secure your LLM app" guide gives the same advice: before a prompt reaches the model, strip the personal data out of it — swap names, emails, and account numbers for [REDACTED] or <PERSON>, then call the model.

It sounds obviously right. I assumed it was, too. Then I went and read the research on what masking actually does to a model, and the papers point the other way. The short version: mask your logs, not your prompts.

Spec-Driven Development: is it the solution to Vibe Coding?

· 13 min read
Rafael Fernandes
NLP Engineer & Tech Writer at WiLine
Share:
Engineering practice · AI News

Spec-driven development, tested

Write the promptWrite the contract

Someone posted spec-driven development on LinkedIn this week as the answer to vibe coding — to prompting an agent, half-understanding what you're building, and ending up with code you can't vouch for. The linked toolkit has 127,000 stars and comes from GitHub itself. The pitch lands.

So I installed it and pointed it at a deliberately trivial task. One of the three principles it wrote for me was a dependency policy I never asked for — hold that thought.

Twenty minutes isn't a verdict, though. Two engineers have tested this properly, on real problems, long enough for the seams to show. They used different tools, on different continents, seven months apart — and both reached for the same comparison, unprompted: the last time our industry tried to generate working code from documents. On the one question that decides whether any of this survives contact with AI features, they flatly contradict each other. Neither has a measurement.

Muse Glimmer: a 30B agentic model that runs on one GPU, no data center required

· 4 min read
Rafael Fernandes
NLP Engineer & Tech Writer at WiLine
Share:
Models · AI News

A serious agent, no data center required

Cloud-only agentsOne GPU, fully local

Most "run it locally" model announcements come with an asterisk — smaller, weaker, a toy version of the real thing. Meta's newest release doesn't: Muse Glimmer, a 30B multimodal model built specifically for agentic work, fits on a single consumer GPU and beats larger models on the benchmarks that actually measure agent behavior.

'Loop Engineering Is Dead' — and the Real Story Is Weirder Than the Obituary

· 13 min read
Rafael Fernandes
NLP Engineer & Tech Writer at WiLine
Share:
Architecture · AI News

Loops, graphs, and the six-week obituary

Loop engineeringGraph engineering

In June 2026, the AI world got a new buzzword: loop engineering — roughly, disciplined design of a single agent's tool-calling loop. Six weeks later it was supposedly replaced by graph engineering — wiring up several agents at once — killed by twelve words that 3.1 million people saw:

Don't know what loop engineering is? Don't worry — neither did most of the people declaring it dead. Here are both terms, how a name became an obituary in six weeks, and the twist nobody checked before writing about it.