AI News
Short, high-signal reads on what is changing in AI infrastructure — new tools, releases, and shifts — with a plain take on what each one means if you self-host on WiLine.
Jev: a decision model you put in front of your LLMs to route traffic
Jev is TypeSafe's first model — it doesn't chat, it decides. Give it a request and it returns one structured label in about 127 ms. LiteLLM wired it into its Auto Router as the classifier and clocked it 5.43x faster and 96% cheaper than a Haiku doing the same job. Here is what Jev is, how routing works, where the LLM classifier actually falls down, and where this fits in front of your models on WEC.
Read more →The bug that halved a gateway's throughput without a single error
An engineer wrote a setting that said do not encrypt this connection. The system encrypted it anyway, then sat waiting for a reply that never came. Throughput fell by half, every request still returned success, and no dashboard showed a problem. Pfizer and LiteLLM published the story — and the cause is four lines of code that anyone can read.
Read more →MCP's Fix for Bloated Tool Catalogues Has a Bill Attached — And It's in the Docs, Not the Roadmap
The MCP roadmap names tool-catalogue bloat as a priority. The client documentation already ships the fix — progressive discovery, with a search_tools meta-tool — and the same page warns that it can cost more than it saves, because loading definitions mid-conversation invalidates the provider's prompt cache. Three layers of caching, only one of which the protocol actually specifies, and the one that bites belongs to your model provider.
Read more →LangChain Just Made MCP First-Class — We Ran It the Same Day, and Three Things Don't Work Yet
MCP support moved into the langchain package today, built on FastMCP, with elicitation as a LangGraph interrupt and a cacheable tool catalog. We installed it the same afternoon and pointed it at a server we wrote against the new spec. The announcement is accurate; the ecosystem around it hasn't caught up, and one of the gaps turns a human-approval gate into a sentence the model makes up.
Read more →An agent can narrow its own web search, but never widen it
AWS gave agent web search a list of allowed sites on 19 August. The admin sets one list, the agent can set another, and when they disagree the agent's list can only make the search smaller. That one rule is the difference between a guardrail and a note in the prompt.
Read more →A router that can't see the conversation can't classify "yes"
LiteLLM ran 5,600 live classifier calls to answer one question: how much of the conversation does a model router need to see? On the follow-ups that only make sense against history, agreement went from 14% to 78% — and the completion bill more than doubled, because two thirds of them had been going to the cheapest model.
Read more →Mask Your Logs, Not Your Prompts
The most common LLM privacy advice — scrub personal data out of the prompt before the model sees it — aims at the wrong risk and quietly makes the model dumber. Here's what the research actually shows, and where the masking really belongs.
Read more →Spec-Driven Development: is it the solution to Vibe Coding?
A GitHub toolkit with 127k stars says you should write the spec before the code, and let the agent build from it. I ran it, then read the two engineers who tested it properly — and both reached for the same comparison: the last time our industry tried generating code from documents.
Read more →Muse Glimmer: a 30B agentic model that runs on one GPU, no data center required
Meta released a 30B multimodal model built for local agent workloads — beats Gemma4 and holds its own against Qwen3.6 on agentic benchmarks, and fits on a single consumer GPU. Here's what's actually new, and what it'd take for it to land on WEC.
Read more →'Loop Engineering Is Dead' — and the Real Story Is Weirder Than the Obituary
In June, 'loop engineering' got a name. Six weeks later it was declared dead — and the industry answered with vendor guides, competing definitions, and a wave of SEO. Here's what loop and graph engineering actually mean, why a free MIT textbook settles the argument on page 189, and what a six-week hype cycle should teach you about what to learn.
Read more →