Skip to main content

AI News

Short, high-signal reads on what is changing in AI infrastructure — new tools, releases, and shifts — with a plain take on what each one means if you self-host on WiLine.

Jev: a decision model you put in front of your LLMs to route traffic

· 7 min read

Jev is TypeSafe's first model — it doesn't chat, it decides. Give it a request and it returns one structured label in about 127 ms. LiteLLM wired it into its Auto Router as the classifier and clocked it 5.43x faster and 96% cheaper than a Haiku doing the same job. Here is what Jev is, how routing works, where the LLM classifier actually falls down, and where this fits in front of your models on WEC.

ai-newsroutingcostlitellmmodels
Read more →

The bug that halved a gateway's throughput without a single error

· 6 min read

An engineer wrote a setting that said do not encrypt this connection. The system encrypted it anyway, then sat waiting for a reply that never came. Throughput fell by half, every request still returned success, and no dashboard showed a problem. Pfizer and LiteLLM published the story — and the cause is four lines of code that anyone can read.

ai-newslitellmgatewayredisobservabilityinfrastructure
Read more →

MCP's Fix for Bloated Tool Catalogues Has a Bill Attached — And It's in the Docs, Not the Roadmap

· 8 min read

The MCP roadmap names tool-catalogue bloat as a priority. The client documentation already ships the fix — progressive discovery, with a search_tools meta-tool — and the same page warns that it can cost more than it saves, because loading definitions mid-conversation invalidates the provider's prompt cache. Three layers of caching, only one of which the protocol actually specifies, and the one that bites belongs to your model provider.

ai-newsmcpagentsprotocolscontext-engineeringinfrastructure
Read more →

LangChain Just Made MCP First-Class — We Ran It the Same Day, and Three Things Don't Work Yet

· 10 min read

MCP support moved into the langchain package today, built on FastMCP, with elicitation as a LangGraph interrupt and a cacheable tool catalog. We installed it the same afternoon and pointed it at a server we wrote against the new spec. The announcement is accurate; the ecosystem around it hasn't caught up, and one of the gaps turns a human-approval gate into a sentence the model makes up.

ai-newsmcplangchainlanggraphagentsprotocols
Read more →

An agent can narrow its own web search, but never widen it

· 4 min read

AWS gave agent web search a list of allowed sites on 19 August. The admin sets one list, the agent can set another, and when they disagree the agent's list can only make the search smaller. That one rule is the difference between a guardrail and a note in the prompt.

ai-newsagentsmcpguardrailsgovernanceweb-search
Read more →

A router that can't see the conversation can't classify "yes"

· 7 min read

LiteLLM ran 5,600 live classifier calls to answer one question: how much of the conversation does a model router need to see? On the follow-ups that only make sense against history, agreement went from 14% to 78% — and the completion bill more than doubled, because two thirds of them had been going to the cheapest model.

ai-newsroutingevalscostlitellm
Read more →

Mask Your Logs, Not Your Prompts

· 7 min read

The most common LLM privacy advice — scrub personal data out of the prompt before the model sees it — aims at the wrong risk and quietly makes the model dumber. Here's what the research actually shows, and where the masking really belongs.

ai-newsprivacypiiguardrailsgateway
Read more →

Spec-Driven Development: is it the solution to Vibe Coding?

· 13 min read

A GitHub toolkit with 127k stars says you should write the spec before the code, and let the agent build from it. I ran it, then read the two engineers who tested it properly — and both reached for the same comparison: the last time our industry tried generating code from documents.

ai-newsengineering-practiceagentsevalstooling
Read more →

Muse Glimmer: a 30B agentic model that runs on one GPU, no data center required

· 4 min read

Meta released a 30B multimodal model built for local agent workloads — beats Gemma4 and holds its own against Qwen3.6 on agentic benchmarks, and fits on a single consumer GPU. Here's what's actually new, and what it'd take for it to land on WEC.

ai-newsmodelsopen-weightagentslocal-inference
Read more →

'Loop Engineering Is Dead' — and the Real Story Is Weirder Than the Obituary

· 13 min read

In June, 'loop engineering' got a name. Six weeks later it was declared dead — and the industry answered with vendor guides, competing definitions, and a wave of SEO. Here's what loop and graph engineering actually mean, why a free MIT textbook settles the argument on page 189, and what a six-week hype cycle should teach you about what to learn.

ai-newsagentsarchitectureorchestrationengineering-practice
Read more →