<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
    <id>https://development-wec.wiline.com/docs/news/</id>
    <title>WiLine Edge Cloud — AI News</title>
    <updated>2026-09-23T00:00:00.000Z</updated>
    <generator>https://github.com/jpmonette/feed</generator>
    <link rel="alternate" href="https://development-wec.wiline.com/docs/news/"/>
    <subtitle>Short, high-signal takes on what is changing in AI infrastructure — and what it means for self-hosting on WiLine.</subtitle>
    <icon>https://development-wec.wiline.com/docs/img/Logo_icon_blue.svg</icon>
    <entry>
        <title type="html"><![CDATA[Jev: a decision model you put in front of your LLMs to route traffic]]></title>
        <id>https://development-wec.wiline.com/docs/news/router-classifier-tax-system-one-jev/</id>
        <link href="https://development-wec.wiline.com/docs/news/router-classifier-tax-system-one-jev/"/>
        <updated>2026-09-23T00:00:00.000Z</updated>
        <summary type="html"><![CDATA[Jev is TypeSafe's first model — it doesn't chat, it decides. Give it a request and it returns one structured label in about 127 ms. LiteLLM wired it into its Auto Router as the classifier and clocked it 5.43x faster and 96% cheaper than a Haiku doing the same job. Here is what Jev is, how routing works, where the LLM classifier actually falls down, and where this fits in front of your models on WEC.]]></summary>
        <content type="html"><![CDATA[<div class="newsHero"><div class="newsHero__glow" aria-hidden="true"></div><span class="newsHero__eyebrow">Routing · AI News</span><h2 class="newsHero__title">Jev decides, your LLMs answer</h2><div class="newsHero__transition"><span class="newsHero__pill newsHero__pill--from">A request comes in</span><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2.5" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-arrow-right newsHero__arrow" aria-hidden="true"><path d="M5 12h14"></path><path d="m12 5 7 7-7 7"></path></svg><span class="newsHero__pill newsHero__pill--to">The right model, in 127 ms</span></div></div>
<p>TypeSafe <a href="https://typesafe.ai/blog/introducing-system-one-models-and-jev" target="_blank" rel="noopener noreferrer" class="">shipped its first model on 15 September</a>,
and the interesting thing about Jev is what it refuses to do. It doesn't write you a
paragraph. You give it a request and it hands back one structured value — a label, a
class, a decision — in, they say, 70 to 500 ms. Founder Diogo Almeida's framing is
the clearest line in the post: <em>"Think of Jev as a frontier-intelligence function
call: unstructured state in, typed probabilistic decisions out."</em></p>
<p>Most models are built to talk to people. Jev is built to be called by code — and the
first job that shape fits is routing.</p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="a-model-built-to-decide-not-to-talk">A model built to decide, not to talk<a href="https://development-wec.wiline.com/docs/news/router-classifier-tax-system-one-jev/#a-model-built-to-decide-not-to-talk" class="hash-link" aria-label="Direct link to A model built to decide, not to talk" title="Direct link to A model built to decide, not to talk" translate="no">​</a></h2>
<p>The name tells the story twice. "System One" is Kahneman's term for fast, automatic
judgement — the snap decision — against "System Two," the slow deliberate reasoning
we reach for a big LLM to do. And "Jev" is for <strong>William Stanley Jevons</strong>, the
economist of the efficiency paradox: make something cheaper and people use far more
of it. TypeSafe is betting cheap machine decisions get used everywhere.</p>
<p>A chat model generates tokens one at a time until a paragraph exists. Jev does the
opposite — TypeSafe says it "gives up string generation," samples all outputs in a
single parallel query, and returns one value from a set you define in advance. Two
properties fall out, and both matter for routing:</p>
<ul>
<li class=""><strong>The output is typed and structured</strong>, from a known set (<code>SIMPLE</code>, <code>MEDIUM</code>,
<code>COMPLEX</code>, <code>REASONING</code>), not prose you have to parse. TypeSafe claims it "never
makes type errors" — worth being precise here, because they are: that 0% is
<em>"not empirical. Schema matching is guaranteed, thus we can confidently add 0%."</em>
It's a property of constraining the output to a schema, not a measured result.</li>
<li class=""><strong>Each decision carries a calibrated probability</strong> — "higher confidence means
higher accuracy," trained via a method they call Reinforcement Learning for
Calibrated Decisions (RLCD). If that holds, code can act on the number: take the
label when confident, escalate when not.</li>
</ul>
<p>Their headline numbers ("40x-200x faster," output "too cheap to meter," a workflow
eval at "193.6x faster, 444.6x cheaper") are vendor claims on their own evals, and
to their credit they flag the bias — the evals were "made by individuals on our
model capabilities team." Treat them as claims until someone independent re-runs
them.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="how-routing-works-and-why-the-decider-is-the-problem">How routing works, and why the decider is the problem<a href="https://development-wec.wiline.com/docs/news/router-classifier-tax-system-one-jev/#how-routing-works-and-why-the-decider-is-the-problem" class="hash-link" aria-label="Direct link to How routing works, and why the decider is the problem" title="Direct link to How routing works, and why the decider is the problem" translate="no">​</a></h2>
<p>Model routing is a cost play. Not every request needs your biggest model, so you put
a cheap model and an expensive one behind one endpoint and something in front decides
which answers. Easy turns take the cheap tier; hard ones take the expensive tier; the
bill drops. That "something in front" is a classifier:</p>
<!-- -->
<p>Build that classifier the usual way — ask <em>another</em> LLM "is this simple or complex?"
— and every request pays for an LLM call <strong>before</strong> it pays for the answer. That call
is slow and costs tokens, on every request, whether it landed on the cheap tier or
not. The decision meant to save money is quietly spending it.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="jev-in-the-decider-slot">Jev in the decider slot<a href="https://development-wec.wiline.com/docs/news/router-classifier-tax-system-one-jev/#jev-in-the-decider-slot" class="hash-link" aria-label="Direct link to Jev in the decider slot" title="Direct link to Jev in the decider slot" translate="no">​</a></h2>
<p>Jev fits that slot. You don't ask it to answer — you ask which tier the request
belongs to, and the gateway routes on the label:</p>
<figure class="stageFlow"><div class="stageFlow__track"><div class="stageFlow__card" style="background:rgba(var(--primary-rgb), 0.050);border-color:rgba(var(--primary-rgb), 0.250)"><span class="stageFlow__stage">1</span><span class="stageFlow__title">Request arrives</span><span class="stageFlow__tag">at the gateway</span></div><svg xmlns="http://www.w3.org/2000/svg" width="22" height="22" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2.5" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-arrow-right stageFlow__arrow" aria-hidden="true"><path d="M5 12h14"></path><path d="m12 5 7 7-7 7"></path></svg><div class="stageFlow__card" style="background:rgba(var(--primary-rgb), 0.123);border-color:rgba(var(--primary-rgb), 0.383)"><span class="stageFlow__stage">2</span><span class="stageFlow__title">Jev labels it</span><span class="stageFlow__tag">~127 ms, one tier</span></div><svg xmlns="http://www.w3.org/2000/svg" width="22" height="22" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2.5" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-arrow-right stageFlow__arrow" aria-hidden="true"><path d="M5 12h14"></path><path d="m12 5 7 7-7 7"></path></svg><div class="stageFlow__card" style="background:rgba(var(--primary-rgb), 0.197);border-color:rgba(var(--primary-rgb), 0.517)"><span class="stageFlow__stage">3</span><span class="stageFlow__title">Gateway routes</span><span class="stageFlow__tag">to the matching model</span></div><svg xmlns="http://www.w3.org/2000/svg" width="22" height="22" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2.5" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-arrow-right stageFlow__arrow" aria-hidden="true"><path d="M5 12h14"></path><path d="m12 5 7 7-7 7"></path></svg><div class="stageFlow__card" style="background:rgba(var(--primary-rgb), 0.270);border-color:rgba(var(--primary-rgb), 0.650)"><span class="stageFlow__stage">4</span><span class="stageFlow__title">Model answers</span><span class="stageFlow__tag">small or large</span></div></div><figcaption class="stageFlow__caption">The classifier is step 2. Make it cheap and the arrangement pays off; make it an LLM and it taxes every call.</figcaption></figure>
<p>In LiteLLM's Auto Router that is a config change, not new code — you name the tiers,
map each to a model, and set <code>classifier_type: jev</code>:</p>
<div class="language-yaml codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#393A34;--prism-background-color:#f6f8fa"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-yaml codeBlock_bY9V thin-scrollbar" style="color:#393A34;background-color:#f6f8fa"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#393A34"><span class="token punctuation" style="color:#393A34">-</span><span class="token plain"> </span><span class="token key atrule" style="color:#00a4db">model_name</span><span class="token punctuation" style="color:#393A34">:</span><span class="token plain"> jev</span><span class="token punctuation" style="color:#393A34">-</span><span class="token plain">router</span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">  </span><span class="token key atrule" style="color:#00a4db">litellm_params</span><span class="token punctuation" style="color:#393A34">:</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">    </span><span class="token key atrule" style="color:#00a4db">model</span><span class="token punctuation" style="color:#393A34">:</span><span class="token plain"> auto_router/complexity_router</span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">    </span><span class="token key atrule" style="color:#00a4db">complexity_router_config</span><span class="token punctuation" style="color:#393A34">:</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">      </span><span class="token key atrule" style="color:#00a4db">tiers</span><span class="token punctuation" style="color:#393A34">:</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">        </span><span class="token key atrule" style="color:#00a4db">SIMPLE</span><span class="token punctuation" style="color:#393A34">:</span><span class="token plain"> </span><span class="token punctuation" style="color:#393A34">{</span><span class="token punctuation" style="color:#393A34">{</span><span class="token plain">openai_small</span><span class="token punctuation" style="color:#393A34">}</span><span class="token punctuation" style="color:#393A34">}</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">        </span><span class="token key atrule" style="color:#00a4db">MEDIUM</span><span class="token punctuation" style="color:#393A34">:</span><span class="token plain"> </span><span class="token punctuation" style="color:#393A34">{</span><span class="token punctuation" style="color:#393A34">{</span><span class="token plain">openai_large</span><span class="token punctuation" style="color:#393A34">}</span><span class="token punctuation" style="color:#393A34">}</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">        </span><span class="token key atrule" style="color:#00a4db">COMPLEX</span><span class="token punctuation" style="color:#393A34">:</span><span class="token plain"> </span><span class="token punctuation" style="color:#393A34">{</span><span class="token punctuation" style="color:#393A34">{</span><span class="token plain">anthropic</span><span class="token punctuation" style="color:#393A34">}</span><span class="token punctuation" style="color:#393A34">}</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">        </span><span class="token key atrule" style="color:#00a4db">REASONING</span><span class="token punctuation" style="color:#393A34">:</span><span class="token plain"> </span><span class="token punctuation" style="color:#393A34">{</span><span class="token punctuation" style="color:#393A34">{</span><span class="token plain">anthropic_large</span><span class="token punctuation" style="color:#393A34">}</span><span class="token punctuation" style="color:#393A34">}</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">      </span><span class="token key atrule" style="color:#00a4db">classifier_type</span><span class="token punctuation" style="color:#393A34">:</span><span class="token plain"> jev</span><br></div></code></pre></div></div>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="what-the-swap-is-worth--and-where-the-llm-falls-down">What the swap is worth — and where the LLM falls down<a href="https://development-wec.wiline.com/docs/news/router-classifier-tax-system-one-jev/#what-the-swap-is-worth--and-where-the-llm-falls-down" class="hash-link" aria-label="Direct link to What the swap is worth — and where the LLM falls down" title="Direct link to What the swap is worth — and where the LLM falls down" translate="no">​</a></h2>
<p>On 20 September LiteLLM's Moe Khalil <a href="https://docs.litellm.ai/blog/jev-auto-router-benchmark" target="_blank" rel="noopener noreferrer" class="">published a benchmark</a>
titled <em>"JEV Classifier: 5.43x as Fast as Haiku, 96% Lower Cost."</em> He ran 80 cases
three times each — 240 classification calls — with Jev (<code>jev-1.13.0</code>) against Claude
Haiku 4.5 as the "classify with an LLM" baseline:</p>
<div class="metricCompare"><div class="metricCard"><span class="metricCard__label">Classifier latency (p50)</span><span class="metricCard__factor">5.43× faster</span><div class="metricCard__rows"><div class="metricCard__row metricCard__row--a"><span class="metricCard__name">Haiku 4.5 classifier</span><span class="metricCard__val">688.40 ms</span></div><div class="metricCard__row metricCard__row--b"><span class="metricCard__name">Jev classifier</span><span class="metricCard__val">126.81 ms</span></div></div></div><div class="metricCard"><span class="metricCard__label">Matched the expected tier</span><span class="metricCard__factor">+21 pts</span><div class="metricCard__rows"><div class="metricCard__row metricCard__row--a"><span class="metricCard__name">Haiku 4.5 classifier</span><span class="metricCard__val">73.75% (177/240)</span></div><div class="metricCard__row metricCard__row--b"><span class="metricCard__name">Jev classifier</span><span class="metricCard__val">95.00% (228/240)</span></div></div></div><div class="metricCard"><span class="metricCard__label">Cost for 240 calls</span><span class="metricCard__factor">~96% cheaper</span><div class="metricCard__rows"><div class="metricCard__row metricCard__row--a"><span class="metricCard__name">Haiku 4.5 classifier</span><span class="metricCard__val">$0.1985</span></div><div class="metricCard__row metricCard__row--b"><span class="metricCard__name">Jev classifier</span><span class="metricCard__val">$0.0077</span></div></div></div></div>
<p>The average hides the real finding, which is in the per-tier numbers. On <code>SIMPLE</code>
both were perfect (100%). But on <code>MEDIUM</code> Haiku collapsed to <strong>38.33%</strong> where Jev
held <strong>85%</strong>, and on <code>REASONING</code> Haiku managed <strong>76.67%</strong> against Jev's <strong>100%</strong>. In
other words, the LLM classifier is fine at telling trivial from non-trivial, and
unreliable exactly in the middle band where routing decisions actually save or cost
you money.</p>
<p>Two caveats, both LiteLLM's own, stated plainly: the expected tiers "were authored
with the synthetic prompts, without independent review," and the benchmark measured
<strong>the routing decision, not the quality of the final answer</strong> — "downstream answer
quality was not measured." To their credit they published a frozen, hash-verified
<a href="https://docs.litellm.ai/blog/jev-auto-router-benchmark" target="_blank" rel="noopener noreferrer" class="">reproduction archive</a> so the
run can be re-analysed. Read the result as "Jev picked the tier fast, cheap, and the
way they expected," not "your answers get 96% cheaper end to end."</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="where-this-fits-on-wec">Where this fits on WEC<a href="https://development-wec.wiline.com/docs/news/router-classifier-tax-system-one-jev/#where-this-fits-on-wec" class="hash-link" aria-label="Direct link to Where this fits on WEC" title="Direct link to Where this fits on WEC" translate="no">​</a></h2>
<p>We build this exact setup in the tutorial on
<a class="" href="https://development-wec.wiline.com/docs/tutorials/litellm-complexity-routing/">routing by complexity</a> — a gateway in front
of the WEC Inference API that scores each request and sends it to the right tier,
Qwen3.5:9B for the simple turns and Qwen3.5:122B for the hard ones. And we have shown
how a classifier gets it <a class="" href="https://development-wec.wiline.com/docs/news/llm-router-cannot-classify-yes/">wrong on a bare "yes"</a>,
sending work to the wrong model.</p>
<p>In both, the classifier is the part nobody prices. Put an LLM in that slot and you
add two-thirds of a second and a token charge to every call before the real model is
even chosen. A model like Jev is a bet that the decision in front of your models can
be near-free — so routing pays off instead of taxing itself.</p>
<p>It is early, it is hosted, and the numbers still need someone independent to re-run
them. For a WEC stack that keeps inference in-house, "hosted" is the real question
mark — you would be sending every request's routing decision to a third party. But
the shape is the part that travels: a model that decides in milliseconds, sitting in
front of the models that answer.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="sources">Sources<a href="https://development-wec.wiline.com/docs/news/router-classifier-tax-system-one-jev/#sources" class="hash-link" aria-label="Direct link to Sources" title="Direct link to Sources" translate="no">​</a></h2>
<ul>
<li class="">Diogo Almeida, TypeSafe — <a href="https://typesafe.ai/blog/introducing-system-one-models-and-jev" target="_blank" rel="noopener noreferrer" class=""><em>Introducing System One Models and Jev</em></a> (15 September 2026)</li>
<li class="">Moe Khalil, LiteLLM — <a href="https://docs.litellm.ai/blog/jev-auto-router-benchmark" target="_blank" rel="noopener noreferrer" class=""><em>JEV Classifier: 5.43x as Fast as Haiku, 96% Lower Cost</em></a> (20 September 2026)</li>
</ul>]]></content>
        <author>
            <name>Rafael Fernandes</name>
            <uri>https://www.linkedin.com/in/rafaelmacariofernandes/</uri>
        </author>
        <category label="ai-news" term="ai-news"/>
        <category label="routing" term="routing"/>
        <category label="cost" term="cost"/>
        <category label="litellm" term="litellm"/>
        <category label="models" term="models"/>
    </entry>
    <entry>
        <title type="html"><![CDATA[The bug that halved a gateway's throughput without a single error]]></title>
        <id>https://development-wec.wiline.com/docs/news/pfizer-gateway-silent-throughput-bug/</id>
        <link href="https://development-wec.wiline.com/docs/news/pfizer-gateway-silent-throughput-bug/"/>
        <updated>2026-09-16T00:00:00.000Z</updated>
        <summary type="html"><![CDATA[An engineer wrote a setting that said do not encrypt this connection. The system encrypted it anyway, then sat waiting for a reply that never came. Throughput fell by half, every request still returned success, and no dashboard showed a problem. Pfizer and LiteLLM published the story — and the cause is four lines of code that anyone can read.]]></summary>
        <content type="html"><![CDATA[<figure class="newsHero newsHero--image"><span class="newsHero__chip">Gateways · AI News</span><img src="https://development-wec.wiline.com/docs/img/news/litellm-cover.webp" alt="A gateway losing half its throughput with no errors reported" loading="eager"></figure>
<p>Pfizer runs an AI gateway — the single service every internal tool talks to when it
wants a model. During a routine upgrade check, its throughput fell by half. No
errors. No failed requests. Nothing wrong on any dashboard.</p>
<p>They published <a href="https://docs.litellm.ai/blog/pfizer-gateway-performance-and-resiliency" target="_blank" rel="noopener noreferrer" class="">what they found</a>
together with the LiteLLM team, and the cause is small enough to fit in a paragraph:
someone wrote a setting that said <em>don't encrypt this connection</em>, and the system
encrypted it anyway.</p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="what-actually-went-wrong">What actually went wrong<a href="https://development-wec.wiline.com/docs/news/pfizer-gateway-silent-throughput-bug/#what-actually-went-wrong" class="hash-link" aria-label="Direct link to What actually went wrong" title="Direct link to What actually went wrong" translate="no">​</a></h2>
<p>The gateway keeps a cache — a fast side-store called Redis that it checks before
doing expensive work. The connection to that cache can be encrypted or not, and
that's a setting you write in a config file.</p>
<p>The operator wrote the setting to <strong>off</strong>. Plain connection, no encryption.</p>
<p>The code that read that setting checked whether the setting <em>existed</em>, rather than
what it was set to. Writing <code>off</code> creates the setting. The setting now exists. So
the check passed, and the gateway opened an encrypted connection — to a cache that
wasn't expecting one.</p>
<p>Here's the part that makes this dangerous rather than merely wrong. When a program
tries to start an encrypted conversation with something that doesn't speak
encryption, it doesn't get turned away. It says hello and waits for a reply that
never comes. Not refused — ignored.</p>
<p>So the cache wasn't <em>down</em>. The cache was <strong>slow</strong>. And slow is much harder to see.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-four-lines">The four lines<a href="https://development-wec.wiline.com/docs/news/pfizer-gateway-silent-throughput-bug/#the-four-lines" class="hash-link" aria-label="Direct link to The four lines" title="Direct link to The four lines" translate="no">​</a></h2>
<p>For anyone who wants to look at it directly, this is
<a href="https://github.com/BerriAI/litellm/blob/v1.89.2/litellm/_redis.py#L701" target="_blank" rel="noopener noreferrer" class=""><code>get_redis_connection_pool</code></a>
as it stood in the affected version:</p>
<div class="language-python codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#393A34;--prism-background-color:#f6f8fa"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-python codeBlock_bY9V thin-scrollbar" style="color:#393A34;background-color:#f6f8fa"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#393A34"><span class="token plain">connection_class </span><span class="token operator" style="color:#393A34">=</span><span class="token plain"> async_redis</span><span class="token punctuation" style="color:#393A34">.</span><span class="token plain">Connection</span><br></div><div class="token-line" style="color:#393A34"><span class="token plain"></span><span class="token keyword" style="color:#00009f">if</span><span class="token plain"> </span><span class="token string" style="color:#e3116c">"ssl"</span><span class="token plain"> </span><span class="token keyword" style="color:#00009f">in</span><span class="token plain"> redis_kwargs</span><span class="token punctuation" style="color:#393A34">:</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">    connection_class </span><span class="token operator" style="color:#393A34">=</span><span class="token plain"> async_redis</span><span class="token punctuation" style="color:#393A34">.</span><span class="token plain">SSLConnection</span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">    redis_kwargs</span><span class="token punctuation" style="color:#393A34">.</span><span class="token plain">pop</span><span class="token punctuation" style="color:#393A34">(</span><span class="token string" style="color:#e3116c">"ssl"</span><span class="token punctuation" style="color:#393A34">,</span><span class="token plain"> </span><span class="token boolean" style="color:#36acaa">None</span><span class="token punctuation" style="color:#393A34">)</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">    redis_kwargs</span><span class="token punctuation" style="color:#393A34">[</span><span class="token string" style="color:#e3116c">"connection_class"</span><span class="token punctuation" style="color:#393A34">]</span><span class="token plain"> </span><span class="token operator" style="color:#393A34">=</span><span class="token plain"> connection_class</span><br></div></code></pre></div></div>
<p><code>"ssl" in redis_kwargs</code> asks <em>does this key exist</em>. It does — you created it when you
set it to false. The branch fires, and you get the encrypted connection you
explicitly turned off.</p>
<p>Leave the setting out entirely and everything works. Write it down and say no, and it
breaks. The careful configuration is the broken one.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="why-nobody-got-an-error">Why nobody got an error<a href="https://development-wec.wiline.com/docs/news/pfizer-gateway-silent-throughput-bug/#why-nobody-got-an-error" class="hash-link" aria-label="Direct link to Why nobody got an error" title="Direct link to Why nobody got an error" translate="no">​</a></h2>
<p>Every layer above the cache is built to tolerate a cache that isn't answering, and
that's correct design — if the cache is unavailable, you do the work the slow way and
carry on. Which is exactly what turns this into silence:</p>
<blockquote>
<p>Redis timeout errors appeared in the proxy's internal logs, but the gateway still
returned HTTP 200 to every caller</p>
</blockquote>
<p>HTTP 200 means <em>success</em>. Every single request reported success, for the entire
duration. The post is precise about where the damage landed:</p>
<blockquote>
<p>Redis operations were stalling on TLS handshake timeouts in the async hot path,
degrading throughput without surfacing HTTP-level errors</p>
</blockquote>
<p>The reported numbers: throughput down from about 300 requests per second to about
156 — roughly half — with median response time under load at 4,200 ms, measured
against 750 simulated concurrent users over a sustained 60-second run.</p>
<p>Those measurements come from Pfizer's internal testing and are reported in the post
rather than independently reproducible. The code is a different matter — you can open
the file at that version and read it yourself.</p>
<p>Their own one-line summary is the best sentence in the writeup:</p>
<blockquote>
<p>A throughput drop with a 0% HTTP error rate is the worst kind of regression to
catch after the fact</p>
</blockquote>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-bug-was-already-there">The bug was already there<a href="https://development-wec.wiline.com/docs/news/pfizer-gateway-silent-throughput-bug/#the-bug-was-already-there" class="hash-link" aria-label="Direct link to The bug was already there" title="Direct link to The bug was already there" translate="no">​</a></h2>
<p>I wanted to know whether this arrived with the upgrade that exposed it. It didn't. I
compared the relevant file between the version that performed fine and the version
that didn't — they're byte-for-byte identical. The post says the same thing:</p>
<blockquote>
<p>The underlying Redis bug existed in older versions too, but surfaced during
Pfizer's v1.89.2 upgrade validation under this workload</p>
</blockquote>
<p>That's the detail worth sitting with. This wasn't a bad release you could roll back.
The flaw had been sitting in the connection code across versions that benchmarked
perfectly well. It needed a particular configuration and a particular load before it
did any visible damage — and then it cost half a production gateway's capacity
without raising anything.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="how-they-actually-found-it">How they actually found it<a href="https://development-wec.wiline.com/docs/news/pfizer-gateway-silent-throughput-bug/#how-they-actually-found-it" class="hash-link" aria-label="Direct link to How they actually found it" title="Direct link to How they actually found it" translate="no">​</a></h2>
<p>Not through monitoring, and not by hunting through recent code changes. By turning
features off and on until the number moved:</p>
<blockquote>
<p>Enabled Redis caching. Median latency jumped to 4,200ms. That isolated it to the
Redis connection path.</p>
</blockquote>
<p>A team with real production traffic, real observability and a direct line to the
vendor found this <em>by hand</em>. There was no signal in the error rate, because there
were no errors. Nothing in the status codes, because they all said success. The only
place it showed up was response time under sustained load — which you see only if
you go looking on purpose.</p>
<p>That matches what I ran into at a far smaller scale when I
<a class="" href="https://development-wec.wiline.com/docs/tutorials/load-test-llm-gateway-blocking-callback/">load tested a gateway to find the stall the median hides</a>:
in this kind of infrastructure, the failures that cost you most are the ones that
never raise an error. A blocking callback and a connection stuck waiting on a
handshake are completely different bugs with the same fingerprint — throughput
collapses, the slowest requests get much slower, and the error rate never moves.</p>
<p>I should be straight about the limits here: I don't run anything at Pfizer's scale,
so I can't reproduce their measurements. What I can do is read the code, and the code
is public.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-fix">The fix<a href="https://development-wec.wiline.com/docs/news/pfizer-gateway-silent-throughput-bug/#the-fix" class="hash-link" aria-label="Direct link to The fix" title="Direct link to The fix" translate="no">​</a></h2>
<p>One line, in the current version:</p>
<div class="language-python codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#393A34;--prism-background-color:#f6f8fa"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-python codeBlock_bY9V thin-scrollbar" style="color:#393A34;background-color:#f6f8fa"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#393A34"><span class="token keyword" style="color:#00009f">if</span><span class="token plain"> redis_kwargs</span><span class="token punctuation" style="color:#393A34">.</span><span class="token plain">pop</span><span class="token punctuation" style="color:#393A34">(</span><span class="token string" style="color:#e3116c">"ssl"</span><span class="token punctuation" style="color:#393A34">,</span><span class="token plain"> </span><span class="token boolean" style="color:#36acaa">None</span><span class="token punctuation" style="color:#393A34">)</span><span class="token punctuation" style="color:#393A34">:</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">    redis_kwargs</span><span class="token punctuation" style="color:#393A34">[</span><span class="token string" style="color:#e3116c">"connection_class"</span><span class="token punctuation" style="color:#393A34">]</span><span class="token plain"> </span><span class="token operator" style="color:#393A34">=</span><span class="token plain"> async_redis</span><span class="token punctuation" style="color:#393A34">.</span><span class="token plain">SSLConnection</span><br></div></code></pre></div></div>
<p>This version reads the setting's <em>value</em> instead of merely noting that it exists. Off
now means off.</p>
<p>It's tracked as <em>"Redis ssl handling — Presence check → value check"</em> under LIT-4307
and <a href="https://github.com/BerriAI/litellm/pull/32590" target="_blank" rel="noopener noreferrer" class="">PR #32590</a>, shipped in v1.93.0.
The pull request title puts it plainly: <em>honor ssl value instead of key presence when
building async connection pool</em>.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="what-to-take-from-it">What to take from it<a href="https://development-wec.wiline.com/docs/news/pfizer-gateway-silent-throughput-bug/#what-to-take-from-it" class="hash-link" aria-label="Direct link to What to take from it" title="Direct link to What to take from it" translate="no">​</a></h2>
<p>Two things, if you run anything with a cache behind it.</p>
<p><strong>Check what your system actually did, not what you told it to do.</strong> Writing a
setting and having that setting take effect are different events, and this is a clean
example of the gap. The connection Pfizer got was the opposite of the one they asked
for, and nothing anywhere said so.</p>
<p><strong>If your testing measures averages, it cannot see this.</strong> Neither can alerting built
on error rates. Sustained load, slowest-request timings, and throughput compared
against a known baseline are what turn a silent stall into something visible.</p>
<p>The dashboards stayed green through the whole thing. That's not a gap you close by
adding another alert on errors — there weren't any to alert on.</p>]]></content>
        <author>
            <name>Rafael Fernandes</name>
            <uri>https://www.linkedin.com/in/rafaelmacariofernandes/</uri>
        </author>
        <category label="ai-news" term="ai-news"/>
        <category label="litellm" term="litellm"/>
        <category label="gateway" term="gateway"/>
        <category label="redis" term="redis"/>
        <category label="observability" term="observability"/>
        <category label="infrastructure" term="infrastructure"/>
    </entry>
    <entry>
        <title type="html"><![CDATA[MCP's Fix for Bloated Tool Catalogues Has a Bill Attached — And It's in the Docs, Not the Roadmap]]></title>
        <id>https://development-wec.wiline.com/docs/news/mcp-progressive-discovery-prompt-cache/</id>
        <link href="https://development-wec.wiline.com/docs/news/mcp-progressive-discovery-prompt-cache/"/>
        <updated>2026-09-07T00:00:00.000Z</updated>
        <summary type="html"><![CDATA[The MCP roadmap names tool-catalogue bloat as a priority. The client documentation already ships the fix — progressive discovery, with a search_tools meta-tool — and the same page warns that it can cost more than it saves, because loading definitions mid-conversation invalidates the provider's prompt cache. Three layers of caching, only one of which the protocol actually specifies, and the one that bites belongs to your model provider.]]></summary>
        <content type="html"><![CDATA[<figure class="newsHero newsHero--image"><span class="newsHero__chip">Protocols · AI News</span><img src="https://development-wec.wiline.com/docs/img/news/mcp-drops-sessions-retires-three-core-features-16x9.webp" alt="Progressive tool discovery and the provider prompt cache" loading="eager"></figure>
<p>The MCP maintainers' <a href="https://blog.modelcontextprotocol.io/posts/mcp-roadmap/" target="_blank" rel="noopener noreferrer" class="">current roadmap</a>
names a problem most people building agents have felt without measuring: a server's
tool catalogue is charged to the model before anyone asks a question. Under
<em>Improved primitives</em>, the post is blunt about it —
<em>"Connecting to a server with a hundred tools means the model pays for that entire
surface before the user has asked a single question, and tool selection tends to get
worse as the list grows."</em></p>
<p>What makes this worth reading is not the roadmap. It's that the fix is already
written down, in the client documentation, alongside a warning that it can cost
more than the problem it solves — and that warning gets far less attention than
the fix it qualifies.</p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-number-the-docs-put-on-it">The number the docs put on it<a href="https://development-wec.wiline.com/docs/news/mcp-progressive-discovery-prompt-cache/#the-number-the-docs-put-on-it" class="hash-link" aria-label="Direct link to The number the docs put on it" title="Direct link to The number the docs put on it" translate="no">​</a></h2>
<p>The <a href="https://modelcontextprotocol.io/docs/2026-07-28/develop/clients/client-best-practices" target="_blank" rel="noopener noreferrer" class="">client best-practices page</a>
that shipped with the 2026-07-28 specification describes <strong>progressive
discovery</strong>: the host still calls <code>tools/list</code> as normal, but defers injecting
those definitions into the model's context. Instead it hands the model one
lightweight <code>search_tools</code> meta-tool, and loads full schemas only for what comes
back.</p>
<p>The page's own diagram puts figures on the difference — roughly <strong>150,000 tokens</strong>
consumed by definitions alone when everything is loaded upfront, against about
<strong>2,000</strong> under progressive discovery. That is the documentation's own
illustration rather than a published benchmark, so treat it as the maintainers'
order-of-magnitude claim, not a measurement you can reproduce. It is still a
useful shape: two orders of magnitude, paid before the conversation starts.</p>
<p>The docs also give a threshold rather than leaving it to taste:</p>
<blockquote>
<p>Implement a threshold as a percentage of the context window. For example, 1%-5%.</p>
</blockquote>
<p>Below that, loading everything is fine. Above it, switch.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-trap">The trap<a href="https://development-wec.wiline.com/docs/news/mcp-progressive-discovery-prompt-cache/#the-trap" class="hash-link" aria-label="Direct link to The trap" title="Direct link to The trap" translate="no">​</a></h2>
<p>Here is the sentence that changes how you'd build this. Still on the same page,
under <em>Interaction with Prompt Caching</em>:</p>
<blockquote>
<p>Adding or removing tool definitions mid-conversation invalidates that cache, and
the resulting miss can cost more tokens than the definitions you removed.</p>
</blockquote>
<p>Most providers cache the prompt prefix, and the <code>tools</code> array sits near the front
of it. So the mechanism that saves you 148,000 tokens of upfront definitions works
by <em>mutating the prefix</em> — which is precisely what a prompt cache cannot tolerate.
A prefix cache is only valid up to the first thing that changed: edit the array at
turn six and everything cached after that point goes, which on a long conversation
is nearly all of it.</p>
<p>The docs' own mitigations are worth reading as design constraints rather than
tips: append new definitions strictly <em>after</em> the cache breakpoint rather than
re-sorting the <code>tools</code> array, or route every call through a single stable
<code>call_tool({name, args})</code> meta-tool so the array never changes at all. And treat
disconnecting a server as a conversation boundary, not a per-turn operation.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="three-caches-and-only-one-is-the-protocols">Three caches, and only one is the protocol's<a href="https://development-wec.wiline.com/docs/news/mcp-progressive-discovery-prompt-cache/#three-caches-and-only-one-is-the-protocols" class="hash-link" aria-label="Direct link to Three caches, and only one is the protocol's" title="Direct link to Three caches, and only one is the protocol's" translate="no">​</a></h2>
<p>This is where it gets genuinely confusing, and it's worth separating the layers
because they are easy to collapse into one.</p>
<p><strong>The transport cache.</strong> <code>tools/list</code> results carry <code>ttlMs</code> and <code>cacheScope</code>
hints, defined in the specification's caching utility. A client that honours them
skips the HTTP round trip. This one is MCP's, in the sense that the protocol
specifies it.</p>
<p><strong>The host-side memo.</strong> Advice rather than protocol — a line in the client
documentation's implementation guidelines, recommending you memoise a fetched
definition so re-injecting it later doesn't need another call. It describes what
a host should do with its own state; nothing crosses the wire. The page is
explicit about what it does <em>not</em> do: <em>"This is separate from what's currently in
the model's context."</em></p>
<p><strong>The provider's prompt cache.</strong> Not MCP's at all. Owned by whoever serves your
model, keyed on the prefix, and invalidated by exactly the thing progressive
discovery does for a living.</p>
<p>We ran into the first of those the hard way while writing
<a class="" href="https://development-wec.wiline.com/docs/tutorials/mcp-governed-tools-agent-elicitation/">Part 3 of the LangGraph series</a> —
a client-side <code>cache=True</code> does nothing unless the server actually advertises a
TTL, and nothing warns you. Having now read this page properly, the more useful
lesson is that even getting that right buys you a round trip and nothing else.
The schemas still land in context. Those are different problems with different
fixes, and conflating them is the default mistake.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="what-this-means-if-youre-scoping-tools-by-hand">What this means if you're scoping tools by hand<a href="https://development-wec.wiline.com/docs/news/mcp-progressive-discovery-prompt-cache/#what-this-means-if-youre-scoping-tools-by-hand" class="hash-link" aria-label="Direct link to What this means if you're scoping tools by hand" title="Direct link to What this means if you're scoping tools by hand" translate="no">​</a></h2>
<p><a class="" href="https://development-wec.wiline.com/docs/tutorials/scoped-mcp-tools-per-agent/">Part 4</a> split one MCP server between a
scheduler and a billing agent by filtering the catalogue per role — a static
allow-list, decided before the run and fixed for its duration.</p>
<p>Read against these docs, that turns out to have a property worth naming: because
the list never changes mid-conversation, it doesn't touch the prompt prefix, so
it sidesteps the cache-invalidation problem entirely. It is a cruder instrument
than <code>search_tools</code> — you decide up front instead of letting the model search —
but for a small catalogue split across known roles, "crude and cache-stable" may
simply be the right trade. The docs say as much in the other direction: below the
1–5% threshold, loading everything is fine.</p>
<p>I would not generalise that further. We have not measured cache-miss cost against
definition cost on the WEC Inference API, and the honest position is that the
crossover depends on your catalogue size, your conversation length, and your
provider's caching behaviour. What the docs establish is that a crossover
<em>exists</em>.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="on-the-roadmap-having-no-dates">On the roadmap having no dates<a href="https://development-wec.wiline.com/docs/news/mcp-progressive-discovery-prompt-cache/#on-the-roadmap-having-no-dates" class="hash-link" aria-label="Direct link to On the roadmap having no dates" title="Direct link to On the roadmap having no dates" translate="no">​</a></h2>
<p>The roadmap names five priority areas and commits to no release date or version
number anywhere. It's fair to ask why, and the fairest answer comes from the
maintainers themselves. Soria Parra, in the
<a href="https://blog.modelcontextprotocol.io/posts/2026-mcp-roadmap/" target="_blank" rel="noopener noreferrer" class="">March roadmap</a>:</p>
<blockquote>
<p>A release-oriented roadmap implies a level of predictability that open-standards
work rarely has.</p>
</blockquote>
<p>That's a reasonable position for a specification developed across working groups, and
March's own record backs it up. "Transport Evolution and Scalability" was one of its four
priority areas, and it named the problem precisely — running MCP at scale had surfaced
<em>"a consistent set of gaps: stateful sessions fight with load balancers, horizontal
scaling requires workarounds"</em> — while committing to no date for a fix. The July
specification delivered that stateless core, five months later.</p>
<p>Caching is a different story. It appears nowhere in the March roadmap, and progressive
discovery arrived as client documentation without having been on a roadmap at all — which
is worth noting before treating either roadmap as a reliable index of what is coming.</p>
<p>The reception hasn't been uniformly warm. The <a href="https://news.ycombinator.com/item?id=49399591" target="_blank" rel="noopener noreferrer" class="">Hacker News thread</a>
on the roadmap ran to 270 points and 161 comments, and the criticism runs in three
directions rather than one. The cost of past churn, from <code>colingauvin</code>:</p>
<blockquote>
<p>It's unreal how bad the initial rollout was between HTTP/streaming and stdio,
bearer auth and OAuth. Virtually every client/MCP server pair had a different
portion of that matrix implemented.</p>
</blockquote>
<p>The design itself, from <code>nprateem</code>:</p>
<blockquote>
<p>The real disaster was making it stateful. Need to get some adults in the room.</p>
</blockquote>
<p>And whether the protocol earns its complexity at all, from <code>zackify</code>: <em>"I think the spec
overcomplicates everything honestly."</em></p>
<p>The first of those is the one that bears on this post. It's a fair grievance about
implementation drift, and it's the risk to watch with progressive discovery too: this is
currently a <em>client-side</em> pattern with recommended strategies rather than a specified one,
which means two hosts can both be reasonable and behave differently.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-short-version">The short version<a href="https://development-wec.wiline.com/docs/news/mcp-progressive-discovery-prompt-cache/#the-short-version" class="hash-link" aria-label="Direct link to The short version" title="Direct link to The short version" translate="no">​</a></h2>
<p>The roadmap tells you tool-catalogue bloat is on the maintainers' list. The
documentation tells you what to do about it now, and — to its credit — tells you
in the same breath that the fix has a bill. If you're loading a large catalogue
into every conversation, read that page before you build a discovery layer, and
check where your provider's cache breakpoint sits before you decide the token
maths is in your favour.</p>
<hr>
<p>📖 <strong>Sources:</strong> <a href="https://blog.modelcontextprotocol.io/posts/mcp-roadmap/" target="_blank" rel="noopener noreferrer" class="">Model Context Protocol Blog — The New MCP Roadmap</a> · <a href="https://modelcontextprotocol.io/docs/2026-07-28/develop/clients/client-best-practices" target="_blank" rel="noopener noreferrer" class="">MCP Docs — Client Best Practices</a> · <a href="https://blog.modelcontextprotocol.io/posts/2026-mcp-roadmap/" target="_blank" rel="noopener noreferrer" class="">Model Context Protocol Blog — The 2026 MCP Roadmap</a> · <a href="https://news.ycombinator.com/item?id=49399591" target="_blank" rel="noopener noreferrer" class="">Hacker News — New MCP Roadmap</a></p>]]></content>
        <author>
            <name>Rafael Fernandes</name>
            <uri>https://www.linkedin.com/in/rafaelmacariofernandes/</uri>
        </author>
        <category label="ai-news" term="ai-news"/>
        <category label="mcp" term="mcp"/>
        <category label="agents" term="agents"/>
        <category label="protocols" term="protocols"/>
        <category label="context-engineering" term="context-engineering"/>
        <category label="infrastructure" term="infrastructure"/>
    </entry>
    <entry>
        <title type="html"><![CDATA[LangChain Just Made MCP First-Class — We Ran It the Same Day, and Three Things Don't Work Yet]]></title>
        <id>https://development-wec.wiline.com/docs/news/langchain-mcp-first-class/</id>
        <link href="https://development-wec.wiline.com/docs/news/langchain-mcp-first-class/"/>
        <updated>2026-09-03T00:00:00.000Z</updated>
        <summary type="html"><![CDATA[MCP support moved into the langchain package today, built on FastMCP, with elicitation as a LangGraph interrupt and a cacheable tool catalog. We installed it the same afternoon and pointed it at a server we wrote against the new spec. The announcement is accurate; the ecosystem around it hasn't caught up, and one of the gaps turns a human-approval gate into a sentence the model makes up.]]></summary>
        <content type="html"><![CDATA[<figure class="newsHero newsHero--image"><span class="newsHero__chip">Agent Frameworks · AI News</span><img src="https://development-wec.wiline.com/docs/img/news/6a99a7d831c3788cf2516870_110.png" alt="MCP support lands in the main LangChain package" loading="eager"></figure>
<p>In July <a class="" href="https://development-wec.wiline.com/docs/news/mcp-2026-07-28-spec/">the MCP specification ripped out sessions</a>. We
wrote at the time that the change was infrastructure, not changelog — that a
stateless core would let servers sit behind ordinary load balancers and survive
redeploys, and that clients would need to catch up.</p>
<p>Today LangChain caught up. MCP support moved out of the separate
<code>langchain-mcp-adapters</code> package and into <code>langchain</code> itself, rebuilt on FastMCP,
with two features the old spec made impossible: <strong>elicitation as a LangGraph
interrupt</strong>, and a <strong>cacheable tool catalog</strong>.</p>
<p>We installed it the same afternoon and pointed it at a server we wrote against the
new spec. The announcement is accurate. The ecosystem around it is not ready — and
one of the gaps quietly converts a human-approval gate into a sentence the model
invents.</p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="what-actually-shipped">What actually shipped<a href="https://development-wec.wiline.com/docs/news/langchain-mcp-first-class/#what-actually-shipped" class="hash-link" aria-label="Direct link to What actually shipped" title="Direct link to What actually shipped" translate="no">​</a></h2>
<p><code>pip install "langchain[mcp]"</code>, requiring <strong>1.4.0 or newer</strong>, and in beta — it says
so on import, which you can see in our captures further down. Python today,
TypeScript "soon to follow."</p>
<ul>
<li class=""><strong>One class, in the main package.</strong> <code>MultiServerMCPClient</code> collapses into
<code>MCPAdapter</code>. <code>async with MCPAdapter(url) as adapter:</code> then
<code>await adapter.list_tools()</code>, and what comes back are ordinary LangChain tools
that go anywhere tools go.</li>
<li class=""><strong>Built on FastMCP</strong>, so transports, auth (bearer, OAuth 2.1, machine-to-machine,
CIMD, or any <code>httpx2.Auth</code>), connection management and protocol negotiation come
from the client underneath. MCP now has two <strong>eras</strong> — the 2025-11-25 handshake
protocol and the 2026-07-28 stateless one — and the FastMCP client picks one
<strong>per connection</strong>: "it tries the new protocol and falls back to the handshake for
a server that hasn't upgraded."</li>
<li class=""><strong>Elicitation via interrupts.</strong> When a tool on the MCP server can't finish without
asking a human, your agent's run pauses as a LangGraph <code>interrupt()</code>, and you resume
it with a structured answer. This is only possible because the stateless spec
turned a mid-call question into a retry-able round rather than something held
open on a socket.</li>
<li class=""><strong>Client-side caching.</strong> <code>fastmcp.Client(url, cache=True)</code> honours the server's
freshness hints so the tool catalog isn't re-fetched every run.</li>
<li class=""><strong><code>ClientGroup</code></strong> for several servers at once, each keeping its own era and
credentials, with tool names prefixed by server — <code>billing_search</code> and
<code>docs_search</code> stay distinct.</li>
</ul>
<p>The framing in the announcement is the same one we used in July: <em>"a redeploy no
longer kills live sessions, because there are none."</em></p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="what-the-stateless-spec-actually-removed">What the stateless spec actually removed<a href="https://development-wec.wiline.com/docs/news/langchain-mcp-first-class/#what-the-stateless-spec-actually-removed" class="hash-link" aria-label="Direct link to What the stateless spec actually removed" title="Direct link to What the stateless spec actually removed" translate="no">​</a></h2>
<p>One tool call, before and after, is the clearest picture of what changed — so here it
is, with a third panel for what we actually got:</p>
<p><span class="zoomImage__wrap"><img alt="Three sequence diagrams comparing a tool call. Under the 2025-11-25 handshake era the agent sends initialize, receives a session id, then sends tools/call with that id. Under the 2026-07-28 stateless spec there is no handshake and the agent sends tools/call directly. Against a default FastMCP server on the new spec, the same request is refused with Bad Request: Missing session ID." src="data:image/svg+xml;base64,PHN2ZyB4bWxucz0iaHR0cDovL3d3dy53My5vcmcvMjAwMC9zdmciIHZpZXdCb3g9IjAgMCA5MDAgNjkwIiB3aWR0aD0iOTAwIiBoZWlnaHQ9IjY5MCIgZm9udC1mYW1pbHk9InN5c3RlbS11aSwtYXBwbGUtc3lzdGVtLFNlZ29lIFVJLHNhbnMtc2VyaWYiPgo8dGl0bGU+Q2FsbGluZyBvbmUgdG9vbDogdGhlIGhhbmRzaGFrZSBlcmEsIHRoZSBzdGF0ZWxlc3Mgc3BlYywgYW5kIGEgZGVmYXVsdCBGYXN0TUNQIHNlcnZlciB0aGF0IHN0aWxsIHdhbnRzIGEgc2Vzc2lvbiBJRDwvdGl0bGU+CjxzdHlsZT4KICAuaW5rICAgeyBmaWxsOiAjMGYxNzJhOyB9CiAgLm11dGVkIHsgZmlsbDogIzY0NzQ4YjsgfQogIC5hY2NlbnR7IGZpbGw6ICMyNTYzZWI7IH0KICAuYmFkICAgeyBmaWxsOiAjZGMyNjI2OyB9CiAgLnBhbmVsIHsgZmlsbDogI2Y4ZmFmYzsgc3Ryb2tlOiAjY2JkNWUxOyB9CiAgLmJveCAgIHsgZmlsbDogI2ZmZmZmZjsgc3Ryb2tlOiAjOTRhM2I4OyB9CiAgLmJveGhpIHsgZmlsbDogI2ZmZmZmZjsgc3Ryb2tlOiAjMjU2M2ViOyB9CiAgLmJveGJhZHsgZmlsbDogI2ZmZmZmZjsgc3Ryb2tlOiAjZGMyNjI2OyB9CiAgLmxpZmUgIHsgc3Ryb2tlOiAjOTRhM2I4OyB9CiAgLmFycm93IHsgc3Ryb2tlOiAjMjU2M2ViOyB9CiAgLmFycm93YmFkIHsgc3Ryb2tlOiAjZGMyNjI2OyB9CiAgLm1vbm8gIHsgZm9udC1mYW1pbHk6IHVpLW1vbm9zcGFjZSxTRk1vbm8tUmVndWxhcixNZW5sbyxDb25zb2xhcyxtb25vc3BhY2U7IH0KICBAbWVkaWEgKHByZWZlcnMtY29sb3Itc2NoZW1lOiBkYXJrKSB7CiAgICAuaW5rICAgeyBmaWxsOiAjZTJlOGYwOyB9CiAgICAubXV0ZWQgeyBmaWxsOiAjOTRhM2I4OyB9CiAgICAuYWNjZW50eyBmaWxsOiAjNjBhNWZhOyB9CiAgICAuYmFkICAgeyBmaWxsOiAjZjg3MTcxOyB9CiAgICAucGFuZWwgeyBmaWxsOiAjMGYxNzJhOyBzdHJva2U6ICMzMzQxNTU7IH0KICAgIC5ib3ggICB7IGZpbGw6ICMxZTI5M2I7IHN0cm9rZTogIzY0NzQ4YjsgfQogICAgLmJveGhpIHsgZmlsbDogIzFlMjkzYjsgc3Ryb2tlOiAjNjBhNWZhOyB9CiAgICAuYm94YmFkeyBmaWxsOiAjMWUyOTNiOyBzdHJva2U6ICNmODcxNzE7IH0KICAgIC5saWZlICB7IHN0cm9rZTogIzY0NzQ4YjsgfQogICAgLmFycm93IHsgc3Ryb2tlOiAjNjBhNWZhOyB9CiAgICAuYXJyb3diYWQgeyBzdHJva2U6ICNmODcxNzE7IH0KICB9Cjwvc3R5bGU+CjxkZWZzPgogIDxtYXJrZXIgaWQ9ImFoIiB2aWV3Qm94PSIwIDAgMTAgMTAiIHJlZlg9IjkiIHJlZlk9IjUiIG1hcmtlcldpZHRoPSI3IiBtYXJrZXJIZWlnaHQ9IjciIG9yaWVudD0iYXV0by1zdGFydC1yZXZlcnNlIj4KICAgIDxwYXRoIGQ9Ik0wLDEgTDksNSBMMCw5IHoiIGZpbGw9IiMyNTYzZWIiLz4KICA8L21hcmtlcj4KICA8bWFya2VyIGlkPSJhaGJhZCIgdmlld0JveD0iMCAwIDEwIDEwIiByZWZYPSI5IiByZWZZPSI1IiBtYXJrZXJXaWR0aD0iNyIgbWFya2VySGVpZ2h0PSI3IiBvcmllbnQ9ImF1dG8tc3RhcnQtcmV2ZXJzZSI+CiAgICA8cGF0aCBkPSJNMCwxIEw5LDUgTDAsOSB6IiBmaWxsPSIjZGMyNjI2Ii8+CiAgPC9tYXJrZXI+CjwvZGVmcz4KCjx0ZXh0IHg9IjI0IiB5PSIzNCIgY2xhc3M9ImluayIgZm9udC1zaXplPSIyMyIgZm9udC13ZWlnaHQ9IjYwMCI+Q2FsbGluZyBvbmUgdG9vbDwvdGV4dD4KPHRleHQgeD0iMjQiIHk9IjU4IiBjbGFzcz0ibXV0ZWQgbW9ubyIgZm9udC1zaXplPSIxNC41Ij53aGF0IGl0IHRha2VzIHRvIHJlYWNoIHRoZSBzZXJ2ZXI8L3RleHQ+Cgo8IS0tID09PT09PT09PT09PT09PT09PT09PSBQYW5lbCAxIOKAlCBoYW5kc2hha2UgZXJhID09PT09PT09PT09PT09PT09PT09PSAtLT4KPHJlY3QgeD0iMTYiIHk9IjgwIiB3aWR0aD0iODY4IiBoZWlnaHQ9IjIxMiIgcng9IjEyIiBjbGFzcz0icGFuZWwiLz4KPHRleHQgeD0iNDAiIHk9IjExMiIgY2xhc3M9ImFjY2VudCBtb25vIiBmb250LXNpemU9IjE1Ij4yMDI1LTExLTI1PC90ZXh0Pgo8dGV4dCB4PSI0MCIgeT0iMTM0IiBjbGFzcz0iaW5rIiBmb250LXNpemU9IjE0LjUiIGZvbnQtd2VpZ2h0PSI2MDAiPmEgaGFuZHNoYWtlIGJlZm9yZSBhbnkgd29yazwvdGV4dD4KPHRleHQgeD0iNDAiIHk9IjE1OCIgY2xhc3M9Im11dGVkIiBmb250LXNpemU9IjEzIj50aGUgc2Vzc2lvbiBwaW5zIHRoZSBhZ2VudCB0byB0aGU8L3RleHQ+Cjx0ZXh0IHg9IjQwIiB5PSIxNzYiIGNsYXNzPSJtdXRlZCIgZm9udC1zaXplPSIxMyI+b25lIGluc3RhbmNlIHRoYXQgaXNzdWVkIGl0PC90ZXh0PgoKPHJlY3QgeD0iMzkyIiB5PSIxMDAiIHdpZHRoPSIxMDQiIGhlaWdodD0iMzYiIHJ4PSI4IiBjbGFzcz0iYm94Ii8+Cjx0ZXh0IHg9IjQ0NCIgeT0iMTI0IiBjbGFzcz0iaW5rIG1vbm8iIGZvbnQtc2l6ZT0iMTQiIHRleHQtYW5jaG9yPSJtaWRkbGUiPmFnZW50PC90ZXh0Pgo8cmVjdCB4PSI3NTIiIHk9IjEwMCIgd2lkdGg9IjEwNCIgaGVpZ2h0PSIzNiIgcng9IjgiIGNsYXNzPSJib3giLz4KPHRleHQgeD0iODA0IiB5PSIxMjQiIGNsYXNzPSJpbmsgbW9ubyIgZm9udC1zaXplPSIxNCIgdGV4dC1hbmNob3I9Im1pZGRsZSI+c2VydmVyPC90ZXh0PgoKPGxpbmUgeDE9IjQ0NCIgeTE9IjEzNiIgeDI9IjQ0NCIgeTI9IjI4MiIgY2xhc3M9ImxpZmUiIHN0cm9rZS1kYXNoYXJyYXk9IjMgNCIvPgo8bGluZSB4MT0iODA0IiB5MT0iMTM2IiB4Mj0iODA0IiB5Mj0iMjgyIiBjbGFzcz0ibGlmZSIgc3Ryb2tlLWRhc2hhcnJheT0iMyA0Ii8+Cgo8dGV4dCB4PSI2MjQiIHk9IjE2MyIgY2xhc3M9ImluayBtb25vIiBmb250LXNpemU9IjEzLjUiIHRleHQtYW5jaG9yPSJtaWRkbGUiPmluaXRpYWxpemU8L3RleHQ+CjxsaW5lIHgxPSI0NDQiIHkxPSIxNzIiIHgyPSI3OTgiIHkyPSIxNzIiIGNsYXNzPSJhcnJvdyIgc3Ryb2tlLXdpZHRoPSIxLjkiIG1hcmtlci1lbmQ9InVybCgjYWgpIi8+Cgo8dGV4dCB4PSI2MjQiIHk9IjE5OSIgY2xhc3M9ImluayBtb25vIiBmb250LXNpemU9IjEzLjUiIHRleHQtYW5jaG9yPSJtaWRkbGUiPnNlc3Npb24gaWQ8L3RleHQ+CjxsaW5lIHgxPSI4MDQiIHkxPSIyMDgiIHgyPSI0NTAiIHkyPSIyMDgiIGNsYXNzPSJhcnJvdyIgc3Ryb2tlLXdpZHRoPSIxLjkiIHN0cm9rZS1kYXNoYXJyYXk9IjYgNSIgbWFya2VyLWVuZD0idXJsKCNhaCkiLz4KCjx0ZXh0IHg9IjYyNCIgeT0iMjM1IiBjbGFzcz0iaW5rIG1vbm8iIGZvbnQtc2l6ZT0iMTMuNSIgdGV4dC1hbmNob3I9Im1pZGRsZSI+dG9vbHMvY2FsbCArIHNlc3Npb24gaWQ8L3RleHQ+CjxsaW5lIHgxPSI0NDQiIHkxPSIyNDQiIHgyPSI3OTgiIHkyPSIyNDQiIGNsYXNzPSJhcnJvdyIgc3Ryb2tlLXdpZHRoPSIxLjkiIG1hcmtlci1lbmQ9InVybCgjYWgpIi8+Cgo8dGV4dCB4PSI2MjQiIHk9IjI3MSIgY2xhc3M9ImluayBtb25vIiBmb250LXNpemU9IjEzLjUiIHRleHQtYW5jaG9yPSJtaWRkbGUiPnJlc3VsdDwvdGV4dD4KPGxpbmUgeDE9IjgwNCIgeTE9IjI4MCIgeDI9IjQ1MCIgeTI9IjI4MCIgY2xhc3M9ImFycm93IiBzdHJva2Utd2lkdGg9IjEuOSIgc3Ryb2tlLWRhc2hhcnJheT0iNiA1IiBtYXJrZXItZW5kPSJ1cmwoI2FoKSIvPgoKPCEtLSA9PT09PT09PT09PT09PT09PT09PT0gUGFuZWwgMiDigJQgc3RhdGVsZXNzIHNwZWMgPT09PT09PT09PT09PT09PT09PT09IC0tPgo8cmVjdCB4PSIxNiIgeT0iMzA0IiB3aWR0aD0iODY4IiBoZWlnaHQ9IjE2NCIgcng9IjEyIiBjbGFzcz0icGFuZWwiLz4KPHRleHQgeD0iNDAiIHk9IjMzNiIgY2xhc3M9ImFjY2VudCBtb25vIiBmb250LXNpemU9IjE1Ij4yMDI2LTA3LTI4PC90ZXh0Pgo8dGV4dCB4PSI0MCIgeT0iMzU4IiBjbGFzcz0iaW5rIiBmb250LXNpemU9IjE0LjUiIGZvbnQtd2VpZ2h0PSI2MDAiPm9uZSByZXF1ZXN0LCBzdHJhaWdodCB0byB3b3JrPC90ZXh0Pgo8dGV4dCB4PSI0MCIgeT0iMzgyIiBjbGFzcz0ibXV0ZWQiIGZvbnQtc2l6ZT0iMTMiPml0IGNhcnJpZXMgaXRzIG93biB2ZXJzaW9uIGFuZDwvdGV4dD4KPHRleHQgeD0iNDAiIHk9IjQwMCIgY2xhc3M9Im11dGVkIiBmb250LXNpemU9IjEzIj5pZGVudGl0eSwgc28gYW55IGluc3RhbmNlIGFuc3dlcnM8L3RleHQ+Cgo8cmVjdCB4PSIzOTIiIHk9IjMyNCIgd2lkdGg9IjEwNCIgaGVpZ2h0PSIzNiIgcng9IjgiIGNsYXNzPSJib3hoaSIvPgo8dGV4dCB4PSI0NDQiIHk9IjM0OCIgY2xhc3M9ImluayBtb25vIiBmb250LXNpemU9IjE0IiB0ZXh0LWFuY2hvcj0ibWlkZGxlIj5hZ2VudDwvdGV4dD4KPHJlY3QgeD0iNzUyIiB5PSIzMjQiIHdpZHRoPSIxMDQiIGhlaWdodD0iMzYiIHJ4PSI4IiBjbGFzcz0iYm94aGkiLz4KPHRleHQgeD0iODA0IiB5PSIzNDgiIGNsYXNzPSJpbmsgbW9ubyIgZm9udC1zaXplPSIxNCIgdGV4dC1hbmNob3I9Im1pZGRsZSI+c2VydmVyPC90ZXh0PgoKPGxpbmUgeDE9IjQ0NCIgeTE9IjM2MCIgeDI9IjQ0NCIgeTI9IjQ1OCIgY2xhc3M9ImxpZmUiIHN0cm9rZS1kYXNoYXJyYXk9IjMgNCIvPgo8bGluZSB4MT0iODA0IiB5MT0iMzYwIiB4Mj0iODA0IiB5Mj0iNDU4IiBjbGFzcz0ibGlmZSIgc3Ryb2tlLWRhc2hhcnJheT0iMyA0Ii8+Cgo8dGV4dCB4PSI2MjQiIHk9IjM4NiIgY2xhc3M9Im11dGVkIG1vbm8iIGZvbnQtc2l6ZT0iMTMuNSIgdGV4dC1hbmNob3I9Im1pZGRsZSI+bm8gaGFuZHNoYWtlPC90ZXh0PgoKPHRleHQgeD0iNjI0IiB5PSI0MTUiIGNsYXNzPSJpbmsgbW9ubyIgZm9udC1zaXplPSIxMy41IiB0ZXh0LWFuY2hvcj0ibWlkZGxlIj50b29scy9jYWxsPC90ZXh0Pgo8bGluZSB4MT0iNDQ0IiB5MT0iNDI0IiB4Mj0iNzk4IiB5Mj0iNDI0IiBjbGFzcz0iYXJyb3ciIHN0cm9rZS13aWR0aD0iMS45IiBtYXJrZXItZW5kPSJ1cmwoI2FoKSIvPgoKPHRleHQgeD0iNjI0IiB5PSI0NDciIGNsYXNzPSJpbmsgbW9ubyIgZm9udC1zaXplPSIxMy41IiB0ZXh0LWFuY2hvcj0ibWlkZGxlIj5yZXN1bHQ8L3RleHQ+CjxsaW5lIHgxPSI4MDQiIHkxPSI0NTYiIHgyPSI0NTAiIHkyPSI0NTYiIGNsYXNzPSJhcnJvdyIgc3Ryb2tlLXdpZHRoPSIxLjkiIHN0cm9rZS1kYXNoYXJyYXk9IjYgNSIgbWFya2VyLWVuZD0idXJsKCNhaCkiLz4KCjwhLS0gPT09PT09PT09PT09PT09PT09PT09IFBhbmVsIDMg4oCUIHdoYXQgd2UgYWN0dWFsbHkgZ290ID09PT09PT09PT09PT09PT09PT09PSAtLT4KPHJlY3QgeD0iMTYiIHk9IjQ4MCIgd2lkdGg9Ijg2OCIgaGVpZ2h0PSIxODgiIHJ4PSIxMiIgY2xhc3M9InBhbmVsIi8+Cjx0ZXh0IHg9IjQwIiB5PSI1MTIiIGNsYXNzPSJiYWQgbW9ubyIgZm9udC1zaXplPSIxNSI+MjAyNi0wNy0yODwvdGV4dD4KPHRleHQgeD0iNDAiIHk9IjUzMiIgY2xhc3M9Im11dGVkIiBmb250LXNpemU9IjEyLjUiPuKApmFnYWluc3QgYSBkZWZhdWx0IEZhc3RNQ1Agc2VydmVyPC90ZXh0Pgo8dGV4dCB4PSI0MCIgeT0iNTU4IiBjbGFzcz0iYmFkIiBmb250LXNpemU9IjE0LjUiIGZvbnQtd2VpZ2h0PSI2MDAiPnRoZSBjbGllbnQgbW92ZWQgb24uIHRoZSBzZXJ2ZXIgZGlkbid0LjwvdGV4dD4KPHRleHQgeD0iNDAiIHk9IjU4MiIgY2xhc3M9Im11dGVkIG1vbm8iIGZvbnQtc2l6ZT0iMTMiPnN0YXRlbGVzc19odHRwPVRydWU8L3RleHQ+Cjx0ZXh0IHg9IjQwIiB5PSI2MDAiIGNsYXNzPSJtdXRlZCIgZm9udC1zaXplPSIxMyI+aXMgb3B0LWluLCBhbmQgb2ZmIGJ5IGRlZmF1bHQ8L3RleHQ+Cgo8cmVjdCB4PSIzOTIiIHk9IjUwMCIgd2lkdGg9IjEwNCIgaGVpZ2h0PSIzNiIgcng9IjgiIGNsYXNzPSJib3hoaSIvPgo8dGV4dCB4PSI0NDQiIHk9IjUyNCIgY2xhc3M9ImluayBtb25vIiBmb250LXNpemU9IjE0IiB0ZXh0LWFuY2hvcj0ibWlkZGxlIj5hZ2VudDwvdGV4dD4KPHJlY3QgeD0iNzUyIiB5PSI1MDAiIHdpZHRoPSIxMDQiIGhlaWdodD0iMzYiIHJ4PSI4IiBjbGFzcz0iYm94YmFkIi8+Cjx0ZXh0IHg9IjgwNCIgeT0iNTI0IiBjbGFzcz0iaW5rIG1vbm8iIGZvbnQtc2l6ZT0iMTQiIHRleHQtYW5jaG9yPSJtaWRkbGUiPnNlcnZlcjwvdGV4dD4KCjxsaW5lIHgxPSI0NDQiIHkxPSI1MzYiIHgyPSI0NDQiIHkyPSI2NTAiIGNsYXNzPSJsaWZlIiBzdHJva2UtZGFzaGFycmF5PSIzIDQiLz4KPGxpbmUgeDE9IjgwNCIgeTE9IjUzNiIgeDI9IjgwNCIgeTI9IjY1MCIgY2xhc3M9ImxpZmUiIHN0cm9rZS1kYXNoYXJyYXk9IjMgNCIvPgoKPHRleHQgeD0iNjI0IiB5PSI1NjIiIGNsYXNzPSJtdXRlZCBtb25vIiBmb250LXNpemU9IjEzLjUiIHRleHQtYW5jaG9yPSJtaWRkbGUiPm5vIGhhbmRzaGFrZTwvdGV4dD4KCjx0ZXh0IHg9IjYyNCIgeT0iNTkxIiBjbGFzcz0iaW5rIG1vbm8iIGZvbnQtc2l6ZT0iMTMuNSIgdGV4dC1hbmNob3I9Im1pZGRsZSI+dG9vbHMvY2FsbDwvdGV4dD4KPGxpbmUgeDE9IjQ0NCIgeTE9IjYwMCIgeDI9Ijc5OCIgeTI9IjYwMCIgY2xhc3M9ImFycm93IiBzdHJva2Utd2lkdGg9IjEuOSIgbWFya2VyLWVuZD0idXJsKCNhaCkiLz4KCjx0ZXh0IHg9IjYyNCIgeT0iNjI3IiBjbGFzcz0iYmFkIG1vbm8iIGZvbnQtc2l6ZT0iMTMuNSIgdGV4dC1hbmNob3I9Im1pZGRsZSI+QmFkIFJlcXVlc3Q6IE1pc3Npbmcgc2Vzc2lvbiBJRDwvdGV4dD4KPGxpbmUgeDE9IjgwNCIgeTE9IjYzNiIgeDI9IjQ1MCIgeTI9IjYzNiIgY2xhc3M9ImFycm93YmFkIiBzdHJva2Utd2lkdGg9IjEuOSIgc3Ryb2tlLWRhc2hhcnJheT0iNiA1IiBtYXJrZXItZW5kPSJ1cmwoI2FoYmFkKSIvPgo8dGV4dCB4PSI2MjQiIHk9IjY1NiIgY2xhc3M9Im11dGVkIG1vbm8iIGZvbnQtc2l6ZT0iMTIiIHRleHQtYW5jaG9yPSJtaWRkbGUiPkhUVFAgNDAwIMK3IEpTT04tUlBDIOKIkjMyNjAwPC90ZXh0Pgo8L3N2Zz4K" width="900" height="690" class="zoomImage " loading="lazy"><span class="zoomImage__badge" aria-hidden="true"><svg viewBox="0 0 24 24" width="16" height="16" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round"><circle cx="11" cy="11" r="7"></circle><path d="M21 21l-4.3-4.3"></path><path d="M11 8v6M8 11h6"></path></svg></span></span></p>
<p>The first two panels are the pitch, and the pitch is real. The third is what a
brand-new server does on the afternoon the client ships.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="three-things-that-dont-work-yet">Three things that don't work yet<a href="https://development-wec.wiline.com/docs/news/langchain-mcp-first-class/#three-things-that-dont-work-yet" class="hash-link" aria-label="Direct link to Three things that don't work yet" title="Direct link to Three things that don't work yet" translate="no">​</a></h2>
<p>We built an MCP server exposing a small booking-and-invoicing database, pointed a
LangChain agent at it, and gated a refund behind a human. (The agent's own model runs
on a self-hosted gateway — that sits between the agent and the LLM, not between the
agent and MCP.) It works — that's
<a class="" href="https://development-wec.wiline.com/docs/tutorials/mcp-governed-tools-agent-elicitation/">the tutorial</a>.
Getting there surfaced three gaps between what's written and what runs.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="1-the-spec-is-stateless-your-server-isnt-by-default">1. The spec is stateless. Your server isn't, by default.<a href="https://development-wec.wiline.com/docs/news/langchain-mcp-first-class/#1-the-spec-is-stateless-your-server-isnt-by-default" class="hash-link" aria-label="Direct link to 1. The spec is stateless. Your server isn't, by default." title="Direct link to 1. The spec is stateless. Your server isn't, by default." translate="no">​</a></h3>
<p>The first request to a brand-new FastMCP 4.0.2 server — here a plain <code>curl</code> asking
for <code>tools/list</code>, so nothing client-side can be blamed for it:</p>
<p><span class="zoomImage__wrap"><img alt="A curl POST to the MCP endpoint returning Bad Request: Missing session ID with JSON-RPC error code -32600" src="https://development-wec.wiline.com/docs/assets/images/mcp-session-error-7897aebee33998b591bd27b4606a6e6b.png" width="666" height="75" class="zoomImage " loading="lazy"><span class="zoomImage__badge" aria-hidden="true"><svg viewBox="0 0 24 24" width="16" height="16" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round"><circle cx="11" cy="11" r="7"></circle><path d="M21 21l-4.3-4.3"></path><path d="M11 8v6M8 11h6"></path></svg></span></span></p>
<p>An error about a concept the specification deleted five weeks ago. The client is
stateless; the server still defaults to the stateful transport.</p>
<p>The fix isn't a header, a client option, or anything you send. It's an argument on
the call that <strong>starts your server</strong> — the last line of the server file, where <code>mcp</code>
is your <code>FastMCP</code> instance:</p>
<div class="language-python codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#393A34;--prism-background-color:#f6f8fa"><div class="codeBlockTitle_OeMC">office_tools.py</div><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-python codeBlock_bY9V thin-scrollbar" style="color:#393A34;background-color:#f6f8fa"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#393A34"><span class="token keyword" style="color:#00009f">from</span><span class="token plain"> fastmcp </span><span class="token keyword" style="color:#00009f">import</span><span class="token plain"> FastMCP</span><br></div><div class="token-line" style="color:#393A34"><span class="token plain" style="display:inline-block"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">mcp </span><span class="token operator" style="color:#393A34">=</span><span class="token plain"> FastMCP</span><span class="token punctuation" style="color:#393A34">(</span><span class="token string" style="color:#e3116c">"office-tools"</span><span class="token punctuation" style="color:#393A34">)</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain" style="display:inline-block"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain"></span><span class="token comment" style="color:#999988;font-style:italic"># ... your @mcp.tool functions ...</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain" style="display:inline-block"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain"></span><span class="token keyword" style="color:#00009f">if</span><span class="token plain"> __name__ </span><span class="token operator" style="color:#393A34">==</span><span class="token plain"> </span><span class="token string" style="color:#e3116c">"__main__"</span><span class="token punctuation" style="color:#393A34">:</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">    </span><span class="token comment" style="color:#999988;font-style:italic"># Without stateless_http=True this server answers a 2026-07-28 client</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">    </span><span class="token comment" style="color:#999988;font-style:italic"># with "Bad Request: Missing session ID".</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">    mcp</span><span class="token punctuation" style="color:#393A34">.</span><span class="token plain">run</span><span class="token punctuation" style="color:#393A34">(</span><span class="token plain">transport</span><span class="token operator" style="color:#393A34">=</span><span class="token string" style="color:#e3116c">"http"</span><span class="token punctuation" style="color:#393A34">,</span><span class="token plain"> host</span><span class="token operator" style="color:#393A34">=</span><span class="token string" style="color:#e3116c">"127.0.0.1"</span><span class="token punctuation" style="color:#393A34">,</span><span class="token plain"> port</span><span class="token operator" style="color:#393A34">=</span><span class="token number" style="color:#36acaa">8770</span><span class="token punctuation" style="color:#393A34">,</span><span class="token plain"> stateless_http</span><span class="token operator" style="color:#393A34">=</span><span class="token boolean" style="color:#36acaa">True</span><span class="token punctuation" style="color:#393A34">)</span><br></div></code></pre></div></div>
<p>Nothing in the announcement says so — reasonably enough, since the post is about the
client. But it means the client half of the upgrade is one <code>pip install</code>, and the
server half is a flag you have to know exists.</p>
<p>Two flags exist, and <code>run_http_async</code> (which <code>mcp.run</code> calls for HTTP) accepts both.
From its own docstring:</p>
<div class="language-text codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#393A34;--prism-background-color:#f6f8fa"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-text codeBlock_bY9V thin-scrollbar" style="color:#393A34;background-color:#f6f8fa"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#393A34"><span class="token plain">stateless_http: Whether to use stateless HTTP (defaults to settings.stateless_http)</span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">stateless: Alias for stateless_http for CLI consistency</span><br></div></code></pre></div></div>
<p>One switch, two names, and the setting it defaults to is off.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="2-ctxelicit-is-the-old-api--and-the-error-goes-to-the-model-not-to-you">2. <code>ctx.elicit</code> is the old API — and the error goes to the model, not to you<a href="https://development-wec.wiline.com/docs/news/langchain-mcp-first-class/#2-ctxelicit-is-the-old-api--and-the-error-goes-to-the-model-not-to-you" class="hash-link" aria-label="Direct link to 2-ctxelicit-is-the-old-api--and-the-error-goes-to-the-model-not-to-you" title="Direct link to 2-ctxelicit-is-the-old-api--and-the-error-goes-to-the-model-not-to-you" translate="no">​</a></h3>
<p>Every elicitation example you can find calls <code>await ctx.elicit(message, response_type)</code>
inside the tool body, where <code>ctx</code> is the <code>Context</code> object FastMCP passes to your tool.
That is the <strong>handshake-era</strong> mechanism: it blocks mid-execution and
speaks over the session's back-channel — the back-channel a stateless connection
doesn't have. FastMCP's own docs are explicit that it's for connections
<code>≤ 2025-11-25</code>, and promise that calling it on a modern one "raises a clear era
error rather than failing obscurely."</p>
<p>It does raise one. That isn't the problem. We rebuilt the refund tool around
<code>ctx.elicit</code> and ran the same agent against it:</p>
<p><span class="zoomImage__wrap"><img alt="The agent run printing interrupt raised? False, three tool messages where issue_refund has status=error carrying &amp;#39;elicitation via server-initiated requests is unavailable on 2026-07-28 connections&amp;#39;, the model replying that it has initiated the refund and a human must approve it, and a sqlite3 query showing invoice 2 still open" src="https://development-wec.wiline.com/docs/assets/images/mcp-elicit-old-api-ec02d789aa6970b70a130f83067c235b.png" width="1038" height="437" class="zoomImage " loading="lazy"><span class="zoomImage__badge" aria-hidden="true"><svg viewBox="0 0 24 24" width="16" height="16" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round"><circle cx="11" cy="11" r="7"></circle><path d="M21 21l-4.3-4.3"></path><path d="M11 8v6M8 11h6"></path></svg></span></span></p>
<p>Read that in order. FastMCP raised the era error, exactly as documented. LangChain
caught it and turned it into a <code>ToolMessage</code> with <code>status="error"</code>. That is deliberate:
<code>langchain/mcp/tools.py</code> wires a handler whose docstring says it exists to hand the
server's own error detail to the model "instead of ending the run." The model read
that error and told the user:</p>
<blockquote>
<p>I have located Maria Alvarez and identified her open invoice (ID: 2) for $80.00. I
have <strong>initiated the refund request</strong> for this invoice. Please note that <strong>a human
must approve</strong> the amount and provide a reason to complete the refund.</p>
</blockquote>
<p>No interrupt was raised. The process exited 0. Nothing is pending, nothing is
waiting for a human, and no one will ever be asked — and the invoice is still <code>open</code>,
so the refund didn't happen either. What you get is not a destructive action slipping
past a gate; it's an approval workflow that silently doesn't exist, described in
fluent English by a model that read the error and paraphrased it as progress.</p>
<p>The fix is that the modern pattern has the tool <strong>return</strong> an <code>InputRequiredResult</code>
describing what it needs, and exit. Your agent surfaces that as the LangGraph interrupt,
a human answers, and the MCP client re-issues the same <code>tools/call</code> with the answer
attached. But the failure mode is the story: an era error is a fine thing to raise
into a program, and a terrible thing to hand to a language model that is rewarded for
sounding helpful.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="3-cachetrue-is-necessary-not-sufficient">3. <code>cache=True</code> is necessary, not sufficient<a href="https://development-wec.wiline.com/docs/news/langchain-mcp-first-class/#3-cachetrue-is-necessary-not-sufficient" class="hash-link" aria-label="Direct link to 3-cachetrue-is-necessary-not-sufficient" title="Direct link to 3-cachetrue-is-necessary-not-sufficient" translate="no">​</a></h3>
<p>The client cache respects the <code>ttlMs</code> and <code>cacheScope</code> hints a <strong>server</strong> attaches to
its <code>tools/list</code> response, and only against modern-era servers that send them. A default
FastMCP server sends neither — we dumped our own <code>tools/list</code> response and it carries no
cache hints at all. So you turn the cache on, call <code>list_tools(cache_mode="use")</code> twice
back to back, time both, and get:</p>
<p><span class="zoomImage__wrap"><img alt="Two tool discoveries of five tools each, timed at 14.6 ms and 12.5 ms, showing no cache effect" src="https://development-wec.wiline.com/docs/assets/images/mcp-cache-bbadf88f6b26a069ec15e3b3db635c02.png" width="801" height="125" class="zoomImage " loading="lazy"><span class="zoomImage__badge" aria-hidden="true"><svg viewBox="0 0 24 24" width="16" height="16" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round"><circle cx="11" cy="11" r="7"></circle><path d="M21 21l-4.3-4.3"></path><path d="M11 8v6M8 11h6"></path></svg></span></span></p>
<p>You conclude the cache is broken. It isn't — there was nothing to cache. Worth noting
the cache belongs to the <code>fastmcp.Client</code>, not to <code>MCPAdapter</code>, and one client per
caller keeps catalogs from crossing between tenants.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="and-two-smaller-ones-in-the-announcement-itself">And two smaller ones, in the announcement itself<a href="https://development-wec.wiline.com/docs/news/langchain-mcp-first-class/#and-two-smaller-ones-in-the-announcement-itself" class="hash-link" aria-label="Direct link to And two smaller ones, in the announcement itself" title="Direct link to And two smaller ones, in the announcement itself" translate="no">​</a></h2>
<p><strong>The elicitation snippet raises <code>AttributeError</code>.</strong> The post reads the interrupt as
<code>paused["__interrupt__"][0].value.requests[0]</code> — attribute access. In
<code>langchain/mcp/elicitation.py</code> on 1.4.0:</p>
<div class="language-python codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#393A34;--prism-background-color:#f6f8fa"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-python codeBlock_bY9V thin-scrollbar" style="color:#393A34;background-color:#f6f8fa"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#393A34"><span class="token keyword" style="color:#00009f">class</span><span class="token plain"> </span><span class="token class-name">MCPElicitationInterrupt</span><span class="token punctuation" style="color:#393A34">(</span><span class="token plain">TypedDict</span><span class="token punctuation" style="color:#393A34">)</span><span class="token punctuation" style="color:#393A34">:</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">    </span><span class="token builtin">type</span><span class="token punctuation" style="color:#393A34">:</span><span class="token plain"> Literal</span><span class="token punctuation" style="color:#393A34">[</span><span class="token string" style="color:#e3116c">"mcp_elicitation"</span><span class="token punctuation" style="color:#393A34">]</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">    tool_name</span><span class="token punctuation" style="color:#393A34">:</span><span class="token plain"> </span><span class="token builtin">str</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">    requests</span><span class="token punctuation" style="color:#393A34">:</span><span class="token plain"> </span><span class="token builtin">list</span><span class="token punctuation" style="color:#393A34">[</span><span class="token plain">MCPElicitationRequest</span><span class="token punctuation" style="color:#393A34">]</span><br></div></code></pre></div></div>
<p>A <code>TypedDict</code> is a dict at runtime, so <code>.requests</code> doesn't resolve. It's
<code>value["requests"]</code> — which the same snippet gets right a few lines later, reading
<code>question["key"]</code> by subscript.</p>
<p><strong>The elicitation docs link 404s.</strong> The post closes the section by pointing at
<code>docs.langchain.com/oss/python/langchain/mcp/elicitation</code> for "declining a question, and
gating destructive tools behind the same approval flow." That page returns 404 as of
publication, while the parent page and its other children — <code>mcp</code>,
<code>mcp/connections</code>, <code>mcp/tools</code>, <code>mcp/auth</code> — all resolve. It is, inconveniently, the
one page that would have documented the approval flow gap 2 shows falling over.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="why-this-matters-for-you-specifically">Why this matters for you, specifically<a href="https://development-wec.wiline.com/docs/news/langchain-mcp-first-class/#why-this-matters-for-you-specifically" class="hash-link" aria-label="Direct link to Why this matters for you, specifically" title="Direct link to Why this matters for you, specifically" translate="no">​</a></h2>
<p>If you consume MCP servers someone else runs, this is straightforwardly good news and
you'll notice mostly the shorter import path.</p>
<p>If you <strong>write</strong> MCP servers, the July revision moved the ground and the tooling is
still settling. Three things are now yours to get right: your server does not become
stateless because the spec did — you set a flag; a tool that needs human input has to
be written to be <strong>re-entered</strong> rather than resumed — <code>interrupt()</code> unwinds the whole
call, so the tool body runs again from the top when you answer; and any
elicitation example predating August is teaching you an API whose failure lands in the
model's context instead of your logs.</p>
<p>That last one generalises past MCP. As frameworks get better at keeping agents alive
through errors, the class of bug that ends a run is shrinking and the class that gets
narrated to a user is growing. A gate that fails closed is a bug you find in testing. A
gate that was never installed, described by a model as awaiting your approval, is one
you find in an audit.</p>
<p>The direction is right. The stateless core is what makes an agent's human-approval pause
survive a redeploy, and that's a real capability rather than a refactor. Just don't
expect your first request to succeed.</p>
<hr>
<p>📖 <strong>Sources:</strong> <a href="https://www.langchain.com/blog/mcp-in-langchain-stateless-protocol-elicitation-and-more" target="_blank" rel="noopener noreferrer" class="">LangChain — MCP in LangChain: stateless protocol, elicitation, and more</a> · <a href="https://docs.langchain.com/oss/python/langchain/mcp" target="_blank" rel="noopener noreferrer" class="">MCP in LangChain docs</a> · <a href="https://docs.langchain.com/oss/python/migrate/langchain-mcp-adapters" target="_blank" rel="noopener noreferrer" class="">Migrating from langchain-mcp-adapters</a> · <a href="https://gofastmcp.com/clients/client" target="_blank" rel="noopener noreferrer" class="">FastMCP client documentation</a> · <a href="https://gofastmcp.com/servers/elicitation" target="_blank" rel="noopener noreferrer" class="">FastMCP elicitation</a> · <a href="https://modelcontextprotocol.io/specification/2026-07-28" target="_blank" rel="noopener noreferrer" class="">MCP 2026-07-28 specification</a></p>
<p><em>Versions under test: <code>langchain</code> 1.4.0, <code>fastmcp</code> 4.0.2, <code>mcp</code> 2.1.1, model served over a self-hosted gateway. The <code>ctx.elicit</code> capture is a render of real captured output from that run, not a screen grab.</em></p>]]></content>
        <author>
            <name>Rafael Fernandes</name>
            <uri>https://www.linkedin.com/in/rafaelmacariofernandes/</uri>
        </author>
        <category label="ai-news" term="ai-news"/>
        <category label="mcp" term="mcp"/>
        <category label="langchain" term="langchain"/>
        <category label="langgraph" term="langgraph"/>
        <category label="agents" term="agents"/>
        <category label="protocols" term="protocols"/>
    </entry>
    <entry>
        <title type="html"><![CDATA[An agent can narrow its own web search, but never widen it]]></title>
        <id>https://development-wec.wiline.com/docs/news/agent-web-search-domain-allowlist/</id>
        <link href="https://development-wec.wiline.com/docs/news/agent-web-search-domain-allowlist/"/>
        <updated>2026-08-26T00:00:00.000Z</updated>
        <summary type="html"><![CDATA[AWS gave agent web search a list of allowed sites on 19 August. The admin sets one list, the agent can set another, and when they disagree the agent's list can only make the search smaller. That one rule is the difference between a guardrail and a note in the prompt.]]></summary>
        <content type="html"><![CDATA[<div class="newsHero"><div class="newsHero__glow" aria-hidden="true"></div><span class="newsHero__eyebrow">Agents · AI News</span><h2 class="newsHero__title">Narrow only, never wider</h2><div class="newsHero__transition"><span class="newsHero__pill newsHero__pill--from">Please use trusted sources</span><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2.5" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-arrow-right newsHero__arrow" aria-hidden="true"><path d="M5 12h14"></path><path d="m12 5 7 7-7 7"></path></svg><span class="newsHero__pill newsHero__pill--to">Trusted sources are all there are</span></div></div>
<p>Your agent can search the web. You write in the prompt: <em>only use sec.gov and the big
financial wires.</em> It usually listens. The times it does not are the times you learn that
a line in a prompt is a request, not a rule.</p>
<p>On 19 August AWS added site filters to the web search tool in Bedrock AgentCore. Small
feature. But the way it handles a disagreement is worth knowing, because that is what
makes it a guardrail instead of a note.</p>
<!-- -->
<div class="theme-admonition theme-admonition-note admonition_xJq3 alert alert--secondary"><div class="admonitionHeading_Gvgb"><span class="admonitionIcon_Rf37"><svg viewBox="0 0 14 16"><path fill-rule="evenodd" d="M6.3 5.69a.942.942 0 0 1-.28-.7c0-.28.09-.52.28-.7.19-.18.42-.28.7-.28.28 0 .52.09.7.28.18.19.28.42.28.7 0 .28-.09.52-.28.7a1 1 0 0 1-.7.3c-.28 0-.52-.11-.7-.3zM8 7.99c-.02-.25-.11-.48-.31-.69-.2-.19-.42-.3-.69-.31H6c-.27.02-.48.13-.69.31-.2.2-.3.44-.31.69h1v3c.02.27.11.5.31.69.2.2.42.31.69.31h1c.27 0 .48-.11.69-.31.2-.19.3-.42.31-.69H8V7.98v.01zM7 2.3c-3.14 0-5.7 2.54-5.7 5.68 0 3.14 2.56 5.7 5.7 5.7s5.7-2.55 5.7-5.7c0-3.15-2.56-5.69-5.7-5.69v.01zM7 .98c3.86 0 7 3.14 7 7s-3.14 7-7 7-7-3.12-7-7 3.14-7 7-7z"></path></svg></span>Whose product this is</div><div class="admonitionContent_BuS1"><p>This is AWS's own documentation of an AWS service, quoted here. It does not run on our
infrastructure and we have not tested it. The design is what interests us.</p></div></div>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="what-you-get">What you get<a href="https://development-wec.wiline.com/docs/news/agent-web-search-domain-allowlist/#what-you-get" class="hash-link" aria-label="Direct link to What you get" title="Direct link to What you get" translate="no">​</a></h2>
<p>Two filters. A list of sites to allow and a list to block, plus a date range for when a
page was published.</p>
<p>The agent can pass them in the search call itself:</p>
<div class="language-json codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#393A34;--prism-background-color:#f6f8fa"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-json codeBlock_bY9V thin-scrollbar" style="color:#393A34;background-color:#f6f8fa"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#393A34"><span class="token punctuation" style="color:#393A34">{</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">  </span><span class="token property" style="color:#36acaa">"method"</span><span class="token operator" style="color:#393A34">:</span><span class="token plain"> </span><span class="token string" style="color:#e3116c">"tools/call"</span><span class="token punctuation" style="color:#393A34">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">  </span><span class="token property" style="color:#36acaa">"params"</span><span class="token operator" style="color:#393A34">:</span><span class="token plain"> </span><span class="token punctuation" style="color:#393A34">{</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">    </span><span class="token property" style="color:#36acaa">"name"</span><span class="token operator" style="color:#393A34">:</span><span class="token plain"> </span><span class="token string" style="color:#e3116c">"WebSearch"</span><span class="token punctuation" style="color:#393A34">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">    </span><span class="token property" style="color:#36acaa">"arguments"</span><span class="token operator" style="color:#393A34">:</span><span class="token plain"> </span><span class="token punctuation" style="color:#393A34">{</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">      </span><span class="token property" style="color:#36acaa">"query"</span><span class="token operator" style="color:#393A34">:</span><span class="token plain"> </span><span class="token string" style="color:#e3116c">"latest SEC enforcement actions 2026"</span><span class="token punctuation" style="color:#393A34">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">      </span><span class="token property" style="color:#36acaa">"filters"</span><span class="token operator" style="color:#393A34">:</span><span class="token plain"> </span><span class="token punctuation" style="color:#393A34">{</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">        </span><span class="token property" style="color:#36acaa">"domainFilter"</span><span class="token operator" style="color:#393A34">:</span><span class="token plain"> </span><span class="token punctuation" style="color:#393A34">{</span><span class="token plain"> </span><span class="token property" style="color:#36acaa">"include"</span><span class="token operator" style="color:#393A34">:</span><span class="token plain"> </span><span class="token punctuation" style="color:#393A34">[</span><span class="token string" style="color:#e3116c">"sec.gov"</span><span class="token punctuation" style="color:#393A34">]</span><span class="token punctuation" style="color:#393A34">,</span><span class="token plain"> </span><span class="token property" style="color:#36acaa">"exclude"</span><span class="token operator" style="color:#393A34">:</span><span class="token plain"> </span><span class="token punctuation" style="color:#393A34">[</span><span class="token punctuation" style="color:#393A34">]</span><span class="token plain"> </span><span class="token punctuation" style="color:#393A34">}</span><span class="token punctuation" style="color:#393A34">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">        </span><span class="token property" style="color:#36acaa">"publishedDateFilter"</span><span class="token operator" style="color:#393A34">:</span><span class="token plain"> </span><span class="token punctuation" style="color:#393A34">{</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">          </span><span class="token property" style="color:#36acaa">"from"</span><span class="token operator" style="color:#393A34">:</span><span class="token plain"> </span><span class="token string" style="color:#e3116c">"2026-07-01T00:00:00Z"</span><span class="token punctuation" style="color:#393A34">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">          </span><span class="token property" style="color:#36acaa">"to"</span><span class="token operator" style="color:#393A34">:</span><span class="token plain"> </span><span class="token string" style="color:#e3116c">"2026-08-04T23:59:59Z"</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">        </span><span class="token punctuation" style="color:#393A34">}</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">      </span><span class="token punctuation" style="color:#393A34">}</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">    </span><span class="token punctuation" style="color:#393A34">}</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">  </span><span class="token punctuation" style="color:#393A34">}</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain"></span><span class="token punctuation" style="color:#393A34">}</span><br></div></code></pre></div></div>
<p>An admin sets the same kind of list on the gateway instead, where the agent cannot see
or change it. Each list holds up to 100 sites.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-rule">The rule<a href="https://development-wec.wiline.com/docs/news/agent-web-search-domain-allowlist/#the-rule" class="hash-link" aria-label="Direct link to The rule" title="Direct link to The rule" translate="no">​</a></h2>
<p>So the admin has a list and the agent has a list. What happens when they disagree?</p>
<blockquote>
<p>Allow lists are intersected. Block lists are combined.
<em>"Runtime filters can narrow but never expand the scope set by an administrator."</em></p>
</blockquote>
<p>The admin's list is the ceiling. The agent can ask for less, never more. And anything
either one blocks stays blocked.</p>
<p>Turn it around and you see why it matters. If the two allow lists were simply added
together, the admin's list would be a starting suggestion. An agent that wanted some
other site could just name it in its own call and get it. The filter would be paperwork.</p>
<p>That is the difference between the two places you can put a rule. Asking a model to stay
inside a boundary means asking it to remember, halfway through a job, after reading who
knows what. A gateway comparing two lists is not remembering anything.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-date-filter-is-the-sneaky-one">The date filter is the sneaky one<a href="https://development-wec.wiline.com/docs/news/agent-web-search-domain-allowlist/#the-date-filter-is-the-sneaky-one" class="hash-link" aria-label="Direct link to The date filter is the sneaky one" title="Direct link to The date filter is the sneaky one" translate="no">​</a></h2>
<p>The site list gets the attention. The date range may matter more.</p>
<p>An old page does not look old. A four-year-old page about a tax rule or a dead API reads
exactly like a current one — same confident tone, and now with a citation stapled to it,
which makes a wrong answer more convincing rather than less. The model cannot judge how
fresh a page is when nothing tells it. Setting a window gives it a fact it otherwise
never had.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-catch">The catch<a href="https://development-wec.wiline.com/docs/news/agent-web-search-domain-allowlist/#the-catch" class="hash-link" aria-label="Direct link to The catch" title="Direct link to The catch" translate="no">​</a></h2>
<p>Search costs <strong>$7 per 1,000 queries</strong> and runs in three regions. AWS's selling point is
that the queries stay inside their network — <em>"without sending user prompts and
retrieval queries to external search API providers outside of AWS."</em> Good if you already
live in AWS. Mostly irrelevant if you do not.</p>
<p>The idea travels, though, because the tool is reached over MCP: <em>"Web Search uses a
built-in connector target on Bedrock AgentCore Gateway using the Model Context Protocol
(MCP)."</em> Nothing about <em>allow lists intersect, block lists combine, the agent can only
narrow</em> needs Amazon. Any gateway can do it, for any tool — which files something may
read, which hosts it may reach, which tables it may query.</p>
<p>So the question to ask about a tool is not whether you can restrict it. It is where the
restriction lives, and whether the agent can move it.</p>
<p>We went through the protocol under all this in
<a class="" href="https://development-wec.wiline.com/docs/news/mcp-2026-07-28-spec/">MCP's biggest update</a>, and gave a model live search without
a managed service in
<a class="" href="https://development-wec.wiline.com/docs/tutorials/web-search-wiline-inference/">web search on WEC Inference</a>. The next tutorial
in the agent series puts an agent's tools behind a gateway that decides what it may
call.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="sources">Sources<a href="https://development-wec.wiline.com/docs/news/agent-web-search-domain-allowlist/#sources" class="hash-link" aria-label="Direct link to Sources" title="Direct link to Sources" translate="no">​</a></h2>
<ul>
<li class=""><a href="https://aws.amazon.com/blogs/machine-learning/domain-and-publish-date-filters-for-web-search-on-agentcore/" target="_blank" rel="noopener noreferrer" class="">Domain and publish date filters for Web Search on AgentCore</a> — 19 August 2026</li>
<li class=""><a href="https://aws.amazon.com/about-aws/whats-new/2026/08/web-search-amazon-bedrock/" target="_blank" rel="noopener noreferrer" class="">Web Search in Amazon Bedrock AgentCore adds domain and published date filtering, expands to Europe and Asia Pacific</a> — 19 August 2026</li>
<li class=""><a href="https://aws.amazon.com/blogs/aws/announcing-web-search-on-amazon-bedrock-agentcore-ground-your-ai-agents-in-current-accurate-web-knowledge/" target="_blank" rel="noopener noreferrer" class="">Announcing Web Search on Amazon Bedrock AgentCore</a> — 17 June 2026</li>
</ul>]]></content>
        <author>
            <name>Rafael Fernandes</name>
            <uri>https://www.linkedin.com/in/rafaelmacariofernandes/</uri>
        </author>
        <category label="ai-news" term="ai-news"/>
        <category label="agents" term="agents"/>
        <category label="mcp" term="mcp"/>
        <category label="guardrails" term="guardrails"/>
        <category label="governance" term="governance"/>
        <category label="web-search" term="web-search"/>
    </entry>
    <entry>
        <title type="html"><![CDATA[A router that can't see the conversation can't classify "yes"]]></title>
        <id>https://development-wec.wiline.com/docs/news/llm-router-cannot-classify-yes/</id>
        <link href="https://development-wec.wiline.com/docs/news/llm-router-cannot-classify-yes/"/>
        <updated>2026-08-25T00:00:00.000Z</updated>
        <summary type="html"><![CDATA[LiteLLM ran 5,600 live classifier calls to answer one question: how much of the conversation does a model router need to see? On the follow-ups that only make sense against history, agreement went from 14% to 78% — and the completion bill more than doubled, because two thirds of them had been going to the cheapest model.]]></summary>
        <content type="html"><![CDATA[<div class="newsHero"><div class="newsHero__glow" aria-hidden="true"></div><span class="newsHero__eyebrow">Routing · AI News</span><h2 class="newsHero__title">Classifying the word "yes"</h2><div class="newsHero__transition"><span class="newsHero__pill newsHero__pill--from">Score this message</span><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2.5" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-arrow-right newsHero__arrow" aria-hidden="true"><path d="M5 12h14"></path><path d="m12 5 7 7-7 7"></path></svg><span class="newsHero__pill newsHero__pill--to">Score what it approves</span></div></div>
<p>A model router's job is to read a request and decide which model should answer it.
Cheap questions go to a small model, hard ones to a large one, and the bill comes
down. The whole arrangement rests on being able to tell the difference.</p>
<p>Then a user types "yes".</p>
<p>Or "continue". Or "do it". Nothing in those two or three characters says whether
the work being approved is a spelling fix or a database migration. A router
scoring the current message in isolation sees a very short string with no
technical vocabulary, and does the obvious thing: cheapest model.</p>
<p><strong>Which means if you route requests to save money, your cheapest tier is probably
absorbing work it should never have seen — and your savings figure is partly
fake.</strong> On 4 August LiteLLM published a benchmark that measures both halves of
that: how wrong the routing gets, and what it costs to fix.</p>
<!-- -->
<p>The question generalises past their implementation — anything classifying a turn
in a conversation has this problem — and the reason it's worth reading is that
they measured the price of the fix, not only the benefit.</p>
<div class="theme-admonition theme-admonition-note admonition_xJq3 alert alert--secondary"><div class="admonitionHeading_Gvgb"><span class="admonitionIcon_Rf37"><svg viewBox="0 0 14 16"><path fill-rule="evenodd" d="M6.3 5.69a.942.942 0 0 1-.28-.7c0-.28.09-.52.28-.7.19-.18.42-.28.7-.28.28 0 .52.09.7.28.18.19.28.42.28.7 0 .28-.09.52-.28.7a1 1 0 0 1-.7.3c-.28 0-.52-.11-.7-.3zM8 7.99c-.02-.25-.11-.48-.31-.69-.2-.19-.42-.3-.69-.31H6c-.27.02-.48.13-.69.31-.2.2-.3.44-.31.69h1v3c.02.27.11.5.31.69.2.2.42.31.69.31h1c.27 0 .48-.11.69-.31.2-.19.3-.42.31-.69H8V7.98v.01zM7 2.3c-3.14 0-5.7 2.54-5.7 5.68 0 3.14 2.56 5.7 5.7 5.7s5.7-2.55 5.7-5.7c0-3.15-2.56-5.69-5.7-5.69v.01zM7 .98c3.86 0 7 3.14 7 7s-3.14 7-7 7-7-3.12-7-7 3.14-7 7-7z"></path></svg></span>Whose numbers these are</div><div class="admonitionContent_BuS1"><p>Everything below is LiteLLM's own measurement, published on their blog and quoted
here: v1.97, classifier <code>gpt-5.4-mini</code>, their three datasets, their reference
labels. None of it was run on our infrastructure, and we have not reproduced it.</p></div></div>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-measurement">The measurement<a href="https://development-wec.wiline.com/docs/news/llm-router-cannot-classify-yes/#the-measurement" class="hash-link" aria-label="Direct link to The measurement" title="Direct link to The measurement" translate="no">​</a></h2>
<p>The sweep: <strong>5,600 live classifier calls against real providers</strong>, described as
<em>"two sweeps of seven configurations each (<code>classifier_context_window_size</code> of 0,
1, 2, 3, 5, 8, 10), one with assistant turns in the window and one without, across
three multi-turn datasets, with two repeats per conversation."</em></p>
<p>The variable is how many prior turns the classifier gets to see. Zero means it
judges the current message alone. Ten means it reads the last ten turns first.</p>
<p>Agreement with reference tiers, by window size:</p>
<table><thead><tr><th>Prior turns</th><th>Short-reply follow-ups</th><th>MT-Bench 2nd turns</th><th>ShareGPT multi-turn</th></tr></thead><tbody><tr><td>0</td><td>50.0%</td><td>49.4%</td><td>83.8%</td></tr><tr><td>1</td><td>71.2%</td><td>53.1%</td><td>84.4%</td></tr><tr><td>2</td><td>87.5%</td><td>53.1%</td><td>90.6%</td></tr><tr><td><strong>3 (default)</strong></td><td><strong>85.0%</strong></td><td><strong>55.0%</strong></td><td><strong>91.9%</strong></td></tr><tr><td>5</td><td>86.2%</td><td>53.8%</td><td>91.9%</td></tr><tr><td>8</td><td>87.5%</td><td>55.6%</td><td>91.2%</td></tr><tr><td>10</td><td>90.0%</td><td>55.6%</td><td>88.8%</td></tr></tbody></table>
<p><span class="zoomImage__wrap"><img alt="Line chart of classifier agreement against the number of prior conversation turns, showing the history-dependent subset rising from 14% to 78% by two turns and flat thereafter" src="data:image/svg+xml;base64,PHN2ZyB4bWxucz0iaHR0cDovL3d3dy53My5vcmcvMjAwMC9zdmciIHZpZXdCb3g9IjAgMCA3NjAgNDMwIiB3aWR0aD0iNzYwIiBoZWlnaHQ9IjQzMCIgZm9udC1mYW1pbHk9InN5c3RlbS11aSwtYXBwbGUtc3lzdGVtLFNlZ29lIFVJLHNhbnMtc2VyaWYiPgo8dGl0bGU+Q2xhc3NpZmllciBhZ3JlZW1lbnQgd2l0aCByZWZlcmVuY2UgdGllcnMsIGJ5IG51bWJlciBvZiBwcmlvciBjb252ZXJzYXRpb24gdHVybnM8L3RpdGxlPgo8bGluZSB4MT0iNzgiIHkxPSIzMzAuMCIgeDI9IjczMCIgeTI9IjMzMC4wIiBzdHJva2U9IiM5NGEzYjgiIHN0cm9rZS1vcGFjaXR5PSIwLjMwIi8+Cjx0ZXh0IHg9IjY2IiB5PSIzMzQuMCIgdGV4dC1hbmNob3I9ImVuZCIgZm9udC1zaXplPSIxMyIgZmlsbD0iIzY0NzQ4YiI+MCU8L3RleHQ+CjxsaW5lIHgxPSI3OCIgeTE9IjI3MC44IiB4Mj0iNzMwIiB5Mj0iMjcwLjgiIHN0cm9rZT0iIzk0YTNiOCIgc3Ryb2tlLW9wYWNpdHk9IjAuMzAiLz4KPHRleHQgeD0iNjYiIHk9IjI3NC44IiB0ZXh0LWFuY2hvcj0iZW5kIiBmb250LXNpemU9IjEzIiBmaWxsPSIjNjQ3NDhiIj4yMCU8L3RleHQ+CjxsaW5lIHgxPSI3OCIgeTE9IjIxMS42IiB4Mj0iNzMwIiB5Mj0iMjExLjYiIHN0cm9rZT0iIzk0YTNiOCIgc3Ryb2tlLW9wYWNpdHk9IjAuMzAiLz4KPHRleHQgeD0iNjYiIHk9IjIxNS42IiB0ZXh0LWFuY2hvcj0iZW5kIiBmb250LXNpemU9IjEzIiBmaWxsPSIjNjQ3NDhiIj40MCU8L3RleHQ+CjxsaW5lIHgxPSI3OCIgeTE9IjE1Mi40IiB4Mj0iNzMwIiB5Mj0iMTUyLjQiIHN0cm9rZT0iIzk0YTNiOCIgc3Ryb2tlLW9wYWNpdHk9IjAuMzAiLz4KPHRleHQgeD0iNjYiIHk9IjE1Ni40IiB0ZXh0LWFuY2hvcj0iZW5kIiBmb250LXNpemU9IjEzIiBmaWxsPSIjNjQ3NDhiIj42MCU8L3RleHQ+CjxsaW5lIHgxPSI3OCIgeTE9IjkzLjIiIHgyPSI3MzAiIHkyPSI5My4yIiBzdHJva2U9IiM5NGEzYjgiIHN0cm9rZS1vcGFjaXR5PSIwLjMwIi8+Cjx0ZXh0IHg9IjY2IiB5PSI5Ny4yIiB0ZXh0LWFuY2hvcj0iZW5kIiBmb250LXNpemU9IjEzIiBmaWxsPSIjNjQ3NDhiIj44MCU8L3RleHQ+CjxsaW5lIHgxPSI3OCIgeTE9IjM0LjAiIHgyPSI3MzAiIHkyPSIzNC4wIiBzdHJva2U9IiM5NGEzYjgiIHN0cm9rZS1vcGFjaXR5PSIwLjMwIi8+Cjx0ZXh0IHg9IjY2IiB5PSIzOC4wIiB0ZXh0LWFuY2hvcj0iZW5kIiBmb250LXNpemU9IjEzIiBmaWxsPSIjNjQ3NDhiIj4xMDAlPC90ZXh0Pgo8bGluZSB4MT0iNzgiIHkxPSIzMzAuMCIgeDI9IjczMCIgeTI9IjMzMC4wIiBzdHJva2U9IiM5NGEzYjgiLz4KPHRleHQgeD0iNzguMCIgeT0iMzU0IiB0ZXh0LWFuY2hvcj0ibWlkZGxlIiBmb250LXNpemU9IjEzIiBmaWxsPSIjNjQ3NDhiIj4wPC90ZXh0Pgo8dGV4dCB4PSIxODYuNyIgeT0iMzU0IiB0ZXh0LWFuY2hvcj0ibWlkZGxlIiBmb250LXNpemU9IjEzIiBmaWxsPSIjNjQ3NDhiIj4xPC90ZXh0Pgo8dGV4dCB4PSIyOTUuMyIgeT0iMzU0IiB0ZXh0LWFuY2hvcj0ibWlkZGxlIiBmb250LXNpemU9IjEzIiBmaWxsPSIjNjQ3NDhiIj4yPC90ZXh0Pgo8dGV4dCB4PSI0MDQuMCIgeT0iMzU0IiB0ZXh0LWFuY2hvcj0ibWlkZGxlIiBmb250LXNpemU9IjEzIiBmaWxsPSIjNjQ3NDhiIj4zPC90ZXh0Pgo8dGV4dCB4PSI1MTIuNyIgeT0iMzU0IiB0ZXh0LWFuY2hvcj0ibWlkZGxlIiBmb250LXNpemU9IjEzIiBmaWxsPSIjNjQ3NDhiIj41PC90ZXh0Pgo8dGV4dCB4PSI2MjEuMyIgeT0iMzU0IiB0ZXh0LWFuY2hvcj0ibWlkZGxlIiBmb250LXNpemU9IjEzIiBmaWxsPSIjNjQ3NDhiIj44PC90ZXh0Pgo8dGV4dCB4PSI3MzAuMCIgeT0iMzU0IiB0ZXh0LWFuY2hvcj0ibWlkZGxlIiBmb250LXNpemU9IjEzIiBmaWxsPSIjNjQ3NDhiIj4xMDwvdGV4dD4KPHRleHQgeD0iNDA0LjAiIHk9IjM4MCIgdGV4dC1hbmNob3I9Im1pZGRsZSIgZm9udC1zaXplPSIxMy41IiBmaWxsPSIjNjQ3NDhiIj5wcmlvciBjb252ZXJzYXRpb24gdHVybnMgdGhlIGNsYXNzaWZpZXIgY2FuIHNlZTwvdGV4dD4KPHBhdGggZD0iTTc4LjAsMjg4LjYgTDE4Ni43LDE5MC45IEwyOTUuMyw5OS4xIiBmaWxsPSJub25lIiBzdHJva2U9IiNkYzI2MjYiIHN0cm9rZS13aWR0aD0iMi42IiBzdHJva2UtbGluZWpvaW49InJvdW5kIi8+CjxwYXRoIGQ9Ik0yOTUuMyw5OS4xIEw3MzAuMCw5OS4xIiBmaWxsPSJub25lIiBzdHJva2U9IiNkYzI2MjYiIHN0cm9rZS13aWR0aD0iMi42IiBzdHJva2UtZGFzaGFycmF5PSI3IDYiLz4KPHRleHQgeD0iNTEyLjciIHk9Ijg4LjEiIHRleHQtYW5jaG9yPSJtaWRkbGUiIGZvbnQtc2l6ZT0iMTIuNSIgZmlsbD0iI2RjMjYyNiI+ZmxhdCB0byBOPTEwICh0aGVpciB3b3Jkcyk8L3RleHQ+CjxjaXJjbGUgY3g9Ijc4LjAiIGN5PSIyODguNiIgcj0iMy42IiBmaWxsPSIjZGMyNjI2Ii8+CjxjaXJjbGUgY3g9IjE4Ni43IiBjeT0iMTkwLjkiIHI9IjMuNiIgZmlsbD0iI2RjMjYyNiIvPgo8Y2lyY2xlIGN4PSIyOTUuMyIgY3k9Ijk5LjEiIHI9IjMuNiIgZmlsbD0iI2RjMjYyNiIvPgo8cGF0aCBkPSJNNzguMCwxODIuMCBMMTg2LjcsMTE5LjIgTDI5NS4zLDcxLjAgTDQwNC4wLDc4LjQgTDUxMi43LDc0LjggTDYyMS4zLDcxLjAgTDczMC4wLDYzLjYiIGZpbGw9Im5vbmUiIHN0cm9rZT0iIzI1NjNlYiIgc3Ryb2tlLXdpZHRoPSIyLjYiIHN0cm9rZS1saW5lam9pbj0icm91bmQiLz4KPGNpcmNsZSBjeD0iNzguMCIgY3k9IjE4Mi4wIiByPSIzLjYiIGZpbGw9IiMyNTYzZWIiLz4KPGNpcmNsZSBjeD0iMTg2LjciIGN5PSIxMTkuMiIgcj0iMy42IiBmaWxsPSIjMjU2M2ViIi8+CjxjaXJjbGUgY3g9IjI5NS4zIiBjeT0iNzEuMCIgcj0iMy42IiBmaWxsPSIjMjU2M2ViIi8+CjxjaXJjbGUgY3g9IjQwNC4wIiBjeT0iNzguNCIgcj0iMy42IiBmaWxsPSIjMjU2M2ViIi8+CjxjaXJjbGUgY3g9IjUxMi43IiBjeT0iNzQuOCIgcj0iMy42IiBmaWxsPSIjMjU2M2ViIi8+CjxjaXJjbGUgY3g9IjYyMS4zIiBjeT0iNzEuMCIgcj0iMy42IiBmaWxsPSIjMjU2M2ViIi8+CjxjaXJjbGUgY3g9IjczMC4wIiBjeT0iNjMuNiIgcj0iMy42IiBmaWxsPSIjMjU2M2ViIi8+CjxwYXRoIGQ9Ik03OC4wLDgyLjAgTDE4Ni43LDgwLjIgTDI5NS4zLDYxLjggTDQwNC4wLDU4LjAgTDUxMi43LDU4LjAgTDYyMS4zLDYwLjAgTDczMC4wLDY3LjIiIGZpbGw9Im5vbmUiIHN0cm9rZT0iIzA1OTY2OSIgc3Ryb2tlLXdpZHRoPSIyLjYiIHN0cm9rZS1saW5lam9pbj0icm91bmQiLz4KPGNpcmNsZSBjeD0iNzguMCIgY3k9IjgyLjAiIHI9IjMuNiIgZmlsbD0iIzA1OTY2OSIvPgo8Y2lyY2xlIGN4PSIxODYuNyIgY3k9IjgwLjIiIHI9IjMuNiIgZmlsbD0iIzA1OTY2OSIvPgo8Y2lyY2xlIGN4PSIyOTUuMyIgY3k9IjYxLjgiIHI9IjMuNiIgZmlsbD0iIzA1OTY2OSIvPgo8Y2lyY2xlIGN4PSI0MDQuMCIgY3k9IjU4LjAiIHI9IjMuNiIgZmlsbD0iIzA1OTY2OSIvPgo8Y2lyY2xlIGN4PSI1MTIuNyIgY3k9IjU4LjAiIHI9IjMuNiIgZmlsbD0iIzA1OTY2OSIvPgo8Y2lyY2xlIGN4PSI2MjEuMyIgY3k9IjYwLjAiIHI9IjMuNiIgZmlsbD0iIzA1OTY2OSIvPgo8Y2lyY2xlIGN4PSI3MzAuMCIgY3k9IjY3LjIiIHI9IjMuNiIgZmlsbD0iIzA1OTY2OSIvPgo8cGF0aCBkPSJNNzguMCwxODMuOCBMMTg2LjcsMTcyLjggTDI5NS4zLDE3Mi44IEw0MDQuMCwxNjcuMiBMNTEyLjcsMTcwLjggTDYyMS4zLDE2NS40IEw3MzAuMCwxNjUuNCIgZmlsbD0ibm9uZSIgc3Ryb2tlPSIjOTRhM2I4IiBzdHJva2Utd2lkdGg9IjIuNiIgc3Ryb2tlLWxpbmVqb2luPSJyb3VuZCIvPgo8Y2lyY2xlIGN4PSI3OC4wIiBjeT0iMTgzLjgiIHI9IjMuNiIgZmlsbD0iIzk0YTNiOCIvPgo8Y2lyY2xlIGN4PSIxODYuNyIgY3k9IjE3Mi44IiByPSIzLjYiIGZpbGw9IiM5NGEzYjgiLz4KPGNpcmNsZSBjeD0iMjk1LjMiIGN5PSIxNzIuOCIgcj0iMy42IiBmaWxsPSIjOTRhM2I4Ii8+CjxjaXJjbGUgY3g9IjQwNC4wIiBjeT0iMTY3LjIiIHI9IjMuNiIgZmlsbD0iIzk0YTNiOCIvPgo8Y2lyY2xlIGN4PSI1MTIuNyIgY3k9IjE3MC44IiByPSIzLjYiIGZpbGw9IiM5NGEzYjgiLz4KPGNpcmNsZSBjeD0iNjIxLjMiIGN5PSIxNjUuNCIgcj0iMy42IiBmaWxsPSIjOTRhM2I4Ii8+CjxjaXJjbGUgY3g9IjczMC4wIiBjeT0iMTY1LjQiIHI9IjMuNiIgZmlsbD0iIzk0YTNiOCIvPgo8cmVjdCB4PSI3OCIgeT0iNDA0IiB3aWR0aD0iMjAiIGhlaWdodD0iMyIgZmlsbD0iI2RjMjYyNiIvPgo8dGV4dCB4PSIxMDUiIHk9IjQxMCIgZm9udC1zaXplPSIxMyIgZmlsbD0iIzY0NzQ4YiI+MzYgaGlzdG9yeS1kZXBlbmRlbnQgZm9sbG93LXVwczwvdGV4dD4KPHJlY3QgeD0iMzUxLjEiIHk9IjQwNCIgd2lkdGg9IjIwIiBoZWlnaHQ9IjMiIGZpbGw9IiMyNTYzZWIiLz4KPHRleHQgeD0iMzc4LjEiIHk9IjQxMCIgZm9udC1zaXplPSIxMyIgZmlsbD0iIzY0NzQ4YiI+U2hvcnQtcmVwbHkgc2V0PC90ZXh0Pgo8cmVjdCB4PSI1MTAuNiIgeT0iNDA0IiB3aWR0aD0iMjAiIGhlaWdodD0iMyIgZmlsbD0iIzA1OTY2OSIvPgo8dGV4dCB4PSI1MzcuNiIgeT0iNDEwIiBmb250LXNpemU9IjEzIiBmaWxsPSIjNjQ3NDhiIj5TaGFyZUdQVDwvdGV4dD4KPHJlY3QgeD0iNjIwLjQiIHk9IjQwNCIgd2lkdGg9IjIwIiBoZWlnaHQ9IjMiIGZpbGw9IiM5NGEzYjgiLz4KPHRleHQgeD0iNjQ3LjQiIHk9IjQxMCIgZm9udC1zaXplPSIxMyIgZmlsbD0iIzY0NzQ4YiI+TVQtQmVuY2g8L3RleHQ+Cjwvc3ZnPg==" width="760" height="430" class="zoomImage " loading="lazy"><span class="zoomImage__badge" aria-hidden="true"><svg viewBox="0 0 24 24" width="16" height="16" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round"><circle cx="11" cy="11" r="7"></circle><path d="M21 21l-4.3-4.3"></path><path d="M11 8v6M8 11h6"></path></svg></span></span></p>
<p><strong>Figure 1.</strong> Agreement with reference tiers by context window, drawn from the figures published in LiteLLM's benchmark. The red line is the 36-follow-up subset whose difficulty only resolves against history.</p>
<p>Three different stories in three columns, which is the first useful thing here.</p>
<p>ShareGPT starts at 83.8% and gains eight points. MT-Bench barely moves at all —
49.4% to 55.6% across the whole sweep — and their own caveat explains why:
<em>"MT-Bench's ceiling reflects its reference labels rather than router behaviour."</em>
The short-reply set is where it bites, 50% to 90%.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-number-worth-quoting-with-its-denominator">The number worth quoting, with its denominator<a href="https://development-wec.wiline.com/docs/news/llm-router-cannot-classify-yes/#the-number-worth-quoting-with-its-denominator" class="hash-link" aria-label="Direct link to The number worth quoting, with its denominator" title="Direct link to The number worth quoting, with its denominator" translate="no">​</a></h2>
<p>Inside that first column sits a subset, and it is the sharpest result in the post:</p>
<blockquote>
<p>Agreement there is <strong>14% at N=0, 47% at N=1, and 78% at N=2</strong>, and flat from
there out to N=10.</p>
</blockquote>
<p>Fourteen per cent to seventy-eight per cent, achieved by showing the classifier
two prior turns.</p>
<p>That figure describes <strong>36 follow-ups</strong> — the ones <em>"whose final turn only
resolves against the history"</em>. It is a deliberately selected hard subset, not a
general accuracy claim, and reading it as "routers are 14% accurate" would be
wrong. What it does show is the shape of the failure: when a turn's difficulty
lives entirely in what came before it, a context-free classifier is worse than
guessing, and almost all of the recovery happens by the second turn of history.</p>
<p>The curve going flat at N=2 is the practically useful part. This is not a
"more context is better" result. It is a "two turns is nearly all of it" result.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="what-it-cost-and-what-it-cost-more">What it cost, and what it cost <em>more</em><a href="https://development-wec.wiline.com/docs/news/llm-router-cannot-classify-yes/#what-it-cost-and-what-it-cost-more" class="hash-link" aria-label="Direct link to what-it-cost-and-what-it-cost-more" title="Direct link to what-it-cost-and-what-it-cost-more" translate="no">​</a></h2>
<p>This is the half that makes the finding actionable, and where the honest answer is
more interesting than "it's cheap".</p>
<p><strong>The window itself is free.</strong> Every paired comparison against the zero-window
configuration has a 95% bootstrap confidence interval straddling zero. Window 1
with assistant turns off comes in at −17.8 ms, interval −68.9 to +29.6. Their
explanation is that the window adds prefill only — the output stays a small fixed
tier label. Over a 318-to-1,043 prompt-token range, tokens and latency correlate
at r = 0.007.</p>
<p><strong>The classifier is not free.</strong> Adding history costs nothing, but asking a model
to classify at all costs <em>"p50 sits near 600 ms in every configuration"</em> — and
that sits in front of the real completion. The classifier here is
<code>gpt-5.4-mini</code>, and it runs at <strong>$0.31–$0.61 per 1,000 requests</strong> across the whole
sweep. Cheap in money, half a second in time.</p>
<p><strong>And accuracy raised the bill.</strong> This is the part worth sitting with. For the
short-reply set, the modelled routed cost <em>rose</em> with the window — from <strong>$2.87 to
$6.49 per 1,000 requests</strong>. Not a regression: the reason is <em>"a tier mix of 66%
SIMPLE at N=0 against 30% at N=2"</em>. Two thirds of those follow-ups were being
served by the cheapest model. Once the classifier could see what they were
approving, they stopped being.</p>
<p>So the routing was cheap because it was wrong. The saving was a bill someone else
was paying, in answer quality, on requests nobody was auditing. On the other two
datasets the effect runs the other way — ShareGPT's modelled cost falls slightly,
because context lets the classifier settle ambiguous prompts into MEDIUM instead
of defaulting up to REASONING.</p>
<p>That is the useful shape of it: better context doesn't reliably cut spend. It moves
spend toward where the work actually was.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="two-things-this-doesnt-settle">Two things this doesn't settle<a href="https://development-wec.wiline.com/docs/news/llm-router-cannot-classify-yes/#two-things-this-doesnt-settle" class="hash-link" aria-label="Direct link to Two things this doesn't settle" title="Direct link to Two things this doesn't settle" translate="no">​</a></h2>
<p>The default is 3, and 3 is not the best number in any column. Ten is best for
short-reply follow-ups at 90.0%, and simultaneously the <em>worst</em> result for
ShareGPT past N=2, dropping to 88.8% from 91.9%. More history is not monotonically
better, and the shipped default is a judgement call across datasets that disagree.</p>
<p>And they name their own limits, which is the mark of a benchmark worth trusting:
<em>"Reference tiers are judgement calls"</em>, <em>"Routed completion cost is modelled
rather than billed"</em>, <em>"One classifier model was swept"</em>, and <em>"Latency was
measured on a single VM at concurrency 10"</em>. Concurrency 10 on one machine is not
production, and a heavier classifier model carries more absolute latency than the
one they used.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="why-it-matters-if-you-route-anything">Why it matters if you route anything<a href="https://development-wec.wiline.com/docs/news/llm-router-cannot-classify-yes/#why-it-matters-if-you-route-anything" class="hash-link" aria-label="Direct link to Why it matters if you route anything" title="Direct link to Why it matters if you route anything" translate="no">​</a></h2>
<p>The lesson survives the specific implementation. If you are choosing models per
request — with a heuristic, a classifier, or a hand-rolled rule — the turns that
will embarrass you are not the long technical ones. They are the short ones. A
router scoring "yes" as a trivial request routes the approval of an architecture
migration to the smallest model you own, and nothing in your logs will flag it,
because a cheap answer to a cheap-looking question is exactly what you asked for.</p>
<p>Their fix is a window of prior turns. The cheaper fix, if your classifier has no
such option, is to notice that short replies in an established conversation are
the population to worry about, and to stop scoring them on their own contents.</p>
<p>We took the same router apart from the other direction recently — the seven
scoring dimensions, the arithmetic on a real prompt, and what a free keyword
scorer misses that a model catches — in
<a class="" href="https://development-wec.wiline.com/docs/tutorials/litellm-complexity-routing/">route by complexity</a>.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="sources">Sources<a href="https://development-wec.wiline.com/docs/news/llm-router-cannot-classify-yes/#sources" class="hash-link" aria-label="Direct link to Sources" title="Direct link to Sources" translate="no">​</a></h2>
<ul>
<li class=""><a href="https://docs.litellm.ai/blog/auto-router-context-and-benchmarks" target="_blank" rel="noopener noreferrer" class="">Auto Router v1.97: usage benchmarks and better quality for lower cost</a> — 4 August 2026</li>
<li class=""><a href="https://docs.litellm.ai/docs/proxy/auto_routing" target="_blank" rel="noopener noreferrer" class="">Auto Routing configuration reference</a></li>
</ul>]]></content>
        <author>
            <name>Rafael Fernandes</name>
            <uri>https://www.linkedin.com/in/rafaelmacariofernandes/</uri>
        </author>
        <category label="ai-news" term="ai-news"/>
        <category label="routing" term="routing"/>
        <category label="evals" term="evals"/>
        <category label="cost" term="cost"/>
        <category label="litellm" term="litellm"/>
    </entry>
    <entry>
        <title type="html"><![CDATA[Mask Your Logs, Not Your Prompts]]></title>
        <id>https://development-wec.wiline.com/docs/news/mask-your-logs-not-your-prompts/</id>
        <link href="https://development-wec.wiline.com/docs/news/mask-your-logs-not-your-prompts/"/>
        <updated>2026-08-14T00:00:00.000Z</updated>
        <summary type="html"><![CDATA[The most common LLM privacy advice — scrub personal data out of the prompt before the model sees it — aims at the wrong risk and quietly makes the model dumber. Here's what the research actually shows, and where the masking really belongs.]]></summary>
        <content type="html"><![CDATA[<div class="newsHero"><div class="newsHero__glow" aria-hidden="true"></div><span class="newsHero__eyebrow">Privacy · AI News</span><h2 class="newsHero__title">Mask your logs, not your prompts</h2><div class="newsHero__transition"><span class="newsHero__pill newsHero__pill--from">Redact before the model</span><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2.5" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-arrow-right newsHero__arrow" aria-hidden="true"><path d="M5 12h14"></path><path d="m12 5 7 7-7 7"></path></svg><span class="newsHero__pill newsHero__pill--to">Redact before the logs</span></div></div>
<p>Almost every "secure your LLM app" guide gives the same advice: before a prompt reaches
the model, strip the personal data out of it — swap names, emails, and account numbers for
<code>[REDACTED]</code> or <code>&lt;PERSON&gt;</code>, <em>then</em> call the model.</p>
<p>It sounds obviously right. I assumed it was, too. Then I went and read the research on what
masking actually does to a model, and the papers point the other way. The short version:
<strong>mask your logs, not your prompts.</strong></p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="first-what-are-we-protecting">First, what are we protecting?<a href="https://development-wec.wiline.com/docs/news/mask-your-logs-not-your-prompts/#first-what-are-we-protecting" class="hash-link" aria-label="Direct link to First, what are we protecting?" title="Direct link to First, what are we protecting?" translate="no">​</a></h2>
<p><strong>PII</strong> is <em>personally identifiable information</em> — a name, an email, a phone number, an
account or government ID. Data that points at a specific person.</p>
<p>When people reach for pre-call masking, they're usually blending two different worries into
one:</p>
<ol>
<li class=""><em>"I don't want to send sensitive data to whoever runs the model."</em></li>
<li class=""><em>"I don't want to store sensitive data in my logs."</em></li>
</ol>
<p>Pre-call masking is aimed at <strong>#1</strong>. Hold on to that — it turns out to matter.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="problem-1-a-masked-prompt-is-a-confused-model">Problem 1: a masked prompt is a confused model<a href="https://development-wec.wiline.com/docs/news/mask-your-logs-not-your-prompts/#problem-1-a-masked-prompt-is-a-confused-model" class="hash-link" aria-label="Direct link to Problem 1: a masked prompt is a confused model" title="Direct link to Problem 1: a masked prompt is a confused model" translate="no">​</a></h2>
<p>Here's the simplest failure. Put three people in a prompt and redact all of them to
<code>[REDACTED]</code>, and the model can no longer tell them apart. Who signed the contract? Who
was cc'd? The words that carried those relationships are gone, so the answer degrades.</p>
<p>This isn't just intuition. A 2026 benchmark called <strong>RedacBench</strong> measured the trade-off
directly, across 514 texts and 187 policies. Even when a <em>capable</em> model does the redacting,
turning the security dial up to ~81% of sensitive content removed leaves you keeping only
<strong>37.6%</strong> of the text's non-sensitive meaning. You throw away roughly <strong>60% of what made
the prompt useful</strong> to buy that privacy.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="but-i-use-smart-tokenization--it-still-bites">"But I use smart tokenization" — it still bites<a href="https://development-wec.wiline.com/docs/news/mask-your-logs-not-your-prompts/#but-i-use-smart-tokenization--it-still-bites" class="hash-link" aria-label="Direct link to &quot;But I use smart tokenization&quot; — it still bites" title="Direct link to &quot;But I use smart tokenization&quot; — it still bites" translate="no">​</a></h2>
<p>The obvious fix is to stop using dumb placeholders. Deterministic tokenization maps the
same value to the same token every time — "John Smith" always becomes <code>PERSON_42</code>,
"Jane Doe" always <code>PERSON_17</code> — so the model can still track who did what. That's genuinely
better.</p>
<p>But it isn't free either. One engineer ran <strong>109 masking tests</strong> across healthcare, legal,
financial, and developer workflows and wrote up where it broke. <em>(Full disclosure: he also
sells a tokenization tool, so take his framing with a grain of salt — but the failures he
logged are concrete and easy to reproduce.)</em></p>
<ul>
<li class=""><strong>Context-phrase refusals.</strong> Tokenize an SSN into <code>GOV_ID_8x3m</code>, but leave the words
"social security number" sitting next to it, and the model's safety filter can refuse the
whole request — it sees a sensitive label beside an opaque token and flags it.</li>
<li class=""><strong>False positives.</strong> The word "Will" in <em>"this will update the record"</em> got caught by the
name detector.</li>
<li class=""><strong>Misses.</strong> A short name like "Li" in a table row, or an SSN buried in code comments, slid
past the detector when there wasn't enough surrounding context.</li>
<li class=""><strong>Streaming corruption.</strong> A name or SSN can be split across two or three streaming chunks;
process them one at a time and you mangle the entity.</li>
</ul>
<p>His overall detection came out to <strong>89%</strong> — and his sharper point was that the missing 11%
is where it hurts: you're forced to either block a legitimate request or leak. There's no
comfortable default.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="why-none-of-this-is-surprising-my-read-not-a-proof">Why none of this is surprising (my read, not a proof)<a href="https://development-wec.wiline.com/docs/news/mask-your-logs-not-your-prompts/#why-none-of-this-is-surprising-my-read-not-a-proof" class="hash-link" aria-label="Direct link to Why none of this is surprising (my read, not a proof)" title="Direct link to Why none of this is surprising (my read, not a proof)" translate="no">​</a></h2>
<p>Step back and it fits a pattern the research keeps finding: <strong>models are fragile to how a
prompt is worded.</strong></p>
<ul>
<li class=""><em>"On the Worst Prompt Performance of LLMs"</em> took the same question, reworded it in
semantically identical ways, and watched one model's accuracy swing by <strong>45 points</strong>
(worst case, 9.38%).</li>
<li class="">The <strong>DETAIL</strong> framework found that <em>more specific</em> prompts reason better, especially on
smaller models and step-by-step tasks.</li>
</ul>
<p>Neither of those papers tested PII masking — so this next step is <em>my inference, not their
claim</em> — but masking <strong>is</strong> a prompt edit. It makes the prompt less specific and changes its
wording, which is exactly the lever these papers show models are sensitive to. You're
rolling dice you don't need to roll.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-hidden-bill">The hidden bill<a href="https://development-wec.wiline.com/docs/news/mask-your-logs-not-your-prompts/#the-hidden-bill" class="hash-link" aria-label="Direct link to The hidden bill" title="Direct link to The hidden bill" translate="no">​</a></h2>
<p>Masking also costs money in a way that's easy to miss, because it changes the prompt on
<strong>every</strong> call. That quietly breaks <strong>prompt caching</strong> — the discount you get when a prompt's
opening is identical to a previous one.</p>
<p>Both major providers cache by prefix, and both say a change up front invalidates it:</p>
<ul>
<li class="">OpenAI: <em>"Cache hits are only possible for exact prefix matches… a change before the
breakpoint will prevent a cache hit."</em></li>
<li class="">Anthropic: <em>"Changes at each level invalidate that level and all subsequent levels."</em></li>
</ul>
<p>A cache read costs about <strong>10% of the normal input-token price</strong>. So every masked prompt
that misses the cache pays close to <strong>ten times more</strong> on those tokens, and gives up the
faster first token too. On top of that, when a masked placeholder like <code>PERSON_42</code> leaks
into a tool call, the tool rejects it and the agent retries — and a study of coding agents
found that small prompt-wording changes can multiply token use <strong>2.4–7.4×</strong> with no gain in
success. (Again: not masking specifically, but the same mechanism.)</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-reframe-the-model-was-never-the-risk">The reframe: the model was never the risk<a href="https://development-wec.wiline.com/docs/news/mask-your-logs-not-your-prompts/#the-reframe-the-model-was-never-the-risk" class="hash-link" aria-label="Direct link to The reframe: the model was never the risk" title="Direct link to The reframe: the model was never the risk" translate="no">​</a></h2>
<p>Here's the part I had backwards. <strong>The problem was never that the model <em>sees</em> the data.</strong>
A model reading your prompt to answer it is just doing its job — that inference pass isn't
where data leaks. The real question is <strong>retention</strong>: what gets <em>stored</em>, and where.</p>
<p>Once you see it that way, the fix is obvious. Put the masking where storage actually happens
and where you're in control: <strong>your logs.</strong></p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="mask-your-logs-not-your-prompts">Mask your logs, not your prompts<a href="https://development-wec.wiline.com/docs/news/mask-your-logs-not-your-prompts/#mask-your-logs-not-your-prompts" class="hash-link" aria-label="Direct link to Mask your logs, not your prompts" title="Direct link to Mask your logs, not your prompts" translate="no">​</a></h2>
<p>The pattern — often called <strong>logging-only</strong> — is simple:</p>
<ul>
<li class="">The <strong>raw</strong> prompt reaches the model, so reasoning stays intact, the cache still hits, and
nothing gets refused.</li>
<li class="">Your PII masking runs <strong>only on the path to storage</strong> — the request/response logs and
traces your gateway writes.</li>
</ul>
<p>Think of it as three rungs, worst to best:</p>
<ol>
<li class=""><strong>Dumb masking</strong> (<code>[REDACTED]</code>) — destroys entity relationships.</li>
<li class=""><strong>Tokenization</strong> (<code>PERSON_42</code>) — keeps identities, but still triggers refusals and misses.</li>
<li class=""><strong>Logging-only</strong> — the model reads the real text; masking happens where the data rests.</li>
</ol>
<p>One honest caveat: whatever provider you send prompts to, it's worth knowing its retention
policy — that's a separate question from what the model reads, and it's the <em>right</em> place to
put that worry.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="if-you-run-your-own-gateway">If you run your own gateway<a href="https://development-wec.wiline.com/docs/news/mask-your-logs-not-your-prompts/#if-you-run-your-own-gateway" class="hash-link" aria-label="Direct link to If you run your own gateway" title="Direct link to If you run your own gateway" translate="no">​</a></h2>
<p>The nice part: this is a <strong>configuration</strong>, not a rewrite. If you've already
<a class="" href="https://development-wec.wiline.com/docs/tutorials/deploy-llm-gateway-wec-instance/">deployed a gateway on a WEC Instance</a>, the
guardrail runs in log-only mode — clean prompt out to the model, masked copy into the logs.
A follow-up tutorial will wire it up end to end.</p>
<p>Masking a prompt before the model buys you a compliance checkbox and quietly sells your
model's intelligence. Put the privacy work where the data actually rests — in the logs — and
let the model read the real thing.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="sources">Sources<a href="https://development-wec.wiline.com/docs/news/mask-your-logs-not-your-prompts/#sources" class="hash-link" aria-label="Direct link to Sources" title="Direct link to Sources" translate="no">​</a></h2>
<ul>
<li class=""><a href="https://arxiv.org/abs/2603.20208" target="_blank" rel="noopener noreferrer" class="">RedacBench: Can AI Erase Your Secrets? — arXiv 2603.20208</a> — 80.9% security / 37.6% utility at aggressive redaction</li>
<li class=""><a href="https://arxiv.org/abs/2406.10248" target="_blank" rel="noopener noreferrer" class="">On the Worst Prompt Performance of Large Language Models — arXiv 2406.10248</a> — 45.48% swing, 9.38% worst case</li>
<li class=""><a href="https://arxiv.org/abs/2512.02246" target="_blank" rel="noopener noreferrer" class="">DETAIL Matters: Prompt Specificity and Reasoning — arXiv 2512.02246</a></li>
<li class=""><a href="https://arxiv.org/abs/2608.01347" target="_blank" rel="noopener noreferrer" class="">Prompt-Induced Waste in Coding Agents — arXiv 2608.01347</a> — 2.4–7.4× token multiplication</li>
<li class=""><a href="https://www.reddit.com/r/LLMDevs/comments/1sgvjrg/deterministic_tokenization_vs_masking_for_pii_in/" target="_blank" rel="noopener noreferrer" class="">Deterministic tokenization vs. masking for PII: 109 tests — r/LLMDevs</a> — practitioner write-up (author sells a tokenization tool)</li>
<li class=""><a href="https://developers.openai.com/api/docs/guides/prompt-caching" target="_blank" rel="noopener noreferrer" class="">OpenAI — Prompt caching</a></li>
<li class=""><a href="https://platform.claude.com/docs/en/build-with-claude/prompt-caching" target="_blank" rel="noopener noreferrer" class="">Anthropic — Prompt caching</a></li>
</ul>]]></content>
        <author>
            <name>Rafael Fernandes</name>
            <uri>https://www.linkedin.com/in/rafaelmacariofernandes/</uri>
        </author>
        <category label="ai-news" term="ai-news"/>
        <category label="privacy" term="privacy"/>
        <category label="pii" term="pii"/>
        <category label="guardrails" term="guardrails"/>
        <category label="gateway" term="gateway"/>
    </entry>
    <entry>
        <title type="html"><![CDATA[Spec-Driven Development: is it the solution to Vibe Coding?]]></title>
        <id>https://development-wec.wiline.com/docs/news/spec-driven-development-solution-to-vibe-coding/</id>
        <link href="https://development-wec.wiline.com/docs/news/spec-driven-development-solution-to-vibe-coding/"/>
        <updated>2026-08-14T00:00:00.000Z</updated>
        <summary type="html"><![CDATA[A GitHub toolkit with 127k stars says you should write the spec before the code, and let the agent build from it. I ran it, then read the two engineers who tested it properly — and both reached for the same comparison: the last time our industry tried generating code from documents.]]></summary>
        <content type="html"><![CDATA[<div class="newsHero"><div class="newsHero__glow" aria-hidden="true"></div><span class="newsHero__eyebrow">Engineering practice · AI News</span><h2 class="newsHero__title">Spec-driven development, tested</h2><div class="newsHero__transition"><span class="newsHero__pill newsHero__pill--from">Write the prompt</span><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2.5" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-arrow-right newsHero__arrow" aria-hidden="true"><path d="M5 12h14"></path><path d="m12 5 7 7-7 7"></path></svg><span class="newsHero__pill newsHero__pill--to">Write the contract</span></div></div>
<p>Someone posted spec-driven development on LinkedIn this week as <em>the</em> answer to vibe coding —
to prompting an agent, half-understanding what you're building, and ending up with code you
can't vouch for. The linked toolkit has 127,000 stars and comes from GitHub itself. The pitch
lands.</p>
<p>So I installed it and pointed it at a deliberately trivial task. One of the three principles it
wrote for me was a dependency policy I never asked for — hold that thought.</p>
<p>Twenty minutes isn't a verdict, though. Two engineers have tested this properly, on real
problems, long enough for the seams to show. They used different tools, on different
continents, seven months apart — and both reached for the same comparison, unprompted: the last
time our industry tried to generate working code from documents. On the one question that
decides whether any of this survives contact with AI features, they flatly contradict each
other. Neither has a measurement.</p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="what-spec-driven-development-is">What spec-driven development is<a href="https://development-wec.wiline.com/docs/news/spec-driven-development-solution-to-vibe-coding/#what-spec-driven-development-is" class="hash-link" aria-label="Direct link to What spec-driven development is" title="Direct link to What spec-driven development is" translate="no">​</a></h2>
<p>If the term is new to you, the idea is simple. Instead of prompting an agent and iterating until
the code looks right, you write a structured specification first — what you're building, why,
and what "done" means — and the agent works from that document rather than from your prompt.
The spec, not the code, becomes the thing everyone points at.</p>
<p><strong>Spec Kit</strong> is GitHub's implementation. It installs into a repo once, not per task, and gives
your agent a set of commands: <code>constitution</code> to set project principles, then <code>specify</code>, <code>plan</code>,
<code>tasks</code>, <code>implement</code>. Nothing runs in the background — it drops templates and command
definitions into your project, and from then on you're talking to your agent. Closer to a linter
config than a product.</p>
<!-- -->
<p>Before any of that, though, comes <code>constitution</code> — written first, before anyone has seen the
problem, and everything downstream gets judged against it. Hold on to that too.</p>
<p>The setup step tells you what it's really aiming at:</p>
<p><span class="zoomImage__wrap"><img alt="Spec Kit&amp;#39;s setup prompt listing more than thirty coding agent integrations, including Claude Code, GitHub Copilot, Cursor, Gemini CLI, Devin, Grok and IBM Bob." src="https://development-wec.wiline.com/docs/assets/images/spec-kit-agents-f508d78c53fe54f1ad55cd522d110859.png" width="2236" height="1282" class="zoomImage " loading="lazy"><span class="zoomImage__badge" aria-hidden="true"><svg viewBox="0 0 24 24" width="16" height="16" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round"><circle cx="11" cy="11" r="7"></circle><path d="M21 21l-4.3-4.3"></path><path d="M11 8v6M8 11h6"></path></svg></span></span></p>
<p>Thirty-plus agents. This isn't a GitHub-only tool — it wants to be the layer every coding agent
plugs into, which goes a long way to explaining the star count.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="what-actually-happens-when-you-use-it">What actually happens when you use it<a href="https://development-wec.wiline.com/docs/news/spec-driven-development-solution-to-vibe-coding/#what-actually-happens-when-you-use-it" class="hash-link" aria-label="Direct link to What actually happens when you use it" title="Direct link to What actually happens when you use it" translate="no">​</a></h2>
<p><strong>Birgitta Böckeler</strong>, Distinguished Engineer at Thoughtworks, trialled three of these tools by
hand — Kiro, Spec Kit and Tessl — and
<a href="https://martinfowler.com/articles/exploring-gen-ai/sdd-3-tools.html" target="_blank" rel="noopener noreferrer" class="">published what she found</a>
in October 2025.</p>
<p>She asked Kiro to fix a small bug. It produced four user stories and sixteen acceptance
criteria, including — verbatim — <em>"As a developer, I want the transformation function to handle
edge cases gracefully, so that the system remains robust when new category formats are
introduced."</em> Her summary: <em>"like using a sledgehammer to crack a nut."</em></p>
<p>On Spec Kit with a real feature — a few days' work, by her own estimate — she never
finished the implementation, and reckons she could have built the thing by hand in the time she
spent reviewing artifacts. The line that will land with anyone who has done a code review:</p>
<blockquote>
<p>To be honest, I'd rather review code than all these markdown files.</p>
</blockquote>
<p>She also found the agent ignoring the documents meant to steer it. Spec Kit's research step
correctly catalogued existing classes; the agent then read those descriptions as a
specification and generated the classes again, as duplicates. And the reverse failure — the
agent going <em>"way overboard because it was too eagerly following instructions (e.g. one of the
constitution articles)."</em></p>
<p><strong>Alex Punnen</strong> went narrower and deeper, and
<a href="https://github.com/alexcpn/speckit_test" target="_blank" rel="noopener noreferrer" class="">published the whole transcript</a> with line citations.
Spec Kit v0.8.9, on a problem at real scale: querying US elevation data across 1,756 map tiles
— 23 billion measurements, 180GB.</p>
<p>His finding is subtler than "the tool invents things." He asked for principles covering code
quality, testing, consistency and performance, and says plainly: <em>"The principles themselves are
reasonable."</em> What went wrong, he argues, is that one of them quietly tilted every later
decision:</p>
<blockquote>
<p>New dependencies MUST be justified in writing: problem solved, alternatives considered,
license verified.</p>
</blockquote>
<p>Sensible in isolation. But as Punnen puts it, <em>"the stdlib option always wins ties because it
costs zero justification entries."</em> Four phases later the plan chose SQLite because — first
reason listed — <em>"Stdlib, zero new dependency… that's the cheapest possible answer."</em> Three of
the four rejected alternatives fell to that same rule. His verdict: <strong>"The constitution did the
rejecting; the agent was just the microphone."</strong></p>
<p>Then comes the part I found hardest to shake. Earlier in the process the agent had written a
five-minute performance budget into the spec — a number it guessed, with no measurement behind
it, filed under the label SC-008. Later, that guess came back as the reason a better design
couldn't be used: <em>"The transcode cost blows past SC-008."</em> Only under direct pushback did it
concede: <em>"You're right that I overweighted reason #2."</em> The better design, Punnen writes, needed
nothing that wasn't already in the spec, <em>"except the willingness to revise an arbitrary number
the spec itself produced."</em></p>
<p>His summary of that phase applies to the whole category:</p>
<blockquote>
<p>The artefact looks done because every template slot is filled — not because the engineering
question is answered.</p>
</blockquote>
<p>Two things make this more than a one-off. His constitution prompt was essentially <strong>GitHub's own
documented example</strong>, lightly adapted — he followed the quickstart. And the resistance he hit is
partly by design: GitHub's methodology document calls the constitution <em>"a set of <strong>immutable
principles</strong>,"</em> with a section headed <em>"The Power of Immutable Principles."</em></p>
<p>To be precise, the five-minute budget was not a constitutional principle — it was a success
criterion the agent generated downstream of one. The document never claims those are immutable.
But the number carried that authority anyway, and it took a human to dislodge it. The
methodology argues for fixing principles and says nothing about what happens when guesses
derived from them inherit the same standing.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-three-levels-nobody-agrees-on">The three levels nobody agrees on<a href="https://development-wec.wiline.com/docs/news/spec-driven-development-solution-to-vibe-coding/#the-three-levels-nobody-agrees-on" class="hash-link" aria-label="Direct link to The three levels nobody agrees on" title="Direct link to The three levels nobody agrees on" translate="no">​</a></h2>
<p>Böckeler's most useful contribution is a distinction the rest of the debate skips. "Spec-driven
development" covers three different practices:</p>
<figure class="stageFlow"><div class="stageFlow__track"><div class="stageFlow__card" style="background:rgba(var(--primary-rgb), 0.050);border-color:rgba(var(--primary-rgb), 0.250)"><span class="stageFlow__stage">Spec-first</span><span class="stageFlow__title">Spec written first, used for the task at hand</span><span class="stageFlow__tag">all tools do this</span></div><svg xmlns="http://www.w3.org/2000/svg" width="22" height="22" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2.5" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-arrow-right stageFlow__arrow" aria-hidden="true"><path d="M5 12h14"></path><path d="m12 5 7 7-7 7"></path></svg><div class="stageFlow__card" style="background:rgba(var(--primary-rgb), 0.160);border-color:rgba(var(--primary-rgb), 0.450)"><span class="stageFlow__stage">Spec-anchored</span><span class="stageFlow__title">Spec kept and maintained after the task</span><span class="stageFlow__tag">few tools reach</span></div><svg xmlns="http://www.w3.org/2000/svg" width="22" height="22" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2.5" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-arrow-right stageFlow__arrow" aria-hidden="true"><path d="M5 12h14"></path><path d="m12 5 7 7-7 7"></path></svg><div class="stageFlow__card" style="background:rgba(var(--primary-rgb), 0.270);border-color:rgba(var(--primary-rgb), 0.650)"><span class="stageFlow__stage">Spec-as-source</span><span class="stageFlow__title">Only the spec is edited; human never touches code</span><span class="stageFlow__tag">Tessl only</span></div></div><figcaption class="stageFlow__caption">Böckeler's taxonomy, October 2025. IBM published the same three names seven months later.</figcaption></figure>
<p>Her verdict on where the tools actually sit: <em>"All SDD approaches and definitions I've found are
spec-first, but not all strive to be spec-anchored or spec-as-source."</em> Including Spec Kit.
GitHub's methodology aspires far higher — <em>"Specifications don't serve code—code serves
specifications"</em> — but Spec Kit creates a <strong>branch per spec</strong>, so a spec lives for the lifetime
of a change request, not a feature. Her conclusion: <em>"spec-kit is still what I would call
spec-first only, not spec-anchored over time."</em></p>
<p>Worth knowing that IBM published <a href="https://www.ibm.com/think/topics/spec-driven-development" target="_blank" rel="noopener noreferrer" class="">the identical three-level
taxonomy</a> seven months later,
footnoted. If you've seen it credited to IBM, it's hers.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="what-i-got-on-a-trivial-task">What I got on a trivial task<a href="https://development-wec.wiline.com/docs/news/spec-driven-development-solution-to-vibe-coding/#what-i-got-on-a-trivial-task" class="hash-link" aria-label="Direct link to What I got on a trivial task" title="Direct link to What I got on a trivial task" translate="no">​</a></h2>
<p>I ran Spec Kit at commit <code>83883a2</code> on a deliberately minimal prompt — <em>"Principles for a small
Python utility. Keep it minimal — I have no strong constraints."</em></p>
<p><span class="zoomImage__wrap"><img alt="The Spec Kit constitution phase running in a terminal, generating principles from a one-line prompt." src="https://development-wec.wiline.com/docs/assets/images/spec-kit-constitution-2e19dee68aab5c7e2a098a49a54edc1b.png" width="2176" height="1288" class="zoomImage " loading="lazy"><span class="zoomImage__badge" aria-hidden="true"><svg viewBox="0 0 24 24" width="16" height="16" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round"><circle cx="11" cy="11" r="7"></circle><path d="M21 21l-4.3-4.3"></path><path d="M11 8v6M8 11h6"></path></svg></span></span></p>
<p>One of the three principles it wrote:</p>
<div class="language-markdown codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#393A34;--prism-background-color:#f6f8fa"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-markdown codeBlock_bY9V thin-scrollbar" style="color:#393A34;background-color:#f6f8fa"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#393A34"><span class="token title important punctuation" style="color:#393A34">###</span><span class="token title important"> II. Minimal Dependencies</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">Prefer the Python standard library. A third-party dependency MAY be added only when it</span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">removes clearly more complexity than it introduces, and MUST be recorded in the project's</span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">dependency file (e.g. </span><span class="token code-snippet code keyword" style="color:#00009f">`requirements.txt`</span><span class="token plain"> or </span><span class="token code-snippet code keyword" style="color:#00009f">`pyproject.toml`</span><span class="token plain">).</span><br></div></code></pre></div></div>
<p>A dependency policy, written in the MUST/MAY language of a formal standard, from a prompt where
I said I had no constraints. It isn't in
the local template or the skill file — but Article I of GitHub's nine constitutional articles
does ask for implementations <em>"with clear boundaries and <strong>minimal dependencies</strong>."</em> So it's
consistent with the published philosophy rather than invented on the spot. I can't tell you the
mechanism, only what went in and what came out.</p>
<p>One run, one trivial task, and it cost me nothing because nothing was at stake. Punnen's case
shows what this kind of bias costs when the problem is hard enough for it to be wrong. Mine only
shows it turns up unasked — which matters because, as Böckeler notes, Spec Kit's constitution is
its <strong>memory bank</strong>, <em>"a very powerful rules file"</em> applied to every change. An unrequested
preference doesn't sit in a document you'll discard. It becomes a standing rule.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="both-of-them-reached-for-the-1990s">Both of them reached for the 1990s<a href="https://development-wec.wiline.com/docs/news/spec-driven-development-solution-to-vibe-coding/#both-of-them-reached-for-the-1990s" class="hash-link" aria-label="Direct link to Both of them reached for the 1990s" title="Direct link to Both of them reached for the 1990s" translate="no">​</a></h2>
<p>Here's what convinced me this is worth taking seriously rather than dismissing or evangelising.</p>
<p>Böckeler, who worked on model-driven development early in her career, sees MDD:</p>
<blockquote>
<p>I wonder if spec-as-source, and even spec-anchoring, might end up with the downsides of both
MDD and LLMs: Inflexibility and non-determinism.</p>
</blockquote>
<p>Punnen, twenty years in telecom, reaches independently for Rational Rose and UML — <em>"treated as
the silver bullet of its decade: draw boxes, arrows, and diagrams, and the tool would magically
turn them into working code."</em></p>
<p>Neither cites the other. Different tools, different problems, different countries. Both land on
the same era: the last time our industry believed a document could be the source and code the
output.</p>
<p>Böckeler is careful about the comparison — <em>"I'm not nostalgic about my MDD experience."</em> Her
point is that today's tools drop the parts that made MDD painful: you no longer need a special
spec language or a purpose-built generator. What she wonders is whether the exchange is a good
one, since the old approach at least produced the same output every time.</p>
<p>Punnen names the trap underneath:</p>
<blockquote>
<p>The specification has to be very rigorous in the first place, but to become rigorous it needs
to be iteratively refined alongside the generated code.</p>
</blockquote>
<p>A Catch-22, in other words: on his account you can't write a rigorous spec for a problem you
don't yet understand, and understanding arrives while building. He traces the thought back to
Fred Brooks, forty years ago: <em>"descriptions of a software entity that abstract away its
complexity often abstract away its essence."</em></p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="where-they-disagree--and-why-it-matters-to-you">Where they disagree — and why it matters to you<a href="https://development-wec.wiline.com/docs/news/spec-driven-development-solution-to-vibe-coding/#where-they-disagree--and-why-it-matters-to-you" class="hash-link" aria-label="Direct link to Where they disagree — and why it matters to you" title="Direct link to Where they disagree — and why it matters to you" translate="no">​</a></h2>
<p>On one question these two are in direct opposition.</p>
<p>Böckeler, generating code repeatedly from one Tessl spec: <em>"I have seen the non-determinism in
action… an interesting exercise to iterate on the spec and make it more and more specific to
increase the repeatability of the code generation."</em></p>
<p>Punnen: <em>"This is not a major problem in practice. SDD frameworks act as structured prompts, and
modern models produce highly consistent outputs when guided by them."</em></p>
<p>GitHub takes Punnen's side and goes further, claiming <em>"Consistency Across LLMs: Different AI
models produce architecturally compatible code."</em> Not the same model twice — different models.
No evidence offered.</p>
<p>Two experienced engineers, opposite conclusions, a vendor claim stronger than either, and not a
number between them. It matters more than it looks. Keeping a spec and its code in step — the
spec-anchored idea — relies on automated tests to catch the drift. That works for deterministic
code, where the same input gives the same output.</p>
<p>Point it at an LLM feature and the check stops working. <code>assert response == expected</code> means
nothing when the response differs every run. Unless you swap tests for <strong>evals</strong>, where each
acceptance criterion becomes a scored assertion: did it classify correctly, did it return valid
JSON against the schema, did it refuse when it should have.</p>
<p>That bridge survives nondeterminism, and it's the harness I've spent several tutorials building
on the WEC Inference API. It also makes the disagreement <em>measurable</em>: write the spec, turn its
criteria into eval assertions, and run them across a long session to see whether adherence holds
or decays.</p>
<p>That's the next post.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="if-youre-going-to-try-it">If you're going to try it<a href="https://development-wec.wiline.com/docs/news/spec-driven-development-solution-to-vibe-coding/#if-youre-going-to-try-it" class="hash-link" aria-label="Direct link to If you're going to try it" title="Direct link to If you're going to try it" translate="no">​</a></h2>
<p>Punnen's practical takeaways are better than anything I'd invent, and he earned them:</p>
<ul>
<li class=""><strong>Treat "no new deps" rules as biases, not neutrals.</strong> If the right answer needs a dependency,
you'll have to defend it — the framework won't.</li>
<li class=""><strong>Treat generated success criteria as guesses</strong> until an engineer ratifies them. The agent will
quote them back at you later as if they were measured.</li>
<li class=""><strong>Read every clarify menu as a design proposal in disguise.</strong> If the option you want isn't
listed, that's the failure — not a prompt to pick the best of three.</li>
<li class=""><strong>Push back during clarify, not during plan.</strong> Plans are long, internally consistent, and
exhausting to revise.</li>
</ul>
<p>One of my own: <strong>write your principles with a date and a rationale, so a later you can supersede
them.</strong> IBM suggests treating specs as <em>"stackable versioned artifacts, like architecture
decision records"</em> — and an ADR is something you can mark superseded. The methodology calls
these principles immutable. Your problem isn't.</p>
<p>Add IBM's cost test, the sharpest practical line any of them wrote: <em>the cost of refining the
spec should always be lower than the cost of fixing misunderstandings in implementation. When
that balance flips, stop polishing and start building.</em></p>
<p>And check the org before you install. The repo doing the rounds on LinkedIn is a <strong>fork</strong> with
eleven stars; the real project is <a href="https://github.com/github/spec-kit" target="_blank" rel="noopener noreferrer" class=""><code>github/spec-kit</code></a>. That
matters more here than for most tools: Spec Kit's job is writing instruction files your agent
then obeys, and on first run you approve that folder in one keystroke without having read them.</p>
<hr>
<p>So does it replace vibe coding? Not the way the pitch suggests. It doesn't help when you don't
understand your problem — it helps when you understand it and communicate it badly. Those are
different failures, and only one of them has a template.</p>
<p>Both found real value in the early phases — Punnen rates the clarification step his most useful
of all — and both found the artifacts looking most authoritative exactly where the decisions
underneath them were most arbitrary. Punnen: <em>"they are at their most dangerous when they look
the most rigorous."</em></p>
<p>Böckeler reaches for a German compound word for it — <strong>Verschlimmbesserung</strong>. Making something
worse in the attempt of making it better.</p>
<p>The hard part was never writing the code.</p>]]></content>
        <author>
            <name>Rafael Fernandes</name>
            <uri>https://www.linkedin.com/in/rafaelmacariofernandes/</uri>
        </author>
        <category label="ai-news" term="ai-news"/>
        <category label="engineering-practice" term="engineering-practice"/>
        <category label="agents" term="agents"/>
        <category label="evals" term="evals"/>
        <category label="tooling" term="tooling"/>
    </entry>
    <entry>
        <title type="html"><![CDATA[Muse Glimmer: a 30B agentic model that runs on one GPU, no data center required]]></title>
        <id>https://development-wec.wiline.com/docs/news/muse-glimmer-30b-local-agentic-model/</id>
        <link href="https://development-wec.wiline.com/docs/news/muse-glimmer-30b-local-agentic-model/"/>
        <updated>2026-08-10T00:00:00.000Z</updated>
        <summary type="html"><![CDATA[Meta released a 30B multimodal model built for local agent workloads — beats Gemma4 and holds its own against Qwen3.6 on agentic benchmarks, and fits on a single consumer GPU. Here's what's actually new, and what it'd take for it to land on WEC.]]></summary>
        <content type="html"><![CDATA[<div class="newsHero"><div class="newsHero__glow" aria-hidden="true"></div><span class="newsHero__eyebrow">Models · AI News</span><h2 class="newsHero__title">A serious agent, no data center required</h2><div class="newsHero__transition"><span class="newsHero__pill newsHero__pill--from">Cloud-only agents</span><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2.5" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-arrow-right newsHero__arrow" aria-hidden="true"><path d="M5 12h14"></path><path d="m12 5 7 7-7 7"></path></svg><span class="newsHero__pill newsHero__pill--to">One GPU, fully local</span></div></div>
<p>Most "run it locally" model announcements come with an asterisk — smaller, weaker, a toy
version of the real thing. Meta's newest release doesn't: <strong>Muse Glimmer</strong>, a 30B
multimodal model built specifically for agentic work, fits on a single consumer GPU and
beats larger models on the benchmarks that actually measure agent behavior.</p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="architecture">Architecture<a href="https://development-wec.wiline.com/docs/news/muse-glimmer-30b-local-agentic-model/#architecture" class="hash-link" aria-label="Direct link to Architecture" title="Direct link to Architecture" translate="no">​</a></h2>
<p>Muse Glimmer is <strong>30B parameters total</strong>: a 2B ViT-style vision encoder bolted onto a
28B-parameter text decoder, 52 transformer layers using a hybrid attention pattern.
It's distilled from Meta's larger Muse Spark model — the capability of a bigger model,
compressed into something a single GPU can hold. Training data spans 100+ languages,
with a January 4, 2026 knowledge cutoff.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="multimodal-understanding">Multimodal understanding<a href="https://development-wec.wiline.com/docs/news/muse-glimmer-30b-local-agentic-model/#multimodal-understanding" class="hash-link" aria-label="Direct link to Multimodal understanding" title="Direct link to Multimodal understanding" translate="no">​</a></h2>
<ul>
<li class=""><strong>Text, image, and video</strong> — video comprehension up to 96 frames at 2 fps (no audio
track processed).</li>
<li class=""><strong>Open-ended object detection</strong> — it can locate and identify objects in a scene without
a predefined label set, rather than only recognizing a fixed category list.</li>
<li class=""><strong>131K+ token context window</strong>, long enough for extended agent sessions or large
documents without external chunking.</li>
</ul>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="agentic-tool-use">Agentic tool use<a href="https://development-wec.wiline.com/docs/news/muse-glimmer-30b-local-agentic-model/#agentic-tool-use" class="hash-link" aria-label="Direct link to Agentic tool use" title="Direct link to Agentic tool use" translate="no">​</a></h2>
<p>The headline feature: <strong>multimodal tool-calling with structured outputs</strong>. It can look at
an image and decide which function to call based on what it sees, not just parse text
instructions — e.g. inspecting a screenshot and calling the right API based on what's
rendered, not a text description of it. It also generates and executes code, and is
built with explicit failure-recovery behavior rather than assuming every tool call
succeeds on the first try.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="running-it-locally">Running it locally<a href="https://development-wec.wiline.com/docs/news/muse-glimmer-30b-local-agentic-model/#running-it-locally" class="hash-link" aria-label="Direct link to Running it locally" title="Direct link to Running it locally" translate="no">​</a></h2>
<ul>
<li class=""><strong>Fits on one consumer GPU</strong> — quantized variants run in 24–32GB of VRAM.</li>
<li class=""><strong>Day-0 support</strong> in <code>transformers</code>, <code>llama.cpp</code>, <code>vLLM</code>, and Inference Endpoints — no
waiting on community ports.</li>
<li class=""><strong>DFlash speculative decoding</strong> — up to 3x faster generation on supported hardware.</li>
<li class="">No cloud round-trip required for any of the above; everything runs on the box you own.</li>
</ul>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="licensing">Licensing<a href="https://development-wec.wiline.com/docs/news/muse-glimmer-30b-local-agentic-model/#licensing" class="hash-link" aria-label="Direct link to Licensing" title="Direct link to Licensing" translate="no">​</a></h2>
<p><strong>Apache 2.0</strong> — no commercial-use gate, no attribution requirement, no separate license
negotiation to run it in a product. Worth stating plainly because it's not a given right
now: some recent open-weight agentic models carry commercial-use restrictions that only
surface once you read the license text closely. This one doesn't.</p>
<p>Meta also ran safety evaluations for chemical/biological, cybersecurity, and
loss-of-control risk — all rated "moderate or lower."</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="how-it-compares">How it compares<a href="https://development-wec.wiline.com/docs/news/muse-glimmer-30b-local-agentic-model/#how-it-compares" class="hash-link" aria-label="Direct link to How it compares" title="Direct link to How it compares" translate="no">​</a></h2>
<p>Meta's own published numbers, against Gemma4-31B and Qwen3.6-27B:</p>
<table><thead><tr><th>Benchmark</th><th>Muse Glimmer</th><th>Gemma4</th><th>Qwen3.6</th></tr></thead><tbody><tr><td>MCP Atlas (agentic)</td><td><strong>75.5</strong></td><td>54.2</td><td>62.5</td></tr><tr><td>DeepSearch QA</td><td><strong>74.6</strong></td><td>61.7</td><td>71.1</td></tr><tr><td>GAIA2</td><td><strong>43.3</strong></td><td>36.4</td><td>40.0</td></tr><tr><td>SWE-Bench Pro</td><td><strong>51.2</strong></td><td>36.9</td><td>50.2</td></tr><tr><td>SWE-Bench Verified</td><td>76.0</td><td>66.6</td><td><strong>77.2</strong></td></tr></tbody></table>
<p>It's not a clean sweep — Qwen3.6 edges it on SWE-Bench Verified — but against Gemma4 the
gap is wide, and against Qwen3.6 it's competitive or ahead on most agentic-specific
benchmarks. <em>(Numbers from
<a href="https://huggingface.co/blog/muse-glimmer" target="_blank" rel="noopener noreferrer" class="">Meta's Muse Glimmer announcement</a>;
benchmarks are directional, not gospel.)</em></p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="why-it-matters">Why it matters<a href="https://development-wec.wiline.com/docs/news/muse-glimmer-30b-local-agentic-model/#why-it-matters" class="hash-link" aria-label="Direct link to Why it matters" title="Direct link to Why it matters" translate="no">​</a></h2>
<p>The self-hosted AI story has always had a quiet tax: the good agentic models needed real
infrastructure, so "run it yourself" often meant "run a worse version of it yourself." A
model that's genuinely built agent-first, ships permissively licensed, and fits on
hardware a single person can own is the gap closing in real time — the same trend this
whole tutorial series has been betting on.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="is-it-on-wec">Is it on WEC?<a href="https://development-wec.wiline.com/docs/news/muse-glimmer-30b-local-agentic-model/#is-it-on-wec" class="hash-link" aria-label="Direct link to Is it on WEC?" title="Direct link to Is it on WEC?" translate="no">​</a></h2>
<p>Not yet — and we're not going to pretend otherwise. We're running it through our own
model-evaluation suite now, the same one that benchmarks everything already on WEC
Models, before it earns a place in the catalog. If it holds up against what's already
there, expect it soon.</p>
<hr>
<p>📖 <strong>Sources:</strong> <a href="https://huggingface.co/blog/muse-glimmer" target="_blank" rel="noopener noreferrer" class="">Meta's Muse Glimmer announcement (Hugging Face)</a> · <a href="https://huggingface.co/meta-models/Muse-Glimmer-30B-GGUF" target="_blank" rel="noopener noreferrer" class="">Model card</a></p>]]></content>
        <author>
            <name>Rafael Fernandes</name>
            <uri>https://www.linkedin.com/in/rafaelmacariofernandes/</uri>
        </author>
        <category label="ai-news" term="ai-news"/>
        <category label="models" term="models"/>
        <category label="open-weight" term="open-weight"/>
        <category label="agents" term="agents"/>
        <category label="local-inference" term="local-inference"/>
    </entry>
    <entry>
        <title type="html"><![CDATA['Loop Engineering Is Dead' — and the Real Story Is Weirder Than the Obituary]]></title>
        <id>https://development-wec.wiline.com/docs/news/loop-engineering-graph-engineering-what-survives/</id>
        <link href="https://development-wec.wiline.com/docs/news/loop-engineering-graph-engineering-what-survives/"/>
        <updated>2026-08-05T00:00:00.000Z</updated>
        <summary type="html"><![CDATA[In June, 'loop engineering' got a name. Six weeks later it was declared dead — and the industry answered with vendor guides, competing definitions, and a wave of SEO. Here's what loop and graph engineering actually mean, why a free MIT textbook settles the argument on page 189, and what a six-week hype cycle should teach you about what to learn.]]></summary>
        <content type="html"><![CDATA[<div class="newsHero"><div class="newsHero__glow" aria-hidden="true"></div><span class="newsHero__eyebrow">Architecture · AI News</span><h2 class="newsHero__title">Loops, graphs, and the six-week obituary</h2><div class="newsHero__transition"><span class="newsHero__pill newsHero__pill--from">Loop engineering</span><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2.5" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-arrow-right newsHero__arrow" aria-hidden="true"><path d="M5 12h14"></path><path d="m12 5 7 7-7 7"></path></svg><span class="newsHero__pill newsHero__pill--to">Graph engineering</span></div></div>
<p>In June 2026, the AI world got a new buzzword: <strong>loop engineering</strong> — roughly, disciplined
design of a single agent's tool-calling loop. Six weeks later it was supposedly replaced by
<strong>graph engineering</strong> — wiring up several agents at once — killed by twelve words that 3.1
million people saw:</p>
<div class="xEmbed"><blockquote class="twitter-tweet" data-dnt="true" data-conversation="none"><p lang="en" dir="ltr"></p><p>Are we still talking loops or did we shift to graphs yet?</p><p></p>— <!-- -->Peter Steinberger 🦞<!-- --> (@<!-- -->steipete<!-- -->) <a href="https://twitter.com/steipete/status/2078277297791189132">July 18, 2026</a></blockquote></div>
<p>Don't know what loop engineering is? Don't worry — neither did most of the people
declaring it dead. Here are both terms, how a name became an obituary in six weeks, and
the twist nobody checked before writing about it.</p>
<!-- -->
<p>The split shows up best on a job with <strong>independent work that can run at once</strong> and where a
wrong answer is expensive — a <strong>research briefing</strong>: <em>"Pull together what we know about X —
the open web, our own docs, and the repo — and give me a one-page brief with sources."</em></p>
<p>The <a class="" href="https://development-wec.wiline.com/docs/tutorials/deploy-openclaw-docker-compose/">OpenClaw</a> agent you stand up in these
tutorials works the <strong>loop</strong> way — one agent cycling through its tools, one turn at a time.
Graph engineering wires up a team of specialists instead. Same job, two shapes:</p>
<div class="newsThemedWrap newsThemedWrap--light"><p><span class="zoomImage__wrap"><img alt="Loop: one agent cycling through its tools in sequence. Graph: many specialist agents wired into a network, each owning one tool." src="https://development-wec.wiline.com/docs/assets/images/loop-vs-graph-light-d1fd8017587aa7c27d81ad36a7ec1f0b.png" width="2564" height="1751" class="zoomImage " loading="lazy"><span class="zoomImage__badge" aria-hidden="true"><svg viewBox="0 0 24 24" width="16" height="16" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round"><circle cx="11" cy="11" r="7"></circle><path d="M21 21l-4.3-4.3"></path><path d="M11 8v6M8 11h6"></path></svg></span></span></p></div>
<div class="newsThemedWrap newsThemedWrap--dark"><p><span class="zoomImage__wrap"><img alt="Loop: one agent cycling through its tools in sequence. Graph: many specialist agents wired into a network, each owning one tool." src="https://development-wec.wiline.com/docs/assets/images/loop-vs-graph-dark-f9ef88bb5be59e4d6a1396b3985067fd.png" width="2564" height="1751" class="zoomImage " loading="lazy"><span class="zoomImage__badge" aria-hidden="true"><svg viewBox="0 0 24 24" width="16" height="16" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round"><circle cx="11" cy="11" r="7"></circle><path d="M21 21l-4.3-4.3"></path><path d="M11 8v6M8 11h6"></path></svg></span></span></p></div>
<p>Read the <strong>loop</strong> on the left as a wheel: one agent visits each tool, then comes back
around. Push it past simple tasks and two cracks appear: it's <strong>slow</strong> (three independent
searches still run back-to-back, one worker at a time), and it's <strong>hard to trust</strong> (the same
agent searched, read, and wrote the brief, so a wrong claim has no address — you can't tell
if the search was thin or the agent invented it).</p>
<p>The <strong>graph</strong> on the right fixes both: specialists split the job, a planner routes it. The
searches now fire <strong>at the same time</strong> — the brief lands in the time of the slowest one, not
the sum — and each agent owns one source, so a bad claim has an address. A dedicated
<strong>fact-check</strong> node, the step a rushed loop skips, becomes a real gate.</p>
<p>None of that is free: one prompt becomes a planner, four agents, and a verifier — a new
failure mode where the <em>coordination itself</em> can break. Trust problem traded for a plumbing
problem. And every box in the graph still runs its own loop inside: <strong>a graph contains
loops</strong> — that's the whole point.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="both-terms-in-one-picture">Both terms, in one picture<a href="https://development-wec.wiline.com/docs/news/loop-engineering-graph-engineering-what-survives/#both-terms-in-one-picture" class="hash-link" aria-label="Direct link to Both terms, in one picture" title="Direct link to Both terms, in one picture" translate="no">​</a></h2>
<!-- -->
<p><strong>Loop engineering</strong> is about one agent. An AI agent is just a model in a <code>while</code> loop
with tools: give it a goal, it picks a tool, your code runs it, the result goes back in,
repeat. ChatGPT answering a question is one call. An agent that edits a file, runs the
tests, sees them fail and tries again — that's the loop, five times over. Loop engineering
is designing that cycle deliberately: what triggers it, who verifies the work, when it
stops, what happens on failure. The slogan is a good one — <em>the intelligence lives in the
model, the reliability lives in the loop.</em></p>
<p><strong>Graph engineering</strong> is about several agents. Nodes are agents, edges are dependencies,
plus the shared state and permissions between them. If loop engineering is "make one agent
reliable," graph engineering is "make ten agents not step on each other."</p>
<p>Different problems, different scale. So <em>"loop engineering is dead, graph engineering
replaced it"</em> was never a sensible sentence — it's like saying wheels replaced cars.</p>
<p>The name for the first one arrived in June 2026, from
<a href="https://x.com/addyosmani/status/2064127981161959567" target="_blank" rel="noopener noreferrer" class="">Addy Osmani</a>. It caught on fast
because everyone was already doing it, badly, with no shared vocabulary. Remember that
date.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-textbook-settles-it">The textbook settles it<a href="https://development-wec.wiline.com/docs/news/loop-engineering-graph-engineering-what-survives/#the-textbook-settles-it" class="hash-link" aria-label="Direct link to The textbook settles it" title="Direct link to The textbook settles it" translate="no">​</a></h2>
<p>Here's the thing nobody checked before writing a guide: "graph" isn't a 2026 coinage, and
the receipts are sitting in a free MIT textbook. From <em>Mathematics for Computer Science</em>
(6.042J), chapter 6, page 189 — <strong>Definition 6.1.1</strong>:</p>
<blockquote>
<p>"A directed graph G = (V, E) consists of a nonempty set of nodes V and a set of directed
edges E… A directed graph is <strong>simple</strong> if it has no <strong>loops</strong> (that is, edges of the
form u→u) and no multiple edges."</p>
</blockquote>
<p>The formal definition of a graph <em>already contains the word loop</em>, as an ordinary feature
of graphs — not their opposite. Definition 6.1.2 adds that a <strong>cycle</strong> is a walk that
returns to where it started, which is what a loop is.</p>
<p><strong>A loop isn't the opposite of a graph. It's a graph that comes back to an earlier node.</strong>
Mathematics has known this since Euler walked around Königsberg in 1736.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="what-actually-happened-in-july">What actually happened in July<a href="https://development-wec.wiline.com/docs/news/loop-engineering-graph-engineering-what-survives/#what-actually-happened-in-july" class="hash-link" aria-label="Direct link to What actually happened in July" title="Direct link to What actually happened in July" translate="no">​</a></h2>
<p>The fuse wasn't an engineering discovery — it was <strong>two product launches colliding.</strong>
DeepLearning.AI released a course on knowledge graphs with multi-agent systems, taught by
Neo4j's Andreas Kollegger. Around the same time, Linear shipped an agent feature called —
of all things — <strong>Loops</strong>. Two companies, two unrelated products, two words that sounded
like rival philosophies.</p>
<p>Then Steinberger, who built <a class="" href="https://development-wec.wiline.com/docs/tutorials/deploy-openclaw-docker-compose/">OpenClaw</a>, posted
his twelve words at 9:34 PM on July 17. <strong>Four and a half hours later</strong> Hamel Husain
published an X Article titled <em>"Loop Engineering Is Dead. Enter Graph Engineering,"</em> and
<a href="https://x.com/svpino/status/2078516761318584774" target="_blank" rel="noopener noreferrer" class="">Santiago Valdarrama</a> picked it up the
same day. Within 48 hours the new term had three competing definitions and a wave of
copycat posts.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-part-that-should-have-ended-it">The part that should have ended it<a href="https://development-wec.wiline.com/docs/news/loop-engineering-graph-engineering-what-survives/#the-part-that-should-have-ended-it" class="hash-link" aria-label="Direct link to The part that should have ended it" title="Direct link to The part that should have ended it" translate="no">​</a></h3>
<p><strong>Neither founding post was serious</strong> — the writers who tracked the cycle say so outright.
Louis-François Bouchard put it plainly: <em>"my whole feed decided we have a new discipline.
To be honest, both tweets were jokes."</em> Husain's own follow-up only widened the wink,
promising the piece was <em>"not what you think it is."</em></p>
<p>And almost nobody could check: the article sat behind X's Premium paywall. A paradigm's
founding text was something few of the people citing it had actually read. What spread
wasn't an argument — it was a shape: a punchy title, a wink right behind it, and an industry
that answered a joke by writing documentation for it.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-replies-were-smarter-than-the-announcements">The replies were smarter than the announcements<a href="https://development-wec.wiline.com/docs/news/loop-engineering-graph-engineering-what-survives/#the-replies-were-smarter-than-the-announcements" class="hash-link" aria-label="Direct link to The replies were smarter than the announcements" title="Direct link to The replies were smarter than the announcements" translate="no">​</a></h2>
<p>The sharpest response came from outside the agent crowd. <strong>David Khourshid</strong> created
<strong>XState</strong> — a widely-used library for modeling exactly this kind of state machine in
code — and has been modelling these structures for a decade:</p>
<div class="xEmbed"><blockquote class="twitter-tweet" data-dnt="true" data-conversation="none"><p lang="en" dir="ltr"></p><p>First it was loops. Now it's graphs. Next month it'll be something else. Here's the thing: we're constantly rediscovering decades-old software engineering patterns and repackaging them as innovations or whatever, instead of just applying what we've already known for a long time.</p><p></p>— <!-- -->David Khourshid<!-- --> (@<!-- -->DavidKPiano<!-- -->) <a href="https://twitter.com/DavidKPiano/status/2079209887158989231">July 20, 2026</a></blockquote></div>
<p>His explanation takes two minutes. A state machine answers one question — <em>given the
current state, when an event occurs, what is the next state?</em> Draw it and states become
nodes, transitions become edges. Then he turns it on the July argument:</p>
<blockquote>
<p>"Surprise… loops are graphs: directed, cyclic ones."</p>
</blockquote>
<p>A loop, he notes, is the most basic state machine there is: two states, <code>looping</code> and
<code>done</code>. He posted a diagram of it captioned <strong>"This is the silly thing you all hyped for
weeks."</strong></p>
<!-- -->
<p>That's the object six weeks of discourse was about. And real agents — the ones
that retry, backtrack, wait for a human — were never sequential and never acyclic to begin
with (<em>"sorry, DAG lovers"</em>).</p>
<p>The creator of LangChain got there from the opposite direction, his LangGraph being the
implementation everyone kept pointing at:</p>
<div class="xEmbed"><blockquote class="twitter-tweet" data-dnt="true" data-conversation="none"><p lang="en" dir="ltr"></p><p>So i didn't really know what graph engineering is, and i still don't really… but it's basically just langgraph?</p><p></p>— <!-- -->Harrison Chase<!-- --> (@<!-- -->hwchase17<!-- -->) <a href="https://twitter.com/hwchase17/status/2079219804951683380">July 20, 2026</a></blockquote></div>
<p>Four days later he and Sydney Runkle published a longer answer,
<a href="https://www.langchain.com/blog/3-years-of-graph-engineering-with-langgraph" target="_blank" rel="noopener noreferrer" class=""><em>3 Years of Graph Engineering with LangGraph</em></a>,
which is the most useful thing written during the whole episode. Their opening line is
the best description of the phenomenon I've read:</p>
<blockquote>
<p>"It's the latest term to come out of X's AI content factory, joining prompt engineering,
context engineering, harness engineering, and loop engineering."</p>
</blockquote>
<p>And then, from the people whose product is literally the graph:</p>
<blockquote>
<p>"Loops are simple graphs. Loop engineering isn't an alternative to graphs, so much as a
simple version of them."</p>
</blockquote>
<p>They also confirm the thing most July posts got backwards: <strong>production agent graphs are
usually not DAGs.</strong> Real agents retry, ask for missing information, revise after
validation, pause for a human. Cycles aren't a design flaw to be engineered out — they're
the job.</p>
<p>The clearest framing of all, though, came from a reply:</p>
<div class="xEmbed"><blockquote class="twitter-tweet" data-dnt="true" data-conversation="none"><p lang="en" dir="ltr"></p><p>Graph engineering is deciding where the work is allowed to go. Loop engineering is making the work get better each time it runs. Graph is the rails. Loop is the motor. Rails keep you from crashing. The motor is what actually moves.</p><p></p>— <!-- -->Eric Osiu<!-- --> (@<!-- -->ericosiu<!-- -->) <a href="https://twitter.com/ericosiu/status/2079991948106957131">July 22, 2026</a></blockquote></div>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="where-it-stands-today">Where it stands today<a href="https://development-wec.wiline.com/docs/news/loop-engineering-graph-engineering-what-survives/#where-it-stands-today" class="hash-link" aria-label="Direct link to Where it stands today" title="Direct link to Where it stands today" translate="no">​</a></h2>
<p>It's August 5, about two and a half weeks after the obituary. Nobody won. The debate
didn't resolve and it didn't die — <strong>it got absorbed by content marketing.</strong> Every
agent-infrastructure vendor now has a "definitive guide to graph engineering," and behind
them a long tail of SEO pages saying the same thing in the same order.</p>
<figure class="stageFlow"><div class="stageFlow__track"><div class="stageFlow__card" style="background:rgba(var(--primary-rgb), 0.050);border-color:rgba(var(--primary-rgb), 0.250)"><span class="stageFlow__stage">June 7</span><span class="stageFlow__title">A name appears</span><span class="stageFlow__tag">it describes something real</span></div><svg xmlns="http://www.w3.org/2000/svg" width="22" height="22" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2.5" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-arrow-right stageFlow__arrow" aria-hidden="true"><path d="M5 12h14"></path><path d="m12 5 7 7-7 7"></path></svg><div class="stageFlow__card" style="background:rgba(var(--primary-rgb), 0.123);border-color:rgba(var(--primary-rgb), 0.383)"><span class="stageFlow__stage">Weeks 2–5</span><span class="stageFlow__title">Everyone adopts it</span><span class="stageFlow__tag">meaning drifts</span></div><svg xmlns="http://www.w3.org/2000/svg" width="22" height="22" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2.5" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-arrow-right stageFlow__arrow" aria-hidden="true"><path d="M5 12h14"></path><path d="m12 5 7 7-7 7"></path></svg><div class="stageFlow__card" style="background:rgba(var(--primary-rgb), 0.197);border-color:rgba(var(--primary-rgb), 0.517)"><span class="stageFlow__stage">July 17</span><span class="stageFlow__title">Declared dead</span><span class="stageFlow__tag">apparently as a joke</span></div><svg xmlns="http://www.w3.org/2000/svg" width="22" height="22" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2.5" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-arrow-right stageFlow__arrow" aria-hidden="true"><path d="M5 12h14"></path><path d="m12 5 7 7-7 7"></path></svg><div class="stageFlow__card" style="background:rgba(var(--primary-rgb), 0.270);border-color:rgba(var(--primary-rgb), 0.650)"><span class="stageFlow__stage">August</span><span class="stageFlow__title">Vendors publish guides</span><span class="stageFlow__tag">the term becomes a funnel</span></div></div><figcaption class="stageFlow__caption">Naming to obituary to lead generation, in six weeks. No benchmark ran at any point.</figcaption></figure>
<p>Nothing was <em>learned</em> in that time. No benchmark ran, no production system proved one
approach beat the other. Two companies shipped unrelated products, two well-followed
developers made a joke, and a lot of people agreed on a word — then agreed on a different
word.</p>
<p>That has a real cost. If you tried to keep up by adopting each term as it trended, you
rewrote your architecture twice in July for reasons that were social, not technical.
Meanwhile the engineer who ignored both posts and spent that month adding stop rules and
typed state came out ahead, because those mattered under any label.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-failure-nobodys-selling-a-guide-for">The failure nobody's selling a guide for<a href="https://development-wec.wiline.com/docs/news/loop-engineering-graph-engineering-what-survives/#the-failure-nobodys-selling-a-guide-for" class="hash-link" aria-label="Direct link to The failure nobody's selling a guide for" title="Direct link to The failure nobody's selling a guide for" translate="no">​</a></h2>
<p>One idea from this month deserves more attention than the naming war, and it comes from
<a href="https://www.linkedin.com/pulse/what-graph-engineering-really-towards-artificial-intelligence-e79ic/" target="_blank" rel="noopener noreferrer" class="">Towards AI's breakdown</a>:
<strong>more agents doesn't mean more judgement.</strong></p>
<p>Twenty agents running the same model, reading the same flawed context and checking against
the same broken metric will agree with each other at industrial scale. Worse is the
circular version: one agent checks a report against another report, an audit agent checks
both against a dashboard, and the dashboard was built from the same data. Every node
agrees. Nothing touched reality. The system <em>looks</em> well governed, because the diagram has
reviewers everywhere.</p>
<p>The fix is what they call <strong>reality anchors</strong> — evidence from outside the agent system.
Tests that actually ran. Money that reached the bank. Customers who stayed. Rules the
optimiser can't quietly rewrite. Their line is the one I'd put on the wall:</p>
<blockquote>
<p>"Without anchors, a graph is a larger hallucination with better project management."</p>
</blockquote>
<p>The same trap exists one level down, inside a single loop: if the agent that does the work
also decides whether the work is good, you've built an expensive machine for agreeing with
itself. A verifier only counts if it can actually fail you.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="what-to-actually-do">What to actually do<a href="https://development-wec.wiline.com/docs/news/loop-engineering-graph-engineering-what-survives/#what-to-actually-do" class="hash-link" aria-label="Direct link to What to actually do" title="Direct link to What to actually do" translate="no">​</a></h2>
<p>Khourshid's line is the filter, and it isn't cynicism: <em>next month it'll be something
else.</em> When the next term lands, the question is never <em>"is this the new paradigm?"</em> It's
<strong>"what specific failure does this name, and do I have that failure yet?"</strong></p>
<p>Diagnose where your bottleneck actually is:</p>
<ul>
<li class=""><strong>One agent degrading over a long session</strong> — forgetting, repeating tool calls, burning
tokens? That's a <strong>loop</strong> problem: stop rules, what tool output re-enters context,
whether errors are surfaced or swallowed. No graph framework fixes any of it.</li>
<li class=""><strong>Several agents duplicating work or deadlocking on shared state?</strong> That's a
<strong>coordination</strong> problem, and it needs an explicit control plane whatever you call it.</li>
</ul>
<p>And a case for neither: if the task is genuinely open-ended — research, exploration —
forcing it into fixed paths is the wrong move. The LangChain team makes this point against
their own product: they built early deep research on predefined LangGraph workflows and
then moved to a looser agentic loop, and GPT Researcher did the same. Structure you
haven't earned costs you.</p>
<p>And whichever you have, these outlive the vocabulary: every loop needs a stop rule; state
needs a shape and save points; acting nodes need permission boundaries and a human gate on
anything irreversible; add complexity only when a real failure asks for it.</p>
<p>If you want the two concepts that pay off across all of it, take Khourshid's
recommendation over any of July's vocabulary: <strong>state machines and the actor model</strong>. Both
predate this argument by decades and will outlive whatever replaces it — probably in about
six weeks.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="sources">Sources<a href="https://development-wec.wiline.com/docs/news/loop-engineering-graph-engineering-what-survives/#sources" class="hash-link" aria-label="Direct link to Sources" title="Direct link to Sources" translate="no">​</a></h2>
<ul>
<li class=""><a href="https://x.com/addyosmani/status/2064127981161959567" target="_blank" rel="noopener noreferrer" class="">Addy Osmani, naming "loop engineering" (June 2026) — X</a></li>
<li class=""><a href="https://x.com/steipete/status/2078277297791189132" target="_blank" rel="noopener noreferrer" class="">Peter Steinberger's post — X</a></li>
<li class="">Hamel Husain, "Loop Engineering Is Dead. Enter Graph Engineering" — X Article, July 18,
2026 (behind X Premium) · <a href="https://x.com/HamelHusain/status/2078348097697263855" target="_blank" rel="noopener noreferrer" class="">his public follow-up</a></li>
<li class=""><a href="https://x.com/svpino/status/2078516761318584774" target="_blank" rel="noopener noreferrer" class="">Santiago Valdarrama — X</a></li>
<li class=""><a href="https://x.com/DavidKPiano/status/2079209887158989231" target="_blank" rel="noopener noreferrer" class="">David Khourshid (XState), "State machines in 2 minutes" — X</a></li>
<li class=""><a href="https://x.com/hwchase17/status/2079219804951683380" target="_blank" rel="noopener noreferrer" class="">Harrison Chase (LangChain) — X</a> · <a href="https://x.com/ericosiu/status/2079991948106957131" target="_blank" rel="noopener noreferrer" class="">Eric Osiu — X</a></li>
<li class=""><a href="https://www.louisbouchard.ai/graph-engineering-explained/" target="_blank" rel="noopener noreferrer" class="">Louis-François Bouchard, "Graph Engineering Explained: What Actually Changed"</a></li>
<li class=""><a href="https://ai.gopubby.com/two-engineers-made-a-joke-about-graph-engineering-six-days-later-it-had-courses-3c082a5900fd" target="_blank" rel="noopener noreferrer" class="">"Two Engineers Made a Joke About Graph Engineering. Six Days Later It Had Courses." — AI Advances</a></li>
<li class=""><a href="https://smartscope.blog/en/blog/graph-engineering-loop-engineering-logic-review/" target="_blank" rel="noopener noreferrer" class="">SmartScope, on whether the "obituary" is true</a></li>
<li class=""><a href="https://ocw.mit.edu/courses/6-042j-mathematics-for-computer-science-fall-2010/e6db7638031b754f5f68012946af4763_MIT6_042JF10_chap06.pdf" target="_blank" rel="noopener noreferrer" class="">MIT 6.042J <em>Mathematics for Computer Science</em>, Ch. 6 "Directed Graphs" (free PDF)</a> — Definitions 6.1.1–6.1.2, pp. 189–191</li>
</ul>]]></content>
        <author>
            <name>Rafael Fernandes</name>
            <uri>https://www.linkedin.com/in/rafaelmacariofernandes/</uri>
        </author>
        <category label="ai-news" term="ai-news"/>
        <category label="agents" term="agents"/>
        <category label="architecture" term="architecture"/>
        <category label="orchestration" term="orchestration"/>
        <category label="engineering-practice" term="engineering-practice"/>
    </entry>
    <entry>
        <title type="html"><![CDATA[MCP Just Shipped Its Biggest Update Ever — Here's What Actually Changes for AI Agent Engineers]]></title>
        <id>https://development-wec.wiline.com/docs/news/mcp-2026-07-28-spec/</id>
        <link href="https://development-wec.wiline.com/docs/news/mcp-2026-07-28-spec/"/>
        <updated>2026-07-28T00:00:00.000Z</updated>
        <summary type="html"><![CDATA[The 2026-07-28 MCP specification rips out sessions and rewrites authorization. If you build or run MCP servers, this changes your infrastructure, your auth flow, and your deprecation clock — whether you asked for it or not.]]></summary>
        <content type="html"><![CDATA[<figure class="newsHero newsHero--image"><span class="newsHero__chip">Protocols · AI News</span><img src="https://development-wec.wiline.com/docs/img/news/mcp-drops-sessions-retires-three-core-features-16x9.webp" alt="Model Context Protocol — the 2026-07-28 specification" loading="eager"></figure>
<p>The Model Context Protocol just had its biggest revision since launch. The
2026-07-28 specification doesn't add a feature — it rewrites how every MCP
server talks to every client. If you've deployed an MCP server anywhere past
"runs on my laptop," this one touches your infrastructure, not just your
changelog.</p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="what-actually-shipped">What actually shipped<a href="https://development-wec.wiline.com/docs/news/mcp-2026-07-28-spec/#what-actually-shipped" class="hash-link" aria-label="Direct link to What actually shipped" title="Direct link to What actually shipped" translate="no">​</a></h2>
<p>The headline change: <strong>MCP is now stateless at the protocol layer.</strong> Every
request carries its own protocol version, client identity, and capabilities —
there's no more <code>initialize</code>/<code>initialized</code> handshake, no session ID, no
requirement that request N+1 lands on the same server instance that handled
request N. Lead maintainer David Soria Parra called it
<a href="https://blog.modelcontextprotocol.io/posts/2026-07-28/" target="_blank" rel="noopener noreferrer" class="">"a leap in serving scalable MCP servers"</a>,
built on 18 months of running MCP past the local-tool stage.</p>
<p>That one change unlocks a chain of practical ones:</p>
<ul>
<li class=""><strong>Plain load balancing.</strong> A server that used to need sticky sessions and a
shared session store can sit behind an ordinary round-robin balancer.</li>
<li class=""><strong>Header-based routing.</strong> Requests now carry <code>Mcp-Method</code> and <code>Mcp-Name</code>
HTTP headers, so gateways and firewalls can route and rate-limit by
inspecting headers instead of parsing every JSON body.</li>
<li class=""><strong>Cacheable list results.</strong> <code>tools/list</code>, <code>prompts/list</code>, <code>resources/list</code>,
and <code>resources/read</code> now return <code>ttlMs</code> and <code>cacheScope</code>, so clients know
how long they're allowed to skip re-fetching.</li>
<li class=""><strong>Multi Round-Trip Requests (MRTR).</strong> The old server-initiated,
held-open-stream pattern for mid-call input (confirmations, missing
parameters) is gone. A server now returns <code>resultType: "input_required"</code>;
the client retries the same call with <code>inputResponses</code> filled in — no
persistent connection required.</li>
</ul>
<p>State didn't disappear, it just became explicit: if your tool genuinely needs
it, it mints a handle and hands it back to the client to pass in on the next
call, instead of hiding it in the transport layer.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="authorization-got-a-real-rewrite-not-a-patch">Authorization got a real rewrite, not a patch<a href="https://development-wec.wiline.com/docs/news/mcp-2026-07-28-spec/#authorization-got-a-real-rewrite-not-a-patch" class="hash-link" aria-label="Direct link to Authorization got a real rewrite, not a patch" title="Direct link to Authorization got a real rewrite, not a patch" translate="no">​</a></h2>
<p>This is the part that should get an AI engineer's attention before the
protocol change does. MCP authorization now aligns with OAuth 2.1 and OpenID
Connect instead of leaving it to each implementer to wire up their own
version, <a href="https://workos.com/blog/mcp-2026-spec-agent-authentication" target="_blank" rel="noopener noreferrer" class="">as WorkOS breaks down</a>:</p>
<ul>
<li class="">Servers must implement <strong>OAuth 2.0 Protected Resource Metadata</strong> (RFC 9728)
for discovery and <strong>Resource Indicators</strong> (RFC 8707) so a token minted for
one MCP server can't be replayed against another.</li>
<li class=""><strong>Issuer verification is now mandatory</strong> — clients must check the <code>iss</code>
parameter before redeeming a code, closing the authorization-server
mix-up class of bugs.</li>
<li class=""><strong>Client ID Metadata Documents (CIMD)</strong> replace Dynamic Client Registration
as the preferred path (DCR still works, for now).</li>
</ul>
<p>The practical upshot: the "confused deputy" problem — a tool call executing
with credentials meant for a different server — gets closed at the protocol
level instead of depending on every server author to remember to check.
If you're running multiple MCP servers behind one gateway (which, per our own
<a class="" href="https://development-wec.wiline.com/docs/news/litellm-rust-gateway/">gateway coverage</a>, is where this is heading for
everyone), this is the part of the update that actually reduces your risk
surface, not just your ops burden.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-honest-part-this-breaks-things-and-the-maintainers-say-so">The honest part: this breaks things, and the maintainers say so<a href="https://development-wec.wiline.com/docs/news/mcp-2026-07-28-spec/#the-honest-part-this-breaks-things-and-the-maintainers-say-so" class="hash-link" aria-label="Direct link to The honest part: this breaks things, and the maintainers say so" title="Direct link to The honest part: this breaks things, and the maintainers say so" translate="no">​</a></h2>
<p>Nothing here is backward compatible by accident. Soria Parra, in
<a href="https://www.theregister.com/devops/2026/07/23/model-context-protocol-prepares-to-break-with-its-stateful-past/5276722" target="_blank" rel="noopener noreferrer" class="">The Register's reporting</a>,
didn't sugarcoat it: <em>"If you built your own implementation, it's going to be
a lot of uplift to make this correct,"</em> and the stateless redesign, by his own
admission, <em>"makes things on the wire a bit more complicated than they used
to be"</em> even as it removes session state. Roots, Sampling, and Logging are
now formally deprecated — Sampling for confusing semantics, Roots as "a very
niche thing," Logging for being excessively verbose — with a <strong>twelve-month
minimum window</strong> before they're actually removed, alongside the legacy
HTTP+SSE transport.</p>
<p>Read charitably, this is a protocol growing up: a formal deprecation policy
means you get a year of notice instead of a surprise break. Read skeptically
— and The Register does — this is fixing problems that only showed up once
MCP left local dev tooling for cloud deployment, which says something about
how much load-bearing infrastructure got built on the earlier design before
anyone stress-tested it at scale. Both readings are true at once. That's not
a reason to panic; it's a reason to actually read the migration guide before
your integration tests do it for you.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="why-this-matters-for-you-specifically">Why this matters for you, specifically<a href="https://development-wec.wiline.com/docs/news/mcp-2026-07-28-spec/#why-this-matters-for-you-specifically" class="hash-link" aria-label="Direct link to Why this matters for you, specifically" title="Direct link to Why this matters for you, specifically" translate="no">​</a></h2>
<p>If you're building agents against MCP servers you don't control, you likely
notice nothing immediately — Tier 1 SDKs (TypeScript, Python, Go, C#, with
Rust in beta) handle the negotiation. If you <strong>run</strong> an MCP server — for a
RAG pipeline, an internal tool bridge, anything past a demo — this changes
three things you own directly: how it scales (stateless means your ops story
gets simpler), how it's secured (OAuth 2.1 alignment means less of your own
auth code to get wrong), and your clock (twelve months to move off Roots,
Sampling, Logging, and SSE transport before they're gone). None of that is
optional just because you didn't ask for the rewrite.</p>
<hr>
<p>📖 <strong>Sources:</strong> <a href="https://blog.modelcontextprotocol.io/posts/2026-07-28/" target="_blank" rel="noopener noreferrer" class="">Model Context Protocol Blog — the 2026-07-28 specification</a> · <a href="https://workos.com/blog/mcp-2026-spec-agent-authentication" target="_blank" rel="noopener noreferrer" class="">WorkOS — authentication changes in the MCP 2026-07-28 spec</a> · <a href="https://www.theregister.com/devops/2026/07/23/model-context-protocol-prepares-to-break-with-its-stateful-past/5276722" target="_blank" rel="noopener noreferrer" class="">The Register — MCP breaks with its stateful past</a> · <a href="https://venturebeat.com/infrastructure/mcp-just-got-its-biggest-update-ever-heres-what-changes-for-ai-agents" target="_blank" rel="noopener noreferrer" class="">VentureBeat — MCP's biggest update, what changes for AI agents</a></p>]]></content>
        <author>
            <name>Rafael Fernandes</name>
            <uri>https://www.linkedin.com/in/rafaelmacariofernandes/</uri>
        </author>
        <category label="ai-news" term="ai-news"/>
        <category label="mcp" term="mcp"/>
        <category label="agents" term="agents"/>
        <category label="protocols" term="protocols"/>
        <category label="infrastructure" term="infrastructure"/>
    </entry>
    <entry>
        <title type="html"><![CDATA[GPT-5.6: OpenAI's new pitch is cheaper per task, not just smarter — verify it on your workload before you switch]]></title>
        <id>https://development-wec.wiline.com/docs/news/gpt-5-6-token-economics/</id>
        <link href="https://development-wec.wiline.com/docs/news/gpt-5-6-token-economics/"/>
        <updated>2026-07-13T00:00:00.000Z</updated>
        <summary type="html"><![CDATA[GPT-5.6 (Luna, Terra, Sol) ships with a claim aimed straight at your bill: frontier coding scores on half the output tokens. What that means for agent economics, why you should verify it on your own workload — and how to build so model choice stays a config change.]]></summary>
        <content type="html"><![CDATA[<div class="newsHero newsHero--bg newsHero--split"><div class="newsHero__bgSplit" aria-hidden="true"><div class="newsHero__panel newsHero__panel--a" style="background-image:url(/docs/img/news/gpt-5.6.png)"></div><div class="newsHero__panel newsHero__panel--b" style="background-image:url(/docs/img/news/chatgpt.avif)"></div><span class="newsHero__seam"></span><div class="newsHero__scrim"></div></div><span class="newsHero__eyebrow">Models · AI News</span><h2 class="newsHero__title">The frontier race just changed lanes: from smarter to cheaper per task</h2><div class="newsHero__transition"><span class="newsHero__pill newsHero__pill--from">Benchmark points</span><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2.5" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-arrow-right newsHero__arrow" aria-hidden="true"><path d="M5 12h14"></path><path d="m12 5 7 7-7 7"></path></svg><span class="newsHero__pill newsHero__pill--to">Token economics</span></div></div>
<p>OpenAI shipped <strong>GPT-5.6</strong> on July 9 — a family of three models (Luna, Terra, Sol) — and the
headline claim isn't a leaderboard score. It's an efficiency number: frontier coding performance
on <strong>less than half the output tokens</strong>. If you build agents, that's a claim about your bill,
not about bragging rights. It's also exactly the kind of claim you should measure yourself.</p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="what-shipped">What shipped<a href="https://development-wec.wiline.com/docs/news/gpt-5-6-token-economics/#what-shipped" class="hash-link" aria-label="Direct link to What shipped" title="Direct link to What shipped" translate="no">​</a></h2>
<p>Three tiers, cheapest to strongest, available in ChatGPT, Codex, and the OpenAI API:</p>
<table><thead><tr><th>Model</th><th>Positioning</th><th>Input / Output (per 1M tokens)</th></tr></thead><tbody><tr><td>Luna</td><td>fastest, budget tier</td><td>$1 / $6</td></tr><tr><td>Terra</td><td>mid tier — "performance competitive with GPT-5.5"</td><td>$2.50 / $15</td></tr><tr><td>Sol</td><td>flagship, "best coding model yet"</td><td>$5 / $30</td></tr></tbody></table>
<p>All three variants ship native tool use and multimodal reasoning, with long-context evals run
out to 1M tokens; Sol also powers the new ChatGPT Work professional tier. The claims worth
knowing (from <a href="https://openai.com/index/gpt-5-6/" target="_blank" rel="noopener noreferrer" class="">OpenAI's announcement</a> — directional, not
gospel): Sol scores <strong>80 on the Artificial Analysis Coding Agent Index</strong>, 2.8 points above
Anthropic's Fable 5 — <em>while using less than half the output tokens, in less than half the time,
at about a third of the estimated cost</em> — and <strong>88.8% on Terminal-Bench 2.1</strong>. On <em>Agents' Last
Exam</em>, an eval of long-running professional workflows across 55 fields, they report Terra and
Luna beating Fable 5 at around <strong>one-sixteenth the estimated cost</strong>. Sam Altman's framing: Sol
is "<strong>54% more token-efficient</strong>" on coding tasks. The release also leans hard on cybersecurity
— "frontier performance with significantly fewer tokens."</p>
<p>Two API features ship under the same efficiency banner, and they're the most developer-relevant
part of the launch: <strong>Programmatic Tool Calling</strong>, where the model writes and runs small
in-memory programs that coordinate tools and process intermediate results — instead of passing
every tool response back through the model, so tool-heavy tasks burn fewer tokens and fewer
round trips — and a <strong>multi-agent beta</strong> in the Responses API (the new Sol Ultra setting runs
four agents in parallel by default).</p>
<p>One more first, buried in the timeline: GPT-5.6 was planned for June and shipped three weeks
late because a <strong>US government review gated the release</strong> — Commerce's Center for AI Standards
and Innovation ran additional testing before OpenAI got permission for a public rollout. More on
why that matters below.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="why-token-efficiency-is-the-real-story">Why token efficiency is the real story<a href="https://development-wec.wiline.com/docs/news/gpt-5-6-token-economics/#why-token-efficiency-is-the-real-story" class="hash-link" aria-label="Direct link to Why token efficiency is the real story" title="Direct link to Why token efficiency is the real story" translate="no">​</a></h2>
<p>For chat, output tokens are a rounding error. For <strong>agents</strong>, they <em>are</em> the bill: an agentic
coding session burns tokens on every step — plans, diffs, retries, tool calls — and output
tokens cost 5–6× input tokens on every pricing card above. A model that solves the same task on
half the output tokens would be effectively <strong>half price and twice as fast at equal quality</strong>,
even if it's only a couple of benchmark points better. <em>If</em> it holds on your tasks — and that
"if" is the whole subject of the next section — that's a real and welcome shift in what vendors
compete on.</p>
<p>That's why "54% more token-efficient" is a more aggressive competitive move than any leaderboard
jump. And OpenAI wasn't alone — the whole frontier spent this week competing on price per task:
<strong>Meta launched Muse Spark 1.1</strong> for agentic coding at <strong>$1.25/$4.25</strong> per 1M tokens (with $20
free credits per account), and <strong>SpaceXAI released Grok 4.5</strong> — co-trained with Cursor — at
<strong>$2/$6</strong>, marketed as roughly 6× cheaper than comparable frontier models. Even OpenAI's
infrastructure news pointed the same direction: engineers reportedly <strong>halved inference costs
through software optimization alone</strong>. The race has visibly changed lanes, from benchmark points
to cost per task.</p>
<p>The three-tier ladder matters for the same reason. The emerging pattern is <strong>routing by task
difficulty</strong>: cheap tier for classification and extraction, mid tier for everyday generation,
flagship only for the hard multi-step work. If you run a gateway (we've
<a class="" href="https://development-wec.wiline.com/docs/news/litellm-rust-gateway/">written about why that layer matters</a>), this is what it's for.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="vendor-claims-are-eval-questions">Vendor claims are eval questions<a href="https://development-wec.wiline.com/docs/news/gpt-5-6-token-economics/#vendor-claims-are-eval-questions" class="hash-link" aria-label="Direct link to Vendor claims are eval questions" title="Direct link to Vendor claims are eval questions" translate="no">​</a></h2>
<p>Here's the thing about "54% more efficient" and "2.8 points above": those numbers come from the
vendor, measured on the vendor's chosen benchmark, with the vendor's harness. That's not an
accusation — the numbers are very likely real <em>on that benchmark</em>, and we apply the same
discount to everyone's, including the open-weight models we like (we said exactly this about
GLM-5.2's tables). It's a structural point: <strong>no vendor benchmark can know your workload.</strong></p>
<p>You don't even have to leave OpenAI's own coding table to see it — and credit to them for
publishing the mixed rows rather than only the flattering ones:</p>
<table><thead><tr><th>Coding eval (one vendor, one table)</th><th>GPT-5.6 Sol</th><th>Claude Fable 5</th><th>Who leads</th></tr></thead><tbody><tr><td>Artificial Analysis Coding Agent Index v1.1</td><td><strong>80</strong></td><td>77.2</td><td>Sol, +2.8</td></tr><tr><td>Terminal-Bench 2.1</td><td><strong>88.8%</strong></td><td>83.1%</td><td>Sol, +5.7</td></tr><tr><td>DeepSWE v1.1</td><td><strong>72.7%</strong></td><td>69.7%</td><td>Sol, +3.0</td></tr><tr><td>SWE-Bench Pro</td><td>64.6%</td><td><strong>80%</strong></td><td>Fable 5, +15.4</td></tr></tbody></table>
<p><em>(Numbers from <a href="https://openai.com/index/gpt-5-6/" target="_blank" rel="noopener noreferrer" class="">OpenAI's own announcement</a>.)</em></p>
<p>Four coding benchmarks, and "which is the better coding model" flips depending on the row — with
the single largest gap pointing the <em>other</em> way. That's not a knock on either model or on the
table; it's what benchmarks are. If one vendor's own page can't produce a single answer, a
launch-day headline certainly can't tell you what happens on <strong>your</strong> codebase. Token efficiency
compounds the problem: it varies wildly by task shape — a model that's terse on Python refactors
can be verbose on SQL or long-form answers, and an agent harness different from the vendor's can
erase (or amplify) the whole advantage.</p>
<p>The good news: this is a solved problem, and you already have the tooling if you followed our
evals series. The same Promptfoo setup from
<a class="" href="https://development-wec.wiline.com/docs/tutorials/eval-models-promptfoo-wiline-inference/">part 1</a> compares any two OpenAI-compatible
endpoints on <em>your</em> prompts — with <strong>cost and latency assertions</strong>, not just quality grades.
And the head-to-head worth running this week isn't GPT-5.6 against its own launch table — it's
GPT-5.6 against the strongest open-weight model you can serve yourself:</p>
<div class="language-yaml codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#393A34;--prism-background-color:#f6f8fa"><div class="codeBlockTitle_OeMC">the experiment worth an afternoon (part-1 skill)</div><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-yaml codeBlock_bY9V thin-scrollbar" style="color:#393A34;background-color:#f6f8fa"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#393A34"><span class="token key atrule" style="color:#00a4db">providers</span><span class="token punctuation" style="color:#393A34">:</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">  </span><span class="token punctuation" style="color:#393A34">-</span><span class="token plain"> openai</span><span class="token punctuation" style="color:#393A34">:</span><span class="token plain">chat</span><span class="token punctuation" style="color:#393A34">:</span><span class="token plain">&lt;new</span><span class="token punctuation" style="color:#393A34">-</span><span class="token plain">model</span><span class="token punctuation" style="color:#393A34">-</span><span class="token plain">id</span><span class="token punctuation" style="color:#393A34">&gt;</span><span class="token plain">          </span><span class="token comment" style="color:#999988;font-style:italic"># the challenger everyone's talking about</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">  </span><span class="token punctuation" style="color:#393A34">-</span><span class="token plain"> </span><span class="token key atrule" style="color:#00a4db">id</span><span class="token punctuation" style="color:#393A34">:</span><span class="token plain"> openai</span><span class="token punctuation" style="color:#393A34">:</span><span class="token plain">chat</span><span class="token punctuation" style="color:#393A34">:</span><span class="token plain">gemma4              </span><span class="token comment" style="color:#999988;font-style:italic"># the open-weight model already on WEC</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">    </span><span class="token key atrule" style="color:#00a4db">config</span><span class="token punctuation" style="color:#393A34">:</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">      </span><span class="token key atrule" style="color:#00a4db">apiBaseUrl</span><span class="token punctuation" style="color:#393A34">:</span><span class="token plain"> https</span><span class="token punctuation" style="color:#393A34">:</span><span class="token plain">//inference.wiline.com/v1</span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">      </span><span class="token key atrule" style="color:#00a4db">apiKeyEnvar</span><span class="token punctuation" style="color:#393A34">:</span><span class="token plain"> WEC_API_KEY</span><br></div><div class="token-line" style="color:#393A34"><span class="token plain"></span><span class="token key atrule" style="color:#00a4db">defaultTest</span><span class="token punctuation" style="color:#393A34">:</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">  </span><span class="token key atrule" style="color:#00a4db">assert</span><span class="token punctuation" style="color:#393A34">:</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">    </span><span class="token punctuation" style="color:#393A34">-</span><span class="token plain"> </span><span class="token key atrule" style="color:#00a4db">type</span><span class="token punctuation" style="color:#393A34">:</span><span class="token plain"> latency</span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">      </span><span class="token key atrule" style="color:#00a4db">threshold</span><span class="token punctuation" style="color:#393A34">:</span><span class="token plain"> </span><span class="token number" style="color:#36acaa">5000</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">    </span><span class="token punctuation" style="color:#393A34">-</span><span class="token plain"> </span><span class="token key atrule" style="color:#00a4db">type</span><span class="token punctuation" style="color:#393A34">:</span><span class="token plain"> cost</span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">      </span><span class="token key atrule" style="color:#00a4db">threshold</span><span class="token punctuation" style="color:#393A34">:</span><span class="token plain"> </span><span class="token number" style="color:#36acaa">0.002</span><br></div></code></pre></div></div>
<p>Same prompts, both models, and the eval reports quality, tokens, latency, and cost side by side.
An afternoon of this tells you what no launch post can: whether the efficiency claim survives
contact with <em>your</em> traffic. (And if the model backs a RAG service or agent, gate it like we
gated ours in <a class="" href="https://development-wec.wiline.com/docs/tutorials/rag-docs-assistant-wiline-inference/">the capstone</a> — judge for
semantics, deterministic regressions for known bugs.)</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-subtext-portability-just-got-more-valuable">The subtext: portability just got more valuable<a href="https://development-wec.wiline.com/docs/news/gpt-5-6-token-economics/#the-subtext-portability-just-got-more-valuable" class="hash-link" aria-label="Direct link to The subtext: portability just got more valuable" title="Direct link to The subtext: portability just got more valuable" translate="no">​</a></h2>
<p>Two things happened around this launch that matter more together than apart.</p>
<p>First, the <strong>government gate</strong>: for the first time, a frontier model's public release needed a
federal review to proceed — and the concern wasn't abstract. The same week, Sysdig documented
<strong>JadePuffer</strong>, the first fully autonomous AI-agent ransomware operation, which exploited a
Langflow CVE and then performed reconnaissance, lateral movement, and extortion on its own,
recovering from failed attempts within seconds. Whatever your politics, the engineering fact is
that access to closed frontier models now has one more valve you don't control — alongside
pricing, deprecations, and rate limits. We made this argument when <a class="" href="https://development-wec.wiline.com/docs/news/glm-5-2-open-weight-top-10/">an open-weight model cracked
the proprietary top 10</a>; this release made it for us.</p>
<p>Second, the <strong>open-weight world had a loud week too</strong>: Google released <strong>Gemma 4</strong>, an
open-weight, natively multimodal family from 2.3B to 31B parameters (dense and MoE variants,
vision and audio input, a thinking mode); Mistral opened early access on a new open-weight
Mixture-of-Experts family aimed squarely at the frontier gap; and Together AI closed an
<strong>$800M Series C</strong> at an $8.3B valuation on the back of open-model inference — citing over
$1B a year in bookings. The money is saying the same thing the GPT-5.6 delay is saying: models
you can download and run are a hedge worth paying for.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-takeaway--and-something-you-can-try-today">The takeaway — and something you can try today<a href="https://development-wec.wiline.com/docs/news/gpt-5-6-token-economics/#the-takeaway--and-something-you-can-try-today" class="hash-link" aria-label="Direct link to The takeaway — and something you can try today" title="Direct link to The takeaway — and something you can try today" translate="no">​</a></h2>
<p>The developer takeaway is the same one that's held all series: <strong>keep model choice a config
change</strong>. Code against the OpenAI-compatible API, put a gateway in front, and swapping models —
closed, open, or whatever lands on the <a class="" href="https://development-wec.wiline.com/docs/cloud_portal/platform/inference/models_hub/">WEC Models catalog</a>
next — is a base-URL and model-ID edit, not a rewrite. We don't offer GPT-5.6 on WEC today;
the point is that if you build portable and eval before you migrate, that fact constrains you
exactly zero.</p>
<p>Because here's the trap in every launch week: <em>new</em> starts feeling like a reason. It isn't —
it's a hypothesis. The question your eval should answer is whether the shiny closed model beats
an open-weight model you control — on your prompts, at your latency budget, per dollar — by
enough to be worth the valve someone else's hand is on. Sometimes it will, and then you switch
with evidence instead of hype. And sometimes the open model holds the line on the tasks you
actually run, and you just saved yourself a migration and a dependency in one afternoon.</p>
<p>You can run that experiment today, because
<a href="https://deepmind.google/models/gemma/gemma-4/" target="_blank" rel="noopener noreferrer" class=""><strong>Gemma 4</strong></a> — the open-weight, natively
multimodal family Google shipped this same week — <strong>is already live on WEC inference</strong>. And it's
a serious counterpart, not a consolation prize. What Gemma 4 brings to the table:</p>
<ul>
<li class=""><strong>Natively multimodal</strong> — image <em>and</em> audio understanding in an open model, so document
screenshots, scanned forms, and call recordings go through the same endpoint as your text
prompts.</li>
<li class=""><strong>Thinking variants</strong> for the harder multi-step chains — the same "spend tokens to reason"
trade GPT-5.6 is selling, on weights you control.</li>
<li class=""><strong>A size ladder of its own</strong> (edge-sized E2B/E4B up to 31B) — the same route-by-difficulty
pattern from above, without a per-tier vendor contract.</li>
<li class=""><strong>140+ languages</strong>, and vendor benchmarks that put the 31B in frontier company — Google claims
performance comparable to models 10–30× larger. Same discount applies as to OpenAI's numbers,
and the same tool settles it: put it in the eval.</li>
</ul>
<p>One model-ID swap and you're on it:</p>
<div class="language-bash codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#393A34;--prism-background-color:#f6f8fa"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-bash codeBlock_bY9V thin-scrollbar" style="color:#393A34;background-color:#f6f8fa"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#393A34"><span class="token function" style="color:#d73a49">curl</span><span class="token plain"> https://inference.wiline.com/v1/chat/completions </span><span class="token punctuation" style="color:#393A34">\</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">  </span><span class="token parameter variable" style="color:#36acaa">-H</span><span class="token plain"> </span><span class="token string" style="color:#e3116c">"Authorization: Bearer </span><span class="token string variable" style="color:#36acaa">$WEC_API_KEY</span><span class="token string" style="color:#e3116c">"</span><span class="token plain"> </span><span class="token punctuation" style="color:#393A34">\</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">  </span><span class="token parameter variable" style="color:#36acaa">-H</span><span class="token plain"> </span><span class="token string" style="color:#e3116c">"Content-Type: application/json"</span><span class="token plain"> </span><span class="token punctuation" style="color:#393A34">\</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">  </span><span class="token parameter variable" style="color:#36acaa">-d</span><span class="token plain"> </span><span class="token string" style="color:#e3116c">'{ "model": "gemma4", "messages": [{"role":"user","content":"Summarize this incident report…"}] }'</span><br></div></code></pre></div></div>
<p>GPT-5.6's token-efficiency push is good news for everyone — even if you never send OpenAI a
single request, it drags the whole market toward pricing honesty. Take the positive at face
value, take the numbers as hypotheses, and let your own eval — GPT-5.6 on one side, Gemma 4 on
WEC on the other — make the call.</p>
<hr>
<p>📖 <strong>Sources:</strong> <a href="https://openai.com/index/gpt-5-6/" target="_blank" rel="noopener noreferrer" class="">OpenAI — GPT-5.6 announcement</a> · <a href="https://deepmind.google/models/gemma/gemma-4/" target="_blank" rel="noopener noreferrer" class="">Google DeepMind — Gemma 4</a> · <a href="https://techcrunch.com/2026/07/09/openai-launches-its-new-family-of-models-with-gpt-5-6/" target="_blank" rel="noopener noreferrer" class="">TechCrunch — OpenAI launches GPT-5.6</a> · <a href="https://www.cnbc.com/2026/07/08/openai-expanding-gpt-5point6-ai-model-release-ending-government-limits.html" target="_blank" rel="noopener noreferrer" class="">CNBC — public release after government limits</a> · <a href="https://www.axios.com/2026/07/09/ai-openai-gpt-release" target="_blank" rel="noopener noreferrer" class="">Axios — GPT-5.6 and ChatGPT Work</a> · <a href="https://www.engadget.com/2210308/openai-rolls-out-gpt5-6-july-9/" target="_blank" rel="noopener noreferrer" class="">Engadget — rollout timing</a> · <a href="https://www.techtimes.com/articles/319798/20260706/mistral-ai-targets-frontier-gap-open-weight-model-entering-july-early-access.htm" target="_blank" rel="noopener noreferrer" class="">TechTimes — Mistral's open-weight MoE early access</a> · <a href="https://asanify.com/blog/news/open-weight-model-funding-july-7-2026/" target="_blank" rel="noopener noreferrer" class="">Asanify — Together AI's $800M round</a> · <a href="https://medium.com/nlplanet/gpt-5-6-is-out-weekly-ai-newsletter-july-13th-2026-4502e4c324a7" target="_blank" rel="noopener noreferrer" class="">NLPlanet — weekly AI newsletter, July 13</a> <em>(Muse Spark 1.1, Grok 4.5, Gemma 4, JadePuffer, inference-cost items)</em></p>]]></content>
        <author>
            <name>Rafael Fernandes</name>
            <uri>https://www.linkedin.com/in/rafaelmacariofernandes/</uri>
        </author>
        <category label="ai-news" term="ai-news"/>
        <category label="models" term="models"/>
        <category label="openai" term="openai"/>
        <category label="token-efficiency" term="token-efficiency"/>
        <category label="evals" term="evals"/>
    </entry>
    <entry>
        <title type="html"><![CDATA[GLM-5.2: the only open-weight model in the top 10 — and you can run it on WEC]]></title>
        <id>https://development-wec.wiline.com/docs/news/glm-5-2-open-weight-top-10/</id>
        <link href="https://development-wec.wiline.com/docs/news/glm-5-2-open-weight-top-10/"/>
        <updated>2026-06-24T00:00:00.000Z</updated>
        <summary type="html"><![CDATA[GLM-5.2 is the lone open-weight, MIT-licensed model holding its own against the proprietary frontier — a 1M-token context and top open-source coding scores. And it's available on WiLine Edge Cloud.]]></summary>
        <content type="html"><![CDATA[<div class="newsHero newsHero--bg" style="background-image:linear-gradient(rgba(2,12,31,0.62), rgba(2,12,31,0.80)), url(/docs/img/news/glm-cover-a.webp)"><span class="newsHero__eyebrow">Models · AI News</span><h2 class="newsHero__title">An open-weight model just cracked the proprietary top 10</h2><div class="newsHero__transition"><span class="newsHero__pill newsHero__pill--from">Closed frontier</span><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2.5" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-arrow-right newsHero__arrow" aria-hidden="true"><path d="M5 12h14"></path><path d="m12 5 7 7-7 7"></path></svg><span class="newsHero__pill newsHero__pill--to">Open weights</span></div></div>
<p>Look at almost any current model leaderboard and the top is a wall of Anthropic and
OpenAI. Then, sitting in the top 10, there's one outlier that isn't proprietary at
all: <strong>GLM-5.2</strong> from Z.ai — open weights, MIT-licensed. That's the story worth
paying attention to.</p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-standing">The standing<a href="https://development-wec.wiline.com/docs/news/glm-5-2-open-weight-top-10/#the-standing" class="hash-link" aria-label="Direct link to The standing" title="Direct link to The standing" translate="no">​</a></h2>
<p>On the <a href="https://arena.ai/leaderboard/agent" target="_blank" rel="noopener noreferrer" class="">Arena.ai agent leaderboard</a>, GLM-5.2
(Max) lands at <strong>#10</strong> — the <strong>only open-weight model in the top 10</strong>, surrounded
entirely by closed frontier models from Anthropic and OpenAI. (Leaderboards move;
this is a snapshot — <a href="https://arena.ai/leaderboard/agent" target="_blank" rel="noopener noreferrer" class="">check the live ranking</a>.)</p>
<p><span class="zoomImage__wrap"><img alt="GLM-5.2 (Max) on the Arena.ai agent leaderboard — the only open-weight model in the top 10" src="https://development-wec.wiline.com/docs/assets/images/glm-5-2-leaderboard-3f2c9fe86e103693aa80fdbcda6b054b.png" width="1400" height="837" class="zoomImage " loading="lazy"><span class="zoomImage__badge" aria-hidden="true"><svg viewBox="0 0 24 24" width="16" height="16" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round"><circle cx="11" cy="11" r="7"></circle><path d="M21 21l-4.3-4.3"></path><path d="M11 8v6M8 11h6"></path></svg></span></span></p>
<p>That's the headline: not that it tops the chart, but that an <strong>MIT-licensed model you
can download, self-host, and ship commercially</strong> is now trading blows with models you
can only rent.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="what-glm-52-actually-is">What GLM-5.2 actually is<a href="https://development-wec.wiline.com/docs/news/glm-5-2-open-weight-top-10/#what-glm-52-actually-is" class="hash-link" aria-label="Direct link to What GLM-5.2 actually is" title="Direct link to What GLM-5.2 actually is" translate="no">​</a></h2>
<ul>
<li class=""><strong>Open weights, MIT-licensed</strong> — no regional limits; download, self-host, fine-tune, and ship it commercially (<a href="https://huggingface.co/zai-org/GLM-5.2" target="_blank" rel="noopener noreferrer" class="">weights on Hugging Face</a>).</li>
<li class=""><strong>A solid 1M-token context</strong> (~750k words), built for long-horizon agent work. Its new <strong>IndexShare</strong> attention reuses one indexer across every four sparse layers — Z.ai reports <strong>~2.9× fewer per-token FLOPs at 1M context</strong>, which is what keeps that window affordable to run.</li>
<li class=""><strong>Two thinking-effort levels (High / Max)</strong> to trade latency for depth — <code>Max</code> for hard multi-step coding, <code>High</code> for lighter, faster work.</li>
<li class=""><strong>Anthropic/OpenAI-compatible API</strong> — drop it into Claude Code, OpenClaw, Cline, and others with a base-URL + model-ID swap; your harness and prompts stay put.</li>
</ul>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="how-it-compares">How it compares<a href="https://development-wec.wiline.com/docs/news/glm-5-2-open-weight-top-10/#how-it-compares" class="hash-link" aria-label="Direct link to How it compares" title="Direct link to How it compares" translate="no">​</a></h2>
<p>Z.ai's published benchmarks put GLM-5.2 shoulder-to-shoulder with the closed frontier on coding, and ahead on some reasoning:</p>
<table><thead><tr><th>Benchmark</th><th>GLM-5.2</th><th>Claude Opus 4.8</th><th>GPT-5.5</th></tr></thead><tbody><tr><td>SWE-bench Pro</td><td>62.1</td><td>69.2</td><td>58.6</td></tr><tr><td>Terminal-Bench 2.1 (best harness)</td><td>82.7</td><td>78.9</td><td>83.4</td></tr><tr><td>FrontierSWE (dominance)</td><td>74.4</td><td>75.1</td><td>72.6</td></tr><tr><td>AIME 2026</td><td>99.2</td><td>95.7</td><td>98.3</td></tr></tbody></table>
<p>It edges Opus 4.8 on Terminal-Bench, beats GPT-5.5 on FrontierSWE, tops both on AIME, and trails Opus on SWE-bench Pro — remarkably close for a model you can simply download. <em>(Numbers from <a href="https://docs.z.ai/guides/llm/glm-5.2" target="_blank" rel="noopener noreferrer" class="">Z.ai's GLM-5.2 benchmarks</a>; benchmarks are directional, not gospel.)</em></p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="why-it-matters">Why it matters<a href="https://development-wec.wiline.com/docs/news/glm-5-2-open-weight-top-10/#why-it-matters" class="hash-link" aria-label="Direct link to Why it matters" title="Direct link to Why it matters" translate="no">​</a></h2>
<p>The gap between open-weight and proprietary frontier models has been closing all
year. What's changed is the <strong>terms</strong>: with an MIT license and a clean API, GLM-5.2
is something you can <em>own and deploy</em>, not just call. When access to closed models
can shift with export controls or pricing overnight, an open-weight model that holds
top-10 quality is a foundation that stays put.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="run-it-on-wiline-edge-cloud">Run it on WiLine Edge Cloud<a href="https://development-wec.wiline.com/docs/news/glm-5-2-open-weight-top-10/#run-it-on-wiline-edge-cloud" class="hash-link" aria-label="Direct link to Run it on WiLine Edge Cloud" title="Direct link to Run it on WiLine Edge Cloud" translate="no">​</a></h2>
<p>You don't need a third-party account to try it — <strong>GLM-5.2 is available on WiLine
Edge Cloud through WEC Models</strong>, our OpenAI-compatible inference. Point any compatible
tool at the WEC inference endpoint and use GLM-5.2 as the model:</p>
<div class="language-bash codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#393A34;--prism-background-color:#f6f8fa"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-bash codeBlock_bY9V thin-scrollbar" style="color:#393A34;background-color:#f6f8fa"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#393A34"><span class="token function" style="color:#d73a49">curl</span><span class="token plain"> https://inference.wiline.com/v1/chat/completions </span><span class="token punctuation" style="color:#393A34">\</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">  </span><span class="token parameter variable" style="color:#36acaa">-H</span><span class="token plain"> </span><span class="token string" style="color:#e3116c">"Authorization: Bearer </span><span class="token string variable" style="color:#36acaa">$WEC_API_KEY</span><span class="token string" style="color:#e3116c">"</span><span class="token plain"> </span><span class="token punctuation" style="color:#393A34">\</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">  </span><span class="token parameter variable" style="color:#36acaa">-H</span><span class="token plain"> </span><span class="token string" style="color:#e3116c">"Content-Type: application/json"</span><span class="token plain"> </span><span class="token punctuation" style="color:#393A34">\</span><span class="token plain"></span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">  </span><span class="token parameter variable" style="color:#36acaa">-d</span><span class="token plain"> </span><span class="token string" style="color:#e3116c">'{ "model": "glm-5.2", "messages": [{"role":"user","content":"Refactor this for performance…"}] }'</span><br></div></code></pre></div></div>
<p>If you followed the <a class="" href="https://development-wec.wiline.com/docs/tutorials/">Self-hosting OpenClaw series</a>, this is the natural
next move: keep your agent, swap the model — point OpenClaw at GLM-5.2 on WEC instead
of a closed provider, and you're running a top-10 model you fully control.</p>
<hr>
<p>📖 <strong>Sources:</strong> <a href="https://arena.ai/leaderboard/agent" target="_blank" rel="noopener noreferrer" class="">Arena.ai agent leaderboard</a> · <a href="https://huggingface.co/zai-org/GLM-5.2" target="_blank" rel="noopener noreferrer" class="">GLM-5.2 on Hugging Face</a> · <a href="https://docs.z.ai/guides/llm/glm-5.2" target="_blank" rel="noopener noreferrer" class="">Z.ai model docs</a> · <a href="https://arxiv.org/abs/2602.15763" target="_blank" rel="noopener noreferrer" class="">GLM-5 technical report (arXiv)</a></p>]]></content>
        <author>
            <name>Rafael Fernandes</name>
            <uri>https://www.linkedin.com/in/rafaelmacariofernandes/</uri>
        </author>
        <category label="ai-news" term="ai-news"/>
        <category label="models" term="models"/>
        <category label="open-weight" term="open-weight"/>
        <category label="glm" term="glm"/>
        <category label="inference" term="inference"/>
    </entry>
    <entry>
        <title type="html"><![CDATA[Why LiteLLM Is Rewriting Its Gateway in Rust — and Why AI Developers Should Care]]></title>
        <id>https://development-wec.wiline.com/docs/news/litellm-rust-gateway/</id>
        <link href="https://development-wec.wiline.com/docs/news/litellm-rust-gateway/"/>
        <updated>2026-06-24T00:00:00.000Z</updated>
        <summary type="html"><![CDATA[LiteLLM is moving its AI gateway from Python to Rust. It's a signal that AI gateways are becoming critical infrastructure — with real consequences for latency, cost, and reliability on WiLine Edge Cloud.]]></summary>
        <content type="html"><![CDATA[<figure class="newsHero newsHero--image"><span class="newsHero__chip">Infrastructure · AI News</span><img src="https://development-wec.wiline.com/docs/img/news/litellm-rust.webp" alt="LiteLLM — migrating the AI gateway to Rust" loading="eager"></figure>
<p>The AI ecosystem is quietly going through the same transition web infrastructure went through years ago: the performance-critical pieces are moving off interpreted runtimes onto systems languages like Rust. <a href="https://docs.litellm.ai/blog/litellm-rust-launch" target="_blank" rel="noopener noreferrer" class="">LiteLLM rewriting its AI gateway in Rust</a> is the clearest evidence yet — and a sign the AI stack is maturing from experiment into production infrastructure.</p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-gateway-is-becoming-critical-infrastructure">The gateway is becoming critical infrastructure<a href="https://development-wec.wiline.com/docs/news/litellm-rust-gateway/#the-gateway-is-becoming-critical-infrastructure" class="hash-link" aria-label="Direct link to The gateway is becoming critical infrastructure" title="Direct link to The gateway is becoming critical infrastructure" translate="no">​</a></h2>
<p>LiteLLM is the open-source proxy a lot of teams put in front of their models to get one OpenAI-compatible endpoint across 100+ providers. If you've run one in production, this line from the announcement will feel familiar:</p>
<blockquote>
<p>Under real load, CPU and memory climb with concurrency, and pods get OOM-killed at the worst time.</p>
</blockquote>
<p>That's the quiet tax of a gateway: it sits on the hot path of <em>every</em> request — every completion, embedding, moderation call, and agent action flows through it — so its own overhead and memory footprint multiply across pods and regions. For years AI conversations were about model quality. As teams ship agents, RAG, and multi-model routing to production, the layer <em>in front</em> of the model is turning into a first-class infrastructure concern. Moving it to Rust is what that realization looks like in code.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-numbers">The numbers<a href="https://development-wec.wiline.com/docs/news/litellm-rust-gateway/#the-numbers" class="hash-link" aria-label="Direct link to The numbers" title="Direct link to The numbers" translate="no">​</a></h2>
<p>From LiteLLM's published benchmarks (reproducible — the harness ships with the post):</p>
<div class="metricCompare"><div class="metricCard"><span class="metricCard__label">Per-request overhead</span><span class="metricCard__factor">~150× lower</span><div class="metricCard__rows"><div class="metricCard__row metricCard__row--a"><span class="metricCard__name">LiteLLM (Python)</span><span class="metricCard__val">~7.5 ms</span></div><div class="metricCard__row metricCard__row--b"><span class="metricCard__name">Rust gateway</span><span class="metricCard__val">~0.05 ms</span></div></div></div><div class="metricCard"><span class="metricCard__label">Throughput under load</span><span class="metricCard__factor">~15× higher</span><div class="metricCard__rows"><div class="metricCard__row metricCard__row--a"><span class="metricCard__name">LiteLLM (Python)</span><span class="metricCard__val">453 req/s</span></div><div class="metricCard__row metricCard__row--b"><span class="metricCard__name">Rust gateway</span><span class="metricCard__val">6,782 req/s</span></div></div></div><div class="metricCard"><span class="metricCard__label">Peak memory under load</span><span class="metricCard__factor">~11× lighter</span><div class="metricCard__rows"><div class="metricCard__row metricCard__row--a"><span class="metricCard__name">LiteLLM (Python)</span><span class="metricCard__val">358.9 MB</span></div><div class="metricCard__row metricCard__row--b"><span class="metricCard__name">Rust gateway</span><span class="metricCard__val">31.7 MB</span></div></div></div></div>
<p>This measures the gateway <em>forwarding path</em> (transform → forward → handle response), not a full production workload — but that's exactly the layer you don't want eating CPU and memory under concurrency.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-real-win-is-memory-not-latency">The real win is memory, not latency<a href="https://development-wec.wiline.com/docs/news/litellm-rust-gateway/#the-real-win-is-memory-not-latency" class="hash-link" aria-label="Direct link to The real win is memory, not latency" title="Direct link to The real win is memory, not latency" translate="no">​</a></h2>
<p>Most readers will fixate on <strong>150× lower overhead</strong>. But for anyone <em>operating</em> a gateway, the more consequential number is <strong>11× less memory</strong>: 359 MB → ~32 MB. Latency is a per-request improvement; memory is what drives your bill and your reliability.</p>
<p>A gateway that holds ~32 MB instead of ~359 MB changes the operational math across the board:</p>
<ul>
<li class=""><strong>Kubernetes sizing</strong> — smaller pods, higher density per node.</li>
<li class=""><strong>Cloud cost</strong> — that footprint multiplies across every pod, region, and replica you run.</li>
<li class=""><strong>Autoscaling</strong> — lower, more predictable memory means less scaling churn.</li>
<li class=""><strong>OOM crashes</strong> — the failure mode that takes you down at peak largely goes away.</li>
</ul>
<p>When a component sits on the hot path of every request, shaving an order of magnitude off its memory compounds at scale far more than the headline latency figure.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="a-low-risk-rollout">A low-risk rollout<a href="https://development-wec.wiline.com/docs/news/litellm-rust-gateway/#a-low-risk-rollout" class="hash-link" aria-label="Direct link to A low-risk rollout" title="Direct link to A low-risk rollout" translate="no">​</a></h2>
<p>This is <strong>not a v2 and not a rewrite you have to migrate to</strong>. Config files, database schema, client APIs, and provider coverage stay the same. They're moving it in careful stages — a pure-Rust core via PyO3 bindings first (data transformation, no I/O), then the full server on axum/hyper — each route shipped to production behind passing parity tests before the next one starts:</p>
<figure class="stageFlow"><div class="stageFlow__track"><div class="stageFlow__card" style="background:rgba(var(--primary-rgb), 0.050);border-color:rgba(var(--primary-rgb), 0.250)"><span class="stageFlow__stage">Stage 0 · Today</span><span class="stageFlow__title">Python proxy</span><span class="stageFlow__tag">0% Rust</span></div><svg xmlns="http://www.w3.org/2000/svg" width="22" height="22" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2.5" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-arrow-right stageFlow__arrow" aria-hidden="true"><path d="M5 12h14"></path><path d="m12 5 7 7-7 7"></path></svg><div class="stageFlow__card" style="background:rgba(var(--primary-rgb), 0.123);border-color:rgba(var(--primary-rgb), 0.383)"><span class="stageFlow__stage">Stage 1 · Core in Rust</span><span class="stageFlow__title">Python drives transforms via PyO3</span><span class="stageFlow__tag">transforms + router</span></div><svg xmlns="http://www.w3.org/2000/svg" width="22" height="22" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2.5" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-arrow-right stageFlow__arrow" aria-hidden="true"><path d="M5 12h14"></path><path d="m12 5 7 7-7 7"></path></svg><div class="stageFlow__card" style="background:rgba(var(--primary-rgb), 0.197);border-color:rgba(var(--primary-rgb), 0.517)"><span class="stageFlow__stage">Stage 2 · Thin shell</span><span class="stageFlow__title">FastAPI shell, hot path in Rust</span><span class="stageFlow__tag">~full forwarding path</span></div><svg xmlns="http://www.w3.org/2000/svg" width="22" height="22" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2.5" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-arrow-right stageFlow__arrow" aria-hidden="true"><path d="M5 12h14"></path><path d="m12 5 7 7-7 7"></path></svg><div class="stageFlow__card" style="background:rgba(var(--primary-rgb), 0.270);border-color:rgba(var(--primary-rgb), 0.650)"><span class="stageFlow__stage">Stage 3 · Pure Rust</span><span class="stageFlow__title">axum server, Python in a sidecar</span><span class="stageFlow__tag">100% Rust</span></div></div><figcaption class="stageFlow__caption">Four stages — each shipped to production behind passing parity tests before the next begins.</figcaption></figure>
<p>Beta signup is open now; the roadmap targets OCR routes by mid-August 2026, <code>/chat/completions</code> and <code>/messages</code> by September, and the full server by <strong>December 1, 2026</strong>.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="what-this-means-for-ai-builders">What this means for AI builders<a href="https://development-wec.wiline.com/docs/news/litellm-rust-gateway/#what-this-means-for-ai-builders" class="hash-link" aria-label="Direct link to What this means for AI builders" title="Direct link to What this means for AI builders" translate="no">​</a></h2>
<p>For developers building on <strong>WiLine Edge Cloud</strong>, the gateway sits directly between your applications and your models — so a leaner, faster gateway flows straight through to the apps you ship:</p>
<ul>
<li class=""><strong>Faster AI APIs.</strong> Less proxy overhead means faster responses where model latency is already low — embeddings, reranking, moderation, classification. On those workloads the gateway <em>was</em> the tax; now it nearly isn't.</li>
<li class=""><strong>Better reliability.</strong> Lower memory pressure reduces OOM kills, request failures, and autoscaling churn — the things that quietly erode an AI product's uptime in production.</li>
<li class=""><strong>More efficient multi-model deployments.</strong> If you route traffic across many providers, gateway cost stops scaling as aggressively with traffic — you serve more without your proxy fleet ballooning.</li>
<li class=""><strong>Stronger infrastructure foundations.</strong> As AI apps become production systems rather than experiments, the infra layers underneath them matter as much as model quality.</li>
</ul>
<p>If you followed the <a class="" href="https://development-wec.wiline.com/docs/tutorials/">OpenClaw series</a>, you already put a gateway-shaped thing on the critical path — a reverse proxy, a model router, an agent runtime. The lesson generalizes.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-bigger-lesson">The bigger lesson<a href="https://development-wec.wiline.com/docs/news/litellm-rust-gateway/#the-bigger-lesson" class="hash-link" aria-label="Direct link to The bigger lesson" title="Direct link to The bigger lesson" translate="no">​</a></h2>
<p>For years, most AI discussion focused on model quality. But as organizations deploy agents, retrieval systems, and multi-model workflows in production, <strong>the layer in front of your models is infrastructure</strong> — and every millisecond and megabyte on the hot path compounds at scale. LiteLLM's move to Rust reflects a broader industry realization: infrastructure efficiency is no longer a footnote to model performance, it's part of it.</p>
<p><strong>Worth watching, not yet worth switching:</strong> it's beta, and the Python proxy isn't going anywhere. But the direction of travel is clear.</p>
<hr>
<p>📖 <strong>Read the full announcement</strong> — the benchmarks, the route-by-route migration plan, and the architecture diagrams are all worth your time: <a href="https://docs.litellm.ai/blog/litellm-rust-launch" target="_blank" rel="noopener noreferrer" class="">LiteLLM — Building the fastest AI gateway in Rust</a>.</p>]]></content>
        <author>
            <name>Rafael Fernandes</name>
            <uri>https://www.linkedin.com/in/rafaelmacariofernandes/</uri>
        </author>
        <category label="ai-news" term="ai-news"/>
        <category label="gateways" term="gateways"/>
        <category label="performance" term="performance"/>
        <category label="self-hosting" term="self-hosting"/>
    </entry>
</feed>