{
    "version": "https://jsonfeed.org/version/1",
    "title": "WiLine Edge Cloud — AI News",
    "home_page_url": "https://development-wec.wiline.com/docs/news/",
    "description": "Short, high-signal takes on what is changing in AI infrastructure — and what it means for self-hosting on WiLine.",
    "items": [
        {
            "id": "https://development-wec.wiline.com/docs/news/router-classifier-tax-system-one-jev/",
            "content_html": "<div class=\"newsHero\"><div class=\"newsHero__glow\" aria-hidden=\"true\"></div><span class=\"newsHero__eyebrow\">Routing · AI News</span><h2 class=\"newsHero__title\">Jev decides, your LLMs answer</h2><div class=\"newsHero__transition\"><span class=\"newsHero__pill newsHero__pill--from\">A request comes in</span><svg xmlns=\"http://www.w3.org/2000/svg\" width=\"20\" height=\"20\" viewBox=\"0 0 24 24\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"2.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\" class=\"lucide lucide-arrow-right newsHero__arrow\" aria-hidden=\"true\"><path d=\"M5 12h14\"></path><path d=\"m12 5 7 7-7 7\"></path></svg><span class=\"newsHero__pill newsHero__pill--to\">The right model, in 127 ms</span></div></div>\n<p>TypeSafe <a href=\"https://typesafe.ai/blog/introducing-system-one-models-and-jev\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">shipped its first model on 15 September</a>,\nand the interesting thing about Jev is what it refuses to do. It doesn't write you a\nparagraph. You give it a request and it hands back one structured value — a label, a\nclass, a decision — in, they say, 70 to 500 ms. Founder Diogo Almeida's framing is\nthe clearest line in the post: <em>\"Think of Jev as a frontier-intelligence function\ncall: unstructured state in, typed probabilistic decisions out.\"</em></p>\n<p>Most models are built to talk to people. Jev is built to be called by code — and the\nfirst job that shape fits is routing.</p>\n<!-- -->\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"a-model-built-to-decide-not-to-talk\">A model built to decide, not to talk<a href=\"https://development-wec.wiline.com/docs/news/router-classifier-tax-system-one-jev/#a-model-built-to-decide-not-to-talk\" class=\"hash-link\" aria-label=\"Direct link to A model built to decide, not to talk\" title=\"Direct link to A model built to decide, not to talk\" translate=\"no\">​</a></h2>\n<p>The name tells the story twice. \"System One\" is Kahneman's term for fast, automatic\njudgement — the snap decision — against \"System Two,\" the slow deliberate reasoning\nwe reach for a big LLM to do. And \"Jev\" is for <strong>William Stanley Jevons</strong>, the\neconomist of the efficiency paradox: make something cheaper and people use far more\nof it. TypeSafe is betting cheap machine decisions get used everywhere.</p>\n<p>A chat model generates tokens one at a time until a paragraph exists. Jev does the\nopposite — TypeSafe says it \"gives up string generation,\" samples all outputs in a\nsingle parallel query, and returns one value from a set you define in advance. Two\nproperties fall out, and both matter for routing:</p>\n<ul>\n<li class=\"\"><strong>The output is typed and structured</strong>, from a known set (<code>SIMPLE</code>, <code>MEDIUM</code>,\n<code>COMPLEX</code>, <code>REASONING</code>), not prose you have to parse. TypeSafe claims it \"never\nmakes type errors\" — worth being precise here, because they are: that 0% is\n<em>\"not empirical. Schema matching is guaranteed, thus we can confidently add 0%.\"</em>\nIt's a property of constraining the output to a schema, not a measured result.</li>\n<li class=\"\"><strong>Each decision carries a calibrated probability</strong> — \"higher confidence means\nhigher accuracy,\" trained via a method they call Reinforcement Learning for\nCalibrated Decisions (RLCD). If that holds, code can act on the number: take the\nlabel when confident, escalate when not.</li>\n</ul>\n<p>Their headline numbers (\"40x-200x faster,\" output \"too cheap to meter,\" a workflow\neval at \"193.6x faster, 444.6x cheaper\") are vendor claims on their own evals, and\nto their credit they flag the bias — the evals were \"made by individuals on our\nmodel capabilities team.\" Treat them as claims until someone independent re-runs\nthem.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"how-routing-works-and-why-the-decider-is-the-problem\">How routing works, and why the decider is the problem<a href=\"https://development-wec.wiline.com/docs/news/router-classifier-tax-system-one-jev/#how-routing-works-and-why-the-decider-is-the-problem\" class=\"hash-link\" aria-label=\"Direct link to How routing works, and why the decider is the problem\" title=\"Direct link to How routing works, and why the decider is the problem\" translate=\"no\">​</a></h2>\n<p>Model routing is a cost play. Not every request needs your biggest model, so you put\na cheap model and an expensive one behind one endpoint and something in front decides\nwhich answers. Easy turns take the cheap tier; hard ones take the expensive tier; the\nbill drops. That \"something in front\" is a classifier:</p>\n<!-- -->\n<p>Build that classifier the usual way — ask <em>another</em> LLM \"is this simple or complex?\"\n— and every request pays for an LLM call <strong>before</strong> it pays for the answer. That call\nis slow and costs tokens, on every request, whether it landed on the cheap tier or\nnot. The decision meant to save money is quietly spending it.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"jev-in-the-decider-slot\">Jev in the decider slot<a href=\"https://development-wec.wiline.com/docs/news/router-classifier-tax-system-one-jev/#jev-in-the-decider-slot\" class=\"hash-link\" aria-label=\"Direct link to Jev in the decider slot\" title=\"Direct link to Jev in the decider slot\" translate=\"no\">​</a></h2>\n<p>Jev fits that slot. You don't ask it to answer — you ask which tier the request\nbelongs to, and the gateway routes on the label:</p>\n<figure class=\"stageFlow\"><div class=\"stageFlow__track\"><div class=\"stageFlow__card\" style=\"background:rgba(var(--primary-rgb), 0.050);border-color:rgba(var(--primary-rgb), 0.250)\"><span class=\"stageFlow__stage\">1</span><span class=\"stageFlow__title\">Request arrives</span><span class=\"stageFlow__tag\">at the gateway</span></div><svg xmlns=\"http://www.w3.org/2000/svg\" width=\"22\" height=\"22\" viewBox=\"0 0 24 24\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"2.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\" class=\"lucide lucide-arrow-right stageFlow__arrow\" aria-hidden=\"true\"><path d=\"M5 12h14\"></path><path d=\"m12 5 7 7-7 7\"></path></svg><div class=\"stageFlow__card\" style=\"background:rgba(var(--primary-rgb), 0.123);border-color:rgba(var(--primary-rgb), 0.383)\"><span class=\"stageFlow__stage\">2</span><span class=\"stageFlow__title\">Jev labels it</span><span class=\"stageFlow__tag\">~127 ms, one tier</span></div><svg xmlns=\"http://www.w3.org/2000/svg\" width=\"22\" height=\"22\" viewBox=\"0 0 24 24\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"2.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\" class=\"lucide lucide-arrow-right stageFlow__arrow\" aria-hidden=\"true\"><path d=\"M5 12h14\"></path><path d=\"m12 5 7 7-7 7\"></path></svg><div class=\"stageFlow__card\" style=\"background:rgba(var(--primary-rgb), 0.197);border-color:rgba(var(--primary-rgb), 0.517)\"><span class=\"stageFlow__stage\">3</span><span class=\"stageFlow__title\">Gateway routes</span><span class=\"stageFlow__tag\">to the matching model</span></div><svg xmlns=\"http://www.w3.org/2000/svg\" width=\"22\" height=\"22\" viewBox=\"0 0 24 24\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"2.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\" class=\"lucide lucide-arrow-right stageFlow__arrow\" aria-hidden=\"true\"><path d=\"M5 12h14\"></path><path d=\"m12 5 7 7-7 7\"></path></svg><div class=\"stageFlow__card\" style=\"background:rgba(var(--primary-rgb), 0.270);border-color:rgba(var(--primary-rgb), 0.650)\"><span class=\"stageFlow__stage\">4</span><span class=\"stageFlow__title\">Model answers</span><span class=\"stageFlow__tag\">small or large</span></div></div><figcaption class=\"stageFlow__caption\">The classifier is step 2. Make it cheap and the arrangement pays off; make it an LLM and it taxes every call.</figcaption></figure>\n<p>In LiteLLM's Auto Router that is a config change, not new code — you name the tiers,\nmap each to a model, and set <code>classifier_type: jev</code>:</p>\n<div class=\"language-yaml codeBlockContainer_Ckt0 theme-code-block\" style=\"--prism-color:#393A34;--prism-background-color:#f6f8fa\"><div class=\"codeBlockContent_QJqH\"><pre tabindex=\"0\" class=\"prism-code language-yaml codeBlock_bY9V thin-scrollbar\" style=\"color:#393A34;background-color:#f6f8fa\"><code class=\"codeBlockLines_e6Vv\"><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token punctuation\" style=\"color:#393A34\">-</span><span class=\"token plain\"> </span><span class=\"token key atrule\" style=\"color:#00a4db\">model_name</span><span class=\"token punctuation\" style=\"color:#393A34\">:</span><span class=\"token plain\"> jev</span><span class=\"token punctuation\" style=\"color:#393A34\">-</span><span class=\"token plain\">router</span><br></div><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\">  </span><span class=\"token key atrule\" style=\"color:#00a4db\">litellm_params</span><span class=\"token punctuation\" style=\"color:#393A34\">:</span><span class=\"token plain\"></span><br></div><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\">    </span><span class=\"token key atrule\" style=\"color:#00a4db\">model</span><span class=\"token punctuation\" style=\"color:#393A34\">:</span><span class=\"token plain\"> auto_router/complexity_router</span><br></div><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\">    </span><span class=\"token key atrule\" style=\"color:#00a4db\">complexity_router_config</span><span class=\"token punctuation\" style=\"color:#393A34\">:</span><span class=\"token plain\"></span><br></div><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\">      </span><span class=\"token key atrule\" style=\"color:#00a4db\">tiers</span><span class=\"token punctuation\" style=\"color:#393A34\">:</span><span class=\"token plain\"></span><br></div><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\">        </span><span class=\"token key atrule\" style=\"color:#00a4db\">SIMPLE</span><span class=\"token punctuation\" style=\"color:#393A34\">:</span><span class=\"token plain\"> </span><span class=\"token punctuation\" style=\"color:#393A34\">{</span><span class=\"token punctuation\" style=\"color:#393A34\">{</span><span class=\"token plain\">openai_small</span><span class=\"token punctuation\" style=\"color:#393A34\">}</span><span class=\"token punctuation\" style=\"color:#393A34\">}</span><span class=\"token plain\"></span><br></div><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\">        </span><span class=\"token key atrule\" style=\"color:#00a4db\">MEDIUM</span><span class=\"token punctuation\" style=\"color:#393A34\">:</span><span class=\"token plain\"> </span><span class=\"token punctuation\" style=\"color:#393A34\">{</span><span class=\"token punctuation\" style=\"color:#393A34\">{</span><span class=\"token plain\">openai_large</span><span class=\"token punctuation\" style=\"color:#393A34\">}</span><span class=\"token punctuation\" style=\"color:#393A34\">}</span><span class=\"token plain\"></span><br></div><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\">        </span><span class=\"token key atrule\" style=\"color:#00a4db\">COMPLEX</span><span class=\"token punctuation\" style=\"color:#393A34\">:</span><span class=\"token plain\"> </span><span class=\"token punctuation\" style=\"color:#393A34\">{</span><span class=\"token punctuation\" style=\"color:#393A34\">{</span><span class=\"token plain\">anthropic</span><span class=\"token punctuation\" style=\"color:#393A34\">}</span><span class=\"token punctuation\" style=\"color:#393A34\">}</span><span class=\"token plain\"></span><br></div><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\">        </span><span class=\"token key atrule\" style=\"color:#00a4db\">REASONING</span><span class=\"token punctuation\" style=\"color:#393A34\">:</span><span class=\"token plain\"> </span><span class=\"token punctuation\" style=\"color:#393A34\">{</span><span class=\"token punctuation\" style=\"color:#393A34\">{</span><span class=\"token plain\">anthropic_large</span><span class=\"token punctuation\" style=\"color:#393A34\">}</span><span class=\"token punctuation\" style=\"color:#393A34\">}</span><span class=\"token plain\"></span><br></div><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\">      </span><span class=\"token key atrule\" style=\"color:#00a4db\">classifier_type</span><span class=\"token punctuation\" style=\"color:#393A34\">:</span><span class=\"token plain\"> jev</span><br></div></code></pre></div></div>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"what-the-swap-is-worth--and-where-the-llm-falls-down\">What the swap is worth — and where the LLM falls down<a href=\"https://development-wec.wiline.com/docs/news/router-classifier-tax-system-one-jev/#what-the-swap-is-worth--and-where-the-llm-falls-down\" class=\"hash-link\" aria-label=\"Direct link to What the swap is worth — and where the LLM falls down\" title=\"Direct link to What the swap is worth — and where the LLM falls down\" translate=\"no\">​</a></h2>\n<p>On 20 September LiteLLM's Moe Khalil <a href=\"https://docs.litellm.ai/blog/jev-auto-router-benchmark\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">published a benchmark</a>\ntitled <em>\"JEV Classifier: 5.43x as Fast as Haiku, 96% Lower Cost.\"</em> He ran 80 cases\nthree times each — 240 classification calls — with Jev (<code>jev-1.13.0</code>) against Claude\nHaiku 4.5 as the \"classify with an LLM\" baseline:</p>\n<div class=\"metricCompare\"><div class=\"metricCard\"><span class=\"metricCard__label\">Classifier latency (p50)</span><span class=\"metricCard__factor\">5.43× faster</span><div class=\"metricCard__rows\"><div class=\"metricCard__row metricCard__row--a\"><span class=\"metricCard__name\">Haiku 4.5 classifier</span><span class=\"metricCard__val\">688.40 ms</span></div><div class=\"metricCard__row metricCard__row--b\"><span class=\"metricCard__name\">Jev classifier</span><span class=\"metricCard__val\">126.81 ms</span></div></div></div><div class=\"metricCard\"><span class=\"metricCard__label\">Matched the expected tier</span><span class=\"metricCard__factor\">+21 pts</span><div class=\"metricCard__rows\"><div class=\"metricCard__row metricCard__row--a\"><span class=\"metricCard__name\">Haiku 4.5 classifier</span><span class=\"metricCard__val\">73.75% (177/240)</span></div><div class=\"metricCard__row metricCard__row--b\"><span class=\"metricCard__name\">Jev classifier</span><span class=\"metricCard__val\">95.00% (228/240)</span></div></div></div><div class=\"metricCard\"><span class=\"metricCard__label\">Cost for 240 calls</span><span class=\"metricCard__factor\">~96% cheaper</span><div class=\"metricCard__rows\"><div class=\"metricCard__row metricCard__row--a\"><span class=\"metricCard__name\">Haiku 4.5 classifier</span><span class=\"metricCard__val\">$0.1985</span></div><div class=\"metricCard__row metricCard__row--b\"><span class=\"metricCard__name\">Jev classifier</span><span class=\"metricCard__val\">$0.0077</span></div></div></div></div>\n<p>The average hides the real finding, which is in the per-tier numbers. On <code>SIMPLE</code>\nboth were perfect (100%). But on <code>MEDIUM</code> Haiku collapsed to <strong>38.33%</strong> where Jev\nheld <strong>85%</strong>, and on <code>REASONING</code> Haiku managed <strong>76.67%</strong> against Jev's <strong>100%</strong>. In\nother words, the LLM classifier is fine at telling trivial from non-trivial, and\nunreliable exactly in the middle band where routing decisions actually save or cost\nyou money.</p>\n<p>Two caveats, both LiteLLM's own, stated plainly: the expected tiers \"were authored\nwith the synthetic prompts, without independent review,\" and the benchmark measured\n<strong>the routing decision, not the quality of the final answer</strong> — \"downstream answer\nquality was not measured.\" To their credit they published a frozen, hash-verified\n<a href=\"https://docs.litellm.ai/blog/jev-auto-router-benchmark\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">reproduction archive</a> so the\nrun can be re-analysed. Read the result as \"Jev picked the tier fast, cheap, and the\nway they expected,\" not \"your answers get 96% cheaper end to end.\"</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"where-this-fits-on-wec\">Where this fits on WEC<a href=\"https://development-wec.wiline.com/docs/news/router-classifier-tax-system-one-jev/#where-this-fits-on-wec\" class=\"hash-link\" aria-label=\"Direct link to Where this fits on WEC\" title=\"Direct link to Where this fits on WEC\" translate=\"no\">​</a></h2>\n<p>We build this exact setup in the tutorial on\n<a class=\"\" href=\"https://development-wec.wiline.com/docs/tutorials/litellm-complexity-routing/\">routing by complexity</a> — a gateway in front\nof the WEC Inference API that scores each request and sends it to the right tier,\nQwen3.5:9B for the simple turns and Qwen3.5:122B for the hard ones. And we have shown\nhow a classifier gets it <a class=\"\" href=\"https://development-wec.wiline.com/docs/news/llm-router-cannot-classify-yes/\">wrong on a bare \"yes\"</a>,\nsending work to the wrong model.</p>\n<p>In both, the classifier is the part nobody prices. Put an LLM in that slot and you\nadd two-thirds of a second and a token charge to every call before the real model is\neven chosen. A model like Jev is a bet that the decision in front of your models can\nbe near-free — so routing pays off instead of taxing itself.</p>\n<p>It is early, it is hosted, and the numbers still need someone independent to re-run\nthem. For a WEC stack that keeps inference in-house, \"hosted\" is the real question\nmark — you would be sending every request's routing decision to a third party. But\nthe shape is the part that travels: a model that decides in milliseconds, sitting in\nfront of the models that answer.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"sources\">Sources<a href=\"https://development-wec.wiline.com/docs/news/router-classifier-tax-system-one-jev/#sources\" class=\"hash-link\" aria-label=\"Direct link to Sources\" title=\"Direct link to Sources\" translate=\"no\">​</a></h2>\n<ul>\n<li class=\"\">Diogo Almeida, TypeSafe — <a href=\"https://typesafe.ai/blog/introducing-system-one-models-and-jev\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\"><em>Introducing System One Models and Jev</em></a> (15 September 2026)</li>\n<li class=\"\">Moe Khalil, LiteLLM — <a href=\"https://docs.litellm.ai/blog/jev-auto-router-benchmark\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\"><em>JEV Classifier: 5.43x as Fast as Haiku, 96% Lower Cost</em></a> (20 September 2026)</li>\n</ul>",
            "url": "https://development-wec.wiline.com/docs/news/router-classifier-tax-system-one-jev/",
            "title": "Jev: a decision model you put in front of your LLMs to route traffic",
            "summary": "Jev is TypeSafe's first model — it doesn't chat, it decides. Give it a request and it returns one structured label in about 127 ms. LiteLLM wired it into its Auto Router as the classifier and clocked it 5.43x faster and 96% cheaper than a Haiku doing the same job. Here is what Jev is, how routing works, where the LLM classifier actually falls down, and where this fits in front of your models on WEC.",
            "date_modified": "2026-09-23T00:00:00.000Z",
            "author": {
                "name": "Rafael Fernandes",
                "url": "https://www.linkedin.com/in/rafaelmacariofernandes/"
            },
            "tags": [
                "ai-news",
                "routing",
                "cost",
                "litellm",
                "models"
            ]
        },
        {
            "id": "https://development-wec.wiline.com/docs/news/pfizer-gateway-silent-throughput-bug/",
            "content_html": "<figure class=\"newsHero newsHero--image\"><span class=\"newsHero__chip\">Gateways · AI News</span><img src=\"https://development-wec.wiline.com/docs/img/news/litellm-cover.webp\" alt=\"A gateway losing half its throughput with no errors reported\" loading=\"eager\"></figure>\n<p>Pfizer runs an AI gateway — the single service every internal tool talks to when it\nwants a model. During a routine upgrade check, its throughput fell by half. No\nerrors. No failed requests. Nothing wrong on any dashboard.</p>\n<p>They published <a href=\"https://docs.litellm.ai/blog/pfizer-gateway-performance-and-resiliency\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">what they found</a>\ntogether with the LiteLLM team, and the cause is small enough to fit in a paragraph:\nsomeone wrote a setting that said <em>don't encrypt this connection</em>, and the system\nencrypted it anyway.</p>\n<!-- -->\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"what-actually-went-wrong\">What actually went wrong<a href=\"https://development-wec.wiline.com/docs/news/pfizer-gateway-silent-throughput-bug/#what-actually-went-wrong\" class=\"hash-link\" aria-label=\"Direct link to What actually went wrong\" title=\"Direct link to What actually went wrong\" translate=\"no\">​</a></h2>\n<p>The gateway keeps a cache — a fast side-store called Redis that it checks before\ndoing expensive work. The connection to that cache can be encrypted or not, and\nthat's a setting you write in a config file.</p>\n<p>The operator wrote the setting to <strong>off</strong>. Plain connection, no encryption.</p>\n<p>The code that read that setting checked whether the setting <em>existed</em>, rather than\nwhat it was set to. Writing <code>off</code> creates the setting. The setting now exists. So\nthe check passed, and the gateway opened an encrypted connection — to a cache that\nwasn't expecting one.</p>\n<p>Here's the part that makes this dangerous rather than merely wrong. When a program\ntries to start an encrypted conversation with something that doesn't speak\nencryption, it doesn't get turned away. It says hello and waits for a reply that\nnever comes. Not refused — ignored.</p>\n<p>So the cache wasn't <em>down</em>. The cache was <strong>slow</strong>. And slow is much harder to see.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"the-four-lines\">The four lines<a href=\"https://development-wec.wiline.com/docs/news/pfizer-gateway-silent-throughput-bug/#the-four-lines\" class=\"hash-link\" aria-label=\"Direct link to The four lines\" title=\"Direct link to The four lines\" translate=\"no\">​</a></h2>\n<p>For anyone who wants to look at it directly, this is\n<a href=\"https://github.com/BerriAI/litellm/blob/v1.89.2/litellm/_redis.py#L701\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\"><code>get_redis_connection_pool</code></a>\nas it stood in the affected version:</p>\n<div class=\"language-python codeBlockContainer_Ckt0 theme-code-block\" style=\"--prism-color:#393A34;--prism-background-color:#f6f8fa\"><div class=\"codeBlockContent_QJqH\"><pre tabindex=\"0\" class=\"prism-code language-python codeBlock_bY9V thin-scrollbar\" style=\"color:#393A34;background-color:#f6f8fa\"><code class=\"codeBlockLines_e6Vv\"><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\">connection_class </span><span class=\"token operator\" style=\"color:#393A34\">=</span><span class=\"token plain\"> async_redis</span><span class=\"token punctuation\" style=\"color:#393A34\">.</span><span class=\"token plain\">Connection</span><br></div><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\"></span><span class=\"token keyword\" style=\"color:#00009f\">if</span><span class=\"token plain\"> </span><span class=\"token string\" style=\"color:#e3116c\">\"ssl\"</span><span class=\"token plain\"> </span><span class=\"token keyword\" style=\"color:#00009f\">in</span><span class=\"token plain\"> redis_kwargs</span><span class=\"token punctuation\" style=\"color:#393A34\">:</span><span class=\"token plain\"></span><br></div><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\">    connection_class </span><span class=\"token operator\" style=\"color:#393A34\">=</span><span class=\"token plain\"> async_redis</span><span class=\"token punctuation\" style=\"color:#393A34\">.</span><span class=\"token plain\">SSLConnection</span><br></div><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\">    redis_kwargs</span><span class=\"token punctuation\" style=\"color:#393A34\">.</span><span class=\"token plain\">pop</span><span class=\"token punctuation\" style=\"color:#393A34\">(</span><span class=\"token string\" style=\"color:#e3116c\">\"ssl\"</span><span class=\"token punctuation\" style=\"color:#393A34\">,</span><span class=\"token plain\"> </span><span class=\"token boolean\" style=\"color:#36acaa\">None</span><span class=\"token punctuation\" style=\"color:#393A34\">)</span><span class=\"token plain\"></span><br></div><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\">    redis_kwargs</span><span class=\"token punctuation\" style=\"color:#393A34\">[</span><span class=\"token string\" style=\"color:#e3116c\">\"connection_class\"</span><span class=\"token punctuation\" style=\"color:#393A34\">]</span><span class=\"token plain\"> </span><span class=\"token operator\" style=\"color:#393A34\">=</span><span class=\"token plain\"> connection_class</span><br></div></code></pre></div></div>\n<p><code>\"ssl\" in redis_kwargs</code> asks <em>does this key exist</em>. It does — you created it when you\nset it to false. The branch fires, and you get the encrypted connection you\nexplicitly turned off.</p>\n<p>Leave the setting out entirely and everything works. Write it down and say no, and it\nbreaks. The careful configuration is the broken one.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"why-nobody-got-an-error\">Why nobody got an error<a href=\"https://development-wec.wiline.com/docs/news/pfizer-gateway-silent-throughput-bug/#why-nobody-got-an-error\" class=\"hash-link\" aria-label=\"Direct link to Why nobody got an error\" title=\"Direct link to Why nobody got an error\" translate=\"no\">​</a></h2>\n<p>Every layer above the cache is built to tolerate a cache that isn't answering, and\nthat's correct design — if the cache is unavailable, you do the work the slow way and\ncarry on. Which is exactly what turns this into silence:</p>\n<blockquote>\n<p>Redis timeout errors appeared in the proxy's internal logs, but the gateway still\nreturned HTTP 200 to every caller</p>\n</blockquote>\n<p>HTTP 200 means <em>success</em>. Every single request reported success, for the entire\nduration. The post is precise about where the damage landed:</p>\n<blockquote>\n<p>Redis operations were stalling on TLS handshake timeouts in the async hot path,\ndegrading throughput without surfacing HTTP-level errors</p>\n</blockquote>\n<p>The reported numbers: throughput down from about 300 requests per second to about\n156 — roughly half — with median response time under load at 4,200 ms, measured\nagainst 750 simulated concurrent users over a sustained 60-second run.</p>\n<p>Those measurements come from Pfizer's internal testing and are reported in the post\nrather than independently reproducible. The code is a different matter — you can open\nthe file at that version and read it yourself.</p>\n<p>Their own one-line summary is the best sentence in the writeup:</p>\n<blockquote>\n<p>A throughput drop with a 0% HTTP error rate is the worst kind of regression to\ncatch after the fact</p>\n</blockquote>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"the-bug-was-already-there\">The bug was already there<a href=\"https://development-wec.wiline.com/docs/news/pfizer-gateway-silent-throughput-bug/#the-bug-was-already-there\" class=\"hash-link\" aria-label=\"Direct link to The bug was already there\" title=\"Direct link to The bug was already there\" translate=\"no\">​</a></h2>\n<p>I wanted to know whether this arrived with the upgrade that exposed it. It didn't. I\ncompared the relevant file between the version that performed fine and the version\nthat didn't — they're byte-for-byte identical. The post says the same thing:</p>\n<blockquote>\n<p>The underlying Redis bug existed in older versions too, but surfaced during\nPfizer's v1.89.2 upgrade validation under this workload</p>\n</blockquote>\n<p>That's the detail worth sitting with. This wasn't a bad release you could roll back.\nThe flaw had been sitting in the connection code across versions that benchmarked\nperfectly well. It needed a particular configuration and a particular load before it\ndid any visible damage — and then it cost half a production gateway's capacity\nwithout raising anything.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"how-they-actually-found-it\">How they actually found it<a href=\"https://development-wec.wiline.com/docs/news/pfizer-gateway-silent-throughput-bug/#how-they-actually-found-it\" class=\"hash-link\" aria-label=\"Direct link to How they actually found it\" title=\"Direct link to How they actually found it\" translate=\"no\">​</a></h2>\n<p>Not through monitoring, and not by hunting through recent code changes. By turning\nfeatures off and on until the number moved:</p>\n<blockquote>\n<p>Enabled Redis caching. Median latency jumped to 4,200ms. That isolated it to the\nRedis connection path.</p>\n</blockquote>\n<p>A team with real production traffic, real observability and a direct line to the\nvendor found this <em>by hand</em>. There was no signal in the error rate, because there\nwere no errors. Nothing in the status codes, because they all said success. The only\nplace it showed up was response time under sustained load — which you see only if\nyou go looking on purpose.</p>\n<p>That matches what I ran into at a far smaller scale when I\n<a class=\"\" href=\"https://development-wec.wiline.com/docs/tutorials/load-test-llm-gateway-blocking-callback/\">load tested a gateway to find the stall the median hides</a>:\nin this kind of infrastructure, the failures that cost you most are the ones that\nnever raise an error. A blocking callback and a connection stuck waiting on a\nhandshake are completely different bugs with the same fingerprint — throughput\ncollapses, the slowest requests get much slower, and the error rate never moves.</p>\n<p>I should be straight about the limits here: I don't run anything at Pfizer's scale,\nso I can't reproduce their measurements. What I can do is read the code, and the code\nis public.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"the-fix\">The fix<a href=\"https://development-wec.wiline.com/docs/news/pfizer-gateway-silent-throughput-bug/#the-fix\" class=\"hash-link\" aria-label=\"Direct link to The fix\" title=\"Direct link to The fix\" translate=\"no\">​</a></h2>\n<p>One line, in the current version:</p>\n<div class=\"language-python codeBlockContainer_Ckt0 theme-code-block\" style=\"--prism-color:#393A34;--prism-background-color:#f6f8fa\"><div class=\"codeBlockContent_QJqH\"><pre tabindex=\"0\" class=\"prism-code language-python codeBlock_bY9V thin-scrollbar\" style=\"color:#393A34;background-color:#f6f8fa\"><code class=\"codeBlockLines_e6Vv\"><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token keyword\" style=\"color:#00009f\">if</span><span class=\"token plain\"> redis_kwargs</span><span class=\"token punctuation\" style=\"color:#393A34\">.</span><span class=\"token plain\">pop</span><span class=\"token punctuation\" style=\"color:#393A34\">(</span><span class=\"token string\" style=\"color:#e3116c\">\"ssl\"</span><span class=\"token punctuation\" style=\"color:#393A34\">,</span><span class=\"token plain\"> </span><span class=\"token boolean\" style=\"color:#36acaa\">None</span><span class=\"token punctuation\" style=\"color:#393A34\">)</span><span class=\"token punctuation\" style=\"color:#393A34\">:</span><span class=\"token plain\"></span><br></div><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\">    redis_kwargs</span><span class=\"token punctuation\" style=\"color:#393A34\">[</span><span class=\"token string\" style=\"color:#e3116c\">\"connection_class\"</span><span class=\"token punctuation\" style=\"color:#393A34\">]</span><span class=\"token plain\"> </span><span class=\"token operator\" style=\"color:#393A34\">=</span><span class=\"token plain\"> async_redis</span><span class=\"token punctuation\" style=\"color:#393A34\">.</span><span class=\"token plain\">SSLConnection</span><br></div></code></pre></div></div>\n<p>This version reads the setting's <em>value</em> instead of merely noting that it exists. Off\nnow means off.</p>\n<p>It's tracked as <em>\"Redis ssl handling — Presence check → value check\"</em> under LIT-4307\nand <a href=\"https://github.com/BerriAI/litellm/pull/32590\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">PR #32590</a>, shipped in v1.93.0.\nThe pull request title puts it plainly: <em>honor ssl value instead of key presence when\nbuilding async connection pool</em>.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"what-to-take-from-it\">What to take from it<a href=\"https://development-wec.wiline.com/docs/news/pfizer-gateway-silent-throughput-bug/#what-to-take-from-it\" class=\"hash-link\" aria-label=\"Direct link to What to take from it\" title=\"Direct link to What to take from it\" translate=\"no\">​</a></h2>\n<p>Two things, if you run anything with a cache behind it.</p>\n<p><strong>Check what your system actually did, not what you told it to do.</strong> Writing a\nsetting and having that setting take effect are different events, and this is a clean\nexample of the gap. The connection Pfizer got was the opposite of the one they asked\nfor, and nothing anywhere said so.</p>\n<p><strong>If your testing measures averages, it cannot see this.</strong> Neither can alerting built\non error rates. Sustained load, slowest-request timings, and throughput compared\nagainst a known baseline are what turn a silent stall into something visible.</p>\n<p>The dashboards stayed green through the whole thing. That's not a gap you close by\nadding another alert on errors — there weren't any to alert on.</p>",
            "url": "https://development-wec.wiline.com/docs/news/pfizer-gateway-silent-throughput-bug/",
            "title": "The bug that halved a gateway's throughput without a single error",
            "summary": "An engineer wrote a setting that said do not encrypt this connection. The system encrypted it anyway, then sat waiting for a reply that never came. Throughput fell by half, every request still returned success, and no dashboard showed a problem. Pfizer and LiteLLM published the story — and the cause is four lines of code that anyone can read.",
            "date_modified": "2026-09-16T00:00:00.000Z",
            "author": {
                "name": "Rafael Fernandes",
                "url": "https://www.linkedin.com/in/rafaelmacariofernandes/"
            },
            "tags": [
                "ai-news",
                "litellm",
                "gateway",
                "redis",
                "observability",
                "infrastructure"
            ]
        },
        {
            "id": "https://development-wec.wiline.com/docs/news/mcp-progressive-discovery-prompt-cache/",
            "content_html": "<figure class=\"newsHero newsHero--image\"><span class=\"newsHero__chip\">Protocols · AI News</span><img src=\"https://development-wec.wiline.com/docs/img/news/mcp-drops-sessions-retires-three-core-features-16x9.webp\" alt=\"Progressive tool discovery and the provider prompt cache\" loading=\"eager\"></figure>\n<p>The MCP maintainers' <a href=\"https://blog.modelcontextprotocol.io/posts/mcp-roadmap/\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">current roadmap</a>\nnames a problem most people building agents have felt without measuring: a server's\ntool catalogue is charged to the model before anyone asks a question. Under\n<em>Improved primitives</em>, the post is blunt about it —\n<em>\"Connecting to a server with a hundred tools means the model pays for that entire\nsurface before the user has asked a single question, and tool selection tends to get\nworse as the list grows.\"</em></p>\n<p>What makes this worth reading is not the roadmap. It's that the fix is already\nwritten down, in the client documentation, alongside a warning that it can cost\nmore than the problem it solves — and that warning gets far less attention than\nthe fix it qualifies.</p>\n<!-- -->\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"the-number-the-docs-put-on-it\">The number the docs put on it<a href=\"https://development-wec.wiline.com/docs/news/mcp-progressive-discovery-prompt-cache/#the-number-the-docs-put-on-it\" class=\"hash-link\" aria-label=\"Direct link to The number the docs put on it\" title=\"Direct link to The number the docs put on it\" translate=\"no\">​</a></h2>\n<p>The <a href=\"https://modelcontextprotocol.io/docs/2026-07-28/develop/clients/client-best-practices\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">client best-practices page</a>\nthat shipped with the 2026-07-28 specification describes <strong>progressive\ndiscovery</strong>: the host still calls <code>tools/list</code> as normal, but defers injecting\nthose definitions into the model's context. Instead it hands the model one\nlightweight <code>search_tools</code> meta-tool, and loads full schemas only for what comes\nback.</p>\n<p>The page's own diagram puts figures on the difference — roughly <strong>150,000 tokens</strong>\nconsumed by definitions alone when everything is loaded upfront, against about\n<strong>2,000</strong> under progressive discovery. That is the documentation's own\nillustration rather than a published benchmark, so treat it as the maintainers'\norder-of-magnitude claim, not a measurement you can reproduce. It is still a\nuseful shape: two orders of magnitude, paid before the conversation starts.</p>\n<p>The docs also give a threshold rather than leaving it to taste:</p>\n<blockquote>\n<p>Implement a threshold as a percentage of the context window. For example, 1%-5%.</p>\n</blockquote>\n<p>Below that, loading everything is fine. Above it, switch.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"the-trap\">The trap<a href=\"https://development-wec.wiline.com/docs/news/mcp-progressive-discovery-prompt-cache/#the-trap\" class=\"hash-link\" aria-label=\"Direct link to The trap\" title=\"Direct link to The trap\" translate=\"no\">​</a></h2>\n<p>Here is the sentence that changes how you'd build this. Still on the same page,\nunder <em>Interaction with Prompt Caching</em>:</p>\n<blockquote>\n<p>Adding or removing tool definitions mid-conversation invalidates that cache, and\nthe resulting miss can cost more tokens than the definitions you removed.</p>\n</blockquote>\n<p>Most providers cache the prompt prefix, and the <code>tools</code> array sits near the front\nof it. So the mechanism that saves you 148,000 tokens of upfront definitions works\nby <em>mutating the prefix</em> — which is precisely what a prompt cache cannot tolerate.\nA prefix cache is only valid up to the first thing that changed: edit the array at\nturn six and everything cached after that point goes, which on a long conversation\nis nearly all of it.</p>\n<p>The docs' own mitigations are worth reading as design constraints rather than\ntips: append new definitions strictly <em>after</em> the cache breakpoint rather than\nre-sorting the <code>tools</code> array, or route every call through a single stable\n<code>call_tool({name, args})</code> meta-tool so the array never changes at all. And treat\ndisconnecting a server as a conversation boundary, not a per-turn operation.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"three-caches-and-only-one-is-the-protocols\">Three caches, and only one is the protocol's<a href=\"https://development-wec.wiline.com/docs/news/mcp-progressive-discovery-prompt-cache/#three-caches-and-only-one-is-the-protocols\" class=\"hash-link\" aria-label=\"Direct link to Three caches, and only one is the protocol's\" title=\"Direct link to Three caches, and only one is the protocol's\" translate=\"no\">​</a></h2>\n<p>This is where it gets genuinely confusing, and it's worth separating the layers\nbecause they are easy to collapse into one.</p>\n<p><strong>The transport cache.</strong> <code>tools/list</code> results carry <code>ttlMs</code> and <code>cacheScope</code>\nhints, defined in the specification's caching utility. A client that honours them\nskips the HTTP round trip. This one is MCP's, in the sense that the protocol\nspecifies it.</p>\n<p><strong>The host-side memo.</strong> Advice rather than protocol — a line in the client\ndocumentation's implementation guidelines, recommending you memoise a fetched\ndefinition so re-injecting it later doesn't need another call. It describes what\na host should do with its own state; nothing crosses the wire. The page is\nexplicit about what it does <em>not</em> do: <em>\"This is separate from what's currently in\nthe model's context.\"</em></p>\n<p><strong>The provider's prompt cache.</strong> Not MCP's at all. Owned by whoever serves your\nmodel, keyed on the prefix, and invalidated by exactly the thing progressive\ndiscovery does for a living.</p>\n<p>We ran into the first of those the hard way while writing\n<a class=\"\" href=\"https://development-wec.wiline.com/docs/tutorials/mcp-governed-tools-agent-elicitation/\">Part 3 of the LangGraph series</a> —\na client-side <code>cache=True</code> does nothing unless the server actually advertises a\nTTL, and nothing warns you. Having now read this page properly, the more useful\nlesson is that even getting that right buys you a round trip and nothing else.\nThe schemas still land in context. Those are different problems with different\nfixes, and conflating them is the default mistake.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"what-this-means-if-youre-scoping-tools-by-hand\">What this means if you're scoping tools by hand<a href=\"https://development-wec.wiline.com/docs/news/mcp-progressive-discovery-prompt-cache/#what-this-means-if-youre-scoping-tools-by-hand\" class=\"hash-link\" aria-label=\"Direct link to What this means if you're scoping tools by hand\" title=\"Direct link to What this means if you're scoping tools by hand\" translate=\"no\">​</a></h2>\n<p><a class=\"\" href=\"https://development-wec.wiline.com/docs/tutorials/scoped-mcp-tools-per-agent/\">Part 4</a> split one MCP server between a\nscheduler and a billing agent by filtering the catalogue per role — a static\nallow-list, decided before the run and fixed for its duration.</p>\n<p>Read against these docs, that turns out to have a property worth naming: because\nthe list never changes mid-conversation, it doesn't touch the prompt prefix, so\nit sidesteps the cache-invalidation problem entirely. It is a cruder instrument\nthan <code>search_tools</code> — you decide up front instead of letting the model search —\nbut for a small catalogue split across known roles, \"crude and cache-stable\" may\nsimply be the right trade. The docs say as much in the other direction: below the\n1–5% threshold, loading everything is fine.</p>\n<p>I would not generalise that further. We have not measured cache-miss cost against\ndefinition cost on the WEC Inference API, and the honest position is that the\ncrossover depends on your catalogue size, your conversation length, and your\nprovider's caching behaviour. What the docs establish is that a crossover\n<em>exists</em>.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"on-the-roadmap-having-no-dates\">On the roadmap having no dates<a href=\"https://development-wec.wiline.com/docs/news/mcp-progressive-discovery-prompt-cache/#on-the-roadmap-having-no-dates\" class=\"hash-link\" aria-label=\"Direct link to On the roadmap having no dates\" title=\"Direct link to On the roadmap having no dates\" translate=\"no\">​</a></h2>\n<p>The roadmap names five priority areas and commits to no release date or version\nnumber anywhere. It's fair to ask why, and the fairest answer comes from the\nmaintainers themselves. Soria Parra, in the\n<a href=\"https://blog.modelcontextprotocol.io/posts/2026-mcp-roadmap/\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">March roadmap</a>:</p>\n<blockquote>\n<p>A release-oriented roadmap implies a level of predictability that open-standards\nwork rarely has.</p>\n</blockquote>\n<p>That's a reasonable position for a specification developed across working groups, and\nMarch's own record backs it up. \"Transport Evolution and Scalability\" was one of its four\npriority areas, and it named the problem precisely — running MCP at scale had surfaced\n<em>\"a consistent set of gaps: stateful sessions fight with load balancers, horizontal\nscaling requires workarounds\"</em> — while committing to no date for a fix. The July\nspecification delivered that stateless core, five months later.</p>\n<p>Caching is a different story. It appears nowhere in the March roadmap, and progressive\ndiscovery arrived as client documentation without having been on a roadmap at all — which\nis worth noting before treating either roadmap as a reliable index of what is coming.</p>\n<p>The reception hasn't been uniformly warm. The <a href=\"https://news.ycombinator.com/item?id=49399591\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">Hacker News thread</a>\non the roadmap ran to 270 points and 161 comments, and the criticism runs in three\ndirections rather than one. The cost of past churn, from <code>colingauvin</code>:</p>\n<blockquote>\n<p>It's unreal how bad the initial rollout was between HTTP/streaming and stdio,\nbearer auth and OAuth. Virtually every client/MCP server pair had a different\nportion of that matrix implemented.</p>\n</blockquote>\n<p>The design itself, from <code>nprateem</code>:</p>\n<blockquote>\n<p>The real disaster was making it stateful. Need to get some adults in the room.</p>\n</blockquote>\n<p>And whether the protocol earns its complexity at all, from <code>zackify</code>: <em>\"I think the spec\novercomplicates everything honestly.\"</em></p>\n<p>The first of those is the one that bears on this post. It's a fair grievance about\nimplementation drift, and it's the risk to watch with progressive discovery too: this is\ncurrently a <em>client-side</em> pattern with recommended strategies rather than a specified one,\nwhich means two hosts can both be reasonable and behave differently.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"the-short-version\">The short version<a href=\"https://development-wec.wiline.com/docs/news/mcp-progressive-discovery-prompt-cache/#the-short-version\" class=\"hash-link\" aria-label=\"Direct link to The short version\" title=\"Direct link to The short version\" translate=\"no\">​</a></h2>\n<p>The roadmap tells you tool-catalogue bloat is on the maintainers' list. The\ndocumentation tells you what to do about it now, and — to its credit — tells you\nin the same breath that the fix has a bill. If you're loading a large catalogue\ninto every conversation, read that page before you build a discovery layer, and\ncheck where your provider's cache breakpoint sits before you decide the token\nmaths is in your favour.</p>\n<hr>\n<p>📖 <strong>Sources:</strong> <a href=\"https://blog.modelcontextprotocol.io/posts/mcp-roadmap/\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">Model Context Protocol Blog — The New MCP Roadmap</a> · <a href=\"https://modelcontextprotocol.io/docs/2026-07-28/develop/clients/client-best-practices\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">MCP Docs — Client Best Practices</a> · <a href=\"https://blog.modelcontextprotocol.io/posts/2026-mcp-roadmap/\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">Model Context Protocol Blog — The 2026 MCP Roadmap</a> · <a href=\"https://news.ycombinator.com/item?id=49399591\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">Hacker News — New MCP Roadmap</a></p>",
            "url": "https://development-wec.wiline.com/docs/news/mcp-progressive-discovery-prompt-cache/",
            "title": "MCP's Fix for Bloated Tool Catalogues Has a Bill Attached — And It's in the Docs, Not the Roadmap",
            "summary": "The MCP roadmap names tool-catalogue bloat as a priority. The client documentation already ships the fix — progressive discovery, with a search_tools meta-tool — and the same page warns that it can cost more than it saves, because loading definitions mid-conversation invalidates the provider's prompt cache. Three layers of caching, only one of which the protocol actually specifies, and the one that bites belongs to your model provider.",
            "date_modified": "2026-09-07T00:00:00.000Z",
            "author": {
                "name": "Rafael Fernandes",
                "url": "https://www.linkedin.com/in/rafaelmacariofernandes/"
            },
            "tags": [
                "ai-news",
                "mcp",
                "agents",
                "protocols",
                "context-engineering",
                "infrastructure"
            ]
        },
        {
            "id": "https://development-wec.wiline.com/docs/news/langchain-mcp-first-class/",
            "content_html": "<figure class=\"newsHero newsHero--image\"><span class=\"newsHero__chip\">Agent Frameworks · AI News</span><img src=\"https://development-wec.wiline.com/docs/img/news/6a99a7d831c3788cf2516870_110.png\" alt=\"MCP support lands in the main LangChain package\" loading=\"eager\"></figure>\n<p>In July <a class=\"\" href=\"https://development-wec.wiline.com/docs/news/mcp-2026-07-28-spec/\">the MCP specification ripped out sessions</a>. We\nwrote at the time that the change was infrastructure, not changelog — that a\nstateless core would let servers sit behind ordinary load balancers and survive\nredeploys, and that clients would need to catch up.</p>\n<p>Today LangChain caught up. MCP support moved out of the separate\n<code>langchain-mcp-adapters</code> package and into <code>langchain</code> itself, rebuilt on FastMCP,\nwith two features the old spec made impossible: <strong>elicitation as a LangGraph\ninterrupt</strong>, and a <strong>cacheable tool catalog</strong>.</p>\n<p>We installed it the same afternoon and pointed it at a server we wrote against the\nnew spec. The announcement is accurate. The ecosystem around it is not ready — and\none of the gaps quietly converts a human-approval gate into a sentence the model\ninvents.</p>\n<!-- -->\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"what-actually-shipped\">What actually shipped<a href=\"https://development-wec.wiline.com/docs/news/langchain-mcp-first-class/#what-actually-shipped\" class=\"hash-link\" aria-label=\"Direct link to What actually shipped\" title=\"Direct link to What actually shipped\" translate=\"no\">​</a></h2>\n<p><code>pip install \"langchain[mcp]\"</code>, requiring <strong>1.4.0 or newer</strong>, and in beta — it says\nso on import, which you can see in our captures further down. Python today,\nTypeScript \"soon to follow.\"</p>\n<ul>\n<li class=\"\"><strong>One class, in the main package.</strong> <code>MultiServerMCPClient</code> collapses into\n<code>MCPAdapter</code>. <code>async with MCPAdapter(url) as adapter:</code> then\n<code>await adapter.list_tools()</code>, and what comes back are ordinary LangChain tools\nthat go anywhere tools go.</li>\n<li class=\"\"><strong>Built on FastMCP</strong>, so transports, auth (bearer, OAuth 2.1, machine-to-machine,\nCIMD, or any <code>httpx2.Auth</code>), connection management and protocol negotiation come\nfrom the client underneath. MCP now has two <strong>eras</strong> — the 2025-11-25 handshake\nprotocol and the 2026-07-28 stateless one — and the FastMCP client picks one\n<strong>per connection</strong>: \"it tries the new protocol and falls back to the handshake for\na server that hasn't upgraded.\"</li>\n<li class=\"\"><strong>Elicitation via interrupts.</strong> When a tool on the MCP server can't finish without\nasking a human, your agent's run pauses as a LangGraph <code>interrupt()</code>, and you resume\nit with a structured answer. This is only possible because the stateless spec\nturned a mid-call question into a retry-able round rather than something held\nopen on a socket.</li>\n<li class=\"\"><strong>Client-side caching.</strong> <code>fastmcp.Client(url, cache=True)</code> honours the server's\nfreshness hints so the tool catalog isn't re-fetched every run.</li>\n<li class=\"\"><strong><code>ClientGroup</code></strong> for several servers at once, each keeping its own era and\ncredentials, with tool names prefixed by server — <code>billing_search</code> and\n<code>docs_search</code> stay distinct.</li>\n</ul>\n<p>The framing in the announcement is the same one we used in July: <em>\"a redeploy no\nlonger kills live sessions, because there are none.\"</em></p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"what-the-stateless-spec-actually-removed\">What the stateless spec actually removed<a href=\"https://development-wec.wiline.com/docs/news/langchain-mcp-first-class/#what-the-stateless-spec-actually-removed\" class=\"hash-link\" aria-label=\"Direct link to What the stateless spec actually removed\" title=\"Direct link to What the stateless spec actually removed\" translate=\"no\">​</a></h2>\n<p>One tool call, before and after, is the clearest picture of what changed — so here it\nis, with a third panel for what we actually got:</p>\n<p><span class=\"zoomImage__wrap\"><img alt=\"Three sequence diagrams comparing a tool call. Under the 2025-11-25 handshake era the agent sends initialize, receives a session id, then sends tools/call with that id. Under the 2026-07-28 stateless spec there is no handshake and the agent sends tools/call directly. Against a default FastMCP server on the new spec, the same request is refused with Bad Request: Missing session ID.\" src=\"data:image/svg+xml;base64,PHN2ZyB4bWxucz0iaHR0cDovL3d3dy53My5vcmcvMjAwMC9zdmciIHZpZXdCb3g9IjAgMCA5MDAgNjkwIiB3aWR0aD0iOTAwIiBoZWlnaHQ9IjY5MCIgZm9udC1mYW1pbHk9InN5c3RlbS11aSwtYXBwbGUtc3lzdGVtLFNlZ29lIFVJLHNhbnMtc2VyaWYiPgo8dGl0bGU+Q2FsbGluZyBvbmUgdG9vbDogdGhlIGhhbmRzaGFrZSBlcmEsIHRoZSBzdGF0ZWxlc3Mgc3BlYywgYW5kIGEgZGVmYXVsdCBGYXN0TUNQIHNlcnZlciB0aGF0IHN0aWxsIHdhbnRzIGEgc2Vzc2lvbiBJRDwvdGl0bGU+CjxzdHlsZT4KICAuaW5rICAgeyBmaWxsOiAjMGYxNzJhOyB9CiAgLm11dGVkIHsgZmlsbDogIzY0NzQ4YjsgfQogIC5hY2NlbnR7IGZpbGw6ICMyNTYzZWI7IH0KICAuYmFkICAgeyBmaWxsOiAjZGMyNjI2OyB9CiAgLnBhbmVsIHsgZmlsbDogI2Y4ZmFmYzsgc3Ryb2tlOiAjY2JkNWUxOyB9CiAgLmJveCAgIHsgZmlsbDogI2ZmZmZmZjsgc3Ryb2tlOiAjOTRhM2I4OyB9CiAgLmJveGhpIHsgZmlsbDogI2ZmZmZmZjsgc3Ryb2tlOiAjMjU2M2ViOyB9CiAgLmJveGJhZHsgZmlsbDogI2ZmZmZmZjsgc3Ryb2tlOiAjZGMyNjI2OyB9CiAgLmxpZmUgIHsgc3Ryb2tlOiAjOTRhM2I4OyB9CiAgLmFycm93IHsgc3Ryb2tlOiAjMjU2M2ViOyB9CiAgLmFycm93YmFkIHsgc3Ryb2tlOiAjZGMyNjI2OyB9CiAgLm1vbm8gIHsgZm9udC1mYW1pbHk6IHVpLW1vbm9zcGFjZSxTRk1vbm8tUmVndWxhcixNZW5sbyxDb25zb2xhcyxtb25vc3BhY2U7IH0KICBAbWVkaWEgKHByZWZlcnMtY29sb3Itc2NoZW1lOiBkYXJrKSB7CiAgICAuaW5rICAgeyBmaWxsOiAjZTJlOGYwOyB9CiAgICAubXV0ZWQgeyBmaWxsOiAjOTRhM2I4OyB9CiAgICAuYWNjZW50eyBmaWxsOiAjNjBhNWZhOyB9CiAgICAuYmFkICAgeyBmaWxsOiAjZjg3MTcxOyB9CiAgICAucGFuZWwgeyBmaWxsOiAjMGYxNzJhOyBzdHJva2U6ICMzMzQxNTU7IH0KICAgIC5ib3ggICB7IGZpbGw6ICMxZTI5M2I7IHN0cm9rZTogIzY0NzQ4YjsgfQogICAgLmJveGhpIHsgZmlsbDogIzFlMjkzYjsgc3Ryb2tlOiAjNjBhNWZhOyB9CiAgICAuYm94YmFkeyBmaWxsOiAjMWUyOTNiOyBzdHJva2U6ICNmODcxNzE7IH0KICAgIC5saWZlICB7IHN0cm9rZTogIzY0NzQ4YjsgfQogICAgLmFycm93IHsgc3Ryb2tlOiAjNjBhNWZhOyB9CiAgICAuYXJyb3diYWQgeyBzdHJva2U6ICNmODcxNzE7IH0KICB9Cjwvc3R5bGU+CjxkZWZzPgogIDxtYXJrZXIgaWQ9ImFoIiB2aWV3Qm94PSIwIDAgMTAgMTAiIHJlZlg9IjkiIHJlZlk9IjUiIG1hcmtlcldpZHRoPSI3IiBtYXJrZXJIZWlnaHQ9IjciIG9yaWVudD0iYXV0by1zdGFydC1yZXZlcnNlIj4KICAgIDxwYXRoIGQ9Ik0wLDEgTDksNSBMMCw5IHoiIGZpbGw9IiMyNTYzZWIiLz4KICA8L21hcmtlcj4KICA8bWFya2VyIGlkPSJhaGJhZCIgdmlld0JveD0iMCAwIDEwIDEwIiByZWZYPSI5IiByZWZZPSI1IiBtYXJrZXJXaWR0aD0iNyIgbWFya2VySGVpZ2h0PSI3IiBvcmllbnQ9ImF1dG8tc3RhcnQtcmV2ZXJzZSI+CiAgICA8cGF0aCBkPSJNMCwxIEw5LDUgTDAsOSB6IiBmaWxsPSIjZGMyNjI2Ii8+CiAgPC9tYXJrZXI+CjwvZGVmcz4KCjx0ZXh0IHg9IjI0IiB5PSIzNCIgY2xhc3M9ImluayIgZm9udC1zaXplPSIyMyIgZm9udC13ZWlnaHQ9IjYwMCI+Q2FsbGluZyBvbmUgdG9vbDwvdGV4dD4KPHRleHQgeD0iMjQiIHk9IjU4IiBjbGFzcz0ibXV0ZWQgbW9ubyIgZm9udC1zaXplPSIxNC41Ij53aGF0IGl0IHRha2VzIHRvIHJlYWNoIHRoZSBzZXJ2ZXI8L3RleHQ+Cgo8IS0tID09PT09PT09PT09PT09PT09PT09PSBQYW5lbCAxIOKAlCBoYW5kc2hha2UgZXJhID09PT09PT09PT09PT09PT09PT09PSAtLT4KPHJlY3QgeD0iMTYiIHk9IjgwIiB3aWR0aD0iODY4IiBoZWlnaHQ9IjIxMiIgcng9IjEyIiBjbGFzcz0icGFuZWwiLz4KPHRleHQgeD0iNDAiIHk9IjExMiIgY2xhc3M9ImFjY2VudCBtb25vIiBmb250LXNpemU9IjE1Ij4yMDI1LTExLTI1PC90ZXh0Pgo8dGV4dCB4PSI0MCIgeT0iMTM0IiBjbGFzcz0iaW5rIiBmb250LXNpemU9IjE0LjUiIGZvbnQtd2VpZ2h0PSI2MDAiPmEgaGFuZHNoYWtlIGJlZm9yZSBhbnkgd29yazwvdGV4dD4KPHRleHQgeD0iNDAiIHk9IjE1OCIgY2xhc3M9Im11dGVkIiBmb250LXNpemU9IjEzIj50aGUgc2Vzc2lvbiBwaW5zIHRoZSBhZ2VudCB0byB0aGU8L3RleHQ+Cjx0ZXh0IHg9IjQwIiB5PSIxNzYiIGNsYXNzPSJtdXRlZCIgZm9udC1zaXplPSIxMyI+b25lIGluc3RhbmNlIHRoYXQgaXNzdWVkIGl0PC90ZXh0PgoKPHJlY3QgeD0iMzkyIiB5PSIxMDAiIHdpZHRoPSIxMDQiIGhlaWdodD0iMzYiIHJ4PSI4IiBjbGFzcz0iYm94Ii8+Cjx0ZXh0IHg9IjQ0NCIgeT0iMTI0IiBjbGFzcz0iaW5rIG1vbm8iIGZvbnQtc2l6ZT0iMTQiIHRleHQtYW5jaG9yPSJtaWRkbGUiPmFnZW50PC90ZXh0Pgo8cmVjdCB4PSI3NTIiIHk9IjEwMCIgd2lkdGg9IjEwNCIgaGVpZ2h0PSIzNiIgcng9IjgiIGNsYXNzPSJib3giLz4KPHRleHQgeD0iODA0IiB5PSIxMjQiIGNsYXNzPSJpbmsgbW9ubyIgZm9udC1zaXplPSIxNCIgdGV4dC1hbmNob3I9Im1pZGRsZSI+c2VydmVyPC90ZXh0PgoKPGxpbmUgeDE9IjQ0NCIgeTE9IjEzNiIgeDI9IjQ0NCIgeTI9IjI4MiIgY2xhc3M9ImxpZmUiIHN0cm9rZS1kYXNoYXJyYXk9IjMgNCIvPgo8bGluZSB4MT0iODA0IiB5MT0iMTM2IiB4Mj0iODA0IiB5Mj0iMjgyIiBjbGFzcz0ibGlmZSIgc3Ryb2tlLWRhc2hhcnJheT0iMyA0Ii8+Cgo8dGV4dCB4PSI2MjQiIHk9IjE2MyIgY2xhc3M9ImluayBtb25vIiBmb250LXNpemU9IjEzLjUiIHRleHQtYW5jaG9yPSJtaWRkbGUiPmluaXRpYWxpemU8L3RleHQ+CjxsaW5lIHgxPSI0NDQiIHkxPSIxNzIiIHgyPSI3OTgiIHkyPSIxNzIiIGNsYXNzPSJhcnJvdyIgc3Ryb2tlLXdpZHRoPSIxLjkiIG1hcmtlci1lbmQ9InVybCgjYWgpIi8+Cgo8dGV4dCB4PSI2MjQiIHk9IjE5OSIgY2xhc3M9ImluayBtb25vIiBmb250LXNpemU9IjEzLjUiIHRleHQtYW5jaG9yPSJtaWRkbGUiPnNlc3Npb24gaWQ8L3RleHQ+CjxsaW5lIHgxPSI4MDQiIHkxPSIyMDgiIHgyPSI0NTAiIHkyPSIyMDgiIGNsYXNzPSJhcnJvdyIgc3Ryb2tlLXdpZHRoPSIxLjkiIHN0cm9rZS1kYXNoYXJyYXk9IjYgNSIgbWFya2VyLWVuZD0idXJsKCNhaCkiLz4KCjx0ZXh0IHg9IjYyNCIgeT0iMjM1IiBjbGFzcz0iaW5rIG1vbm8iIGZvbnQtc2l6ZT0iMTMuNSIgdGV4dC1hbmNob3I9Im1pZGRsZSI+dG9vbHMvY2FsbCArIHNlc3Npb24gaWQ8L3RleHQ+CjxsaW5lIHgxPSI0NDQiIHkxPSIyNDQiIHgyPSI3OTgiIHkyPSIyNDQiIGNsYXNzPSJhcnJvdyIgc3Ryb2tlLXdpZHRoPSIxLjkiIG1hcmtlci1lbmQ9InVybCgjYWgpIi8+Cgo8dGV4dCB4PSI2MjQiIHk9IjI3MSIgY2xhc3M9ImluayBtb25vIiBmb250LXNpemU9IjEzLjUiIHRleHQtYW5jaG9yPSJtaWRkbGUiPnJlc3VsdDwvdGV4dD4KPGxpbmUgeDE9IjgwNCIgeTE9IjI4MCIgeDI9IjQ1MCIgeTI9IjI4MCIgY2xhc3M9ImFycm93IiBzdHJva2Utd2lkdGg9IjEuOSIgc3Ryb2tlLWRhc2hhcnJheT0iNiA1IiBtYXJrZXItZW5kPSJ1cmwoI2FoKSIvPgoKPCEtLSA9PT09PT09PT09PT09PT09PT09PT0gUGFuZWwgMiDigJQgc3RhdGVsZXNzIHNwZWMgPT09PT09PT09PT09PT09PT09PT09IC0tPgo8cmVjdCB4PSIxNiIgeT0iMzA0IiB3aWR0aD0iODY4IiBoZWlnaHQ9IjE2NCIgcng9IjEyIiBjbGFzcz0icGFuZWwiLz4KPHRleHQgeD0iNDAiIHk9IjMzNiIgY2xhc3M9ImFjY2VudCBtb25vIiBmb250LXNpemU9IjE1Ij4yMDI2LTA3LTI4PC90ZXh0Pgo8dGV4dCB4PSI0MCIgeT0iMzU4IiBjbGFzcz0iaW5rIiBmb250LXNpemU9IjE0LjUiIGZvbnQtd2VpZ2h0PSI2MDAiPm9uZSByZXF1ZXN0LCBzdHJhaWdodCB0byB3b3JrPC90ZXh0Pgo8dGV4dCB4PSI0MCIgeT0iMzgyIiBjbGFzcz0ibXV0ZWQiIGZvbnQtc2l6ZT0iMTMiPml0IGNhcnJpZXMgaXRzIG93biB2ZXJzaW9uIGFuZDwvdGV4dD4KPHRleHQgeD0iNDAiIHk9IjQwMCIgY2xhc3M9Im11dGVkIiBmb250LXNpemU9IjEzIj5pZGVudGl0eSwgc28gYW55IGluc3RhbmNlIGFuc3dlcnM8L3RleHQ+Cgo8cmVjdCB4PSIzOTIiIHk9IjMyNCIgd2lkdGg9IjEwNCIgaGVpZ2h0PSIzNiIgcng9IjgiIGNsYXNzPSJib3hoaSIvPgo8dGV4dCB4PSI0NDQiIHk9IjM0OCIgY2xhc3M9ImluayBtb25vIiBmb250LXNpemU9IjE0IiB0ZXh0LWFuY2hvcj0ibWlkZGxlIj5hZ2VudDwvdGV4dD4KPHJlY3QgeD0iNzUyIiB5PSIzMjQiIHdpZHRoPSIxMDQiIGhlaWdodD0iMzYiIHJ4PSI4IiBjbGFzcz0iYm94aGkiLz4KPHRleHQgeD0iODA0IiB5PSIzNDgiIGNsYXNzPSJpbmsgbW9ubyIgZm9udC1zaXplPSIxNCIgdGV4dC1hbmNob3I9Im1pZGRsZSI+c2VydmVyPC90ZXh0PgoKPGxpbmUgeDE9IjQ0NCIgeTE9IjM2MCIgeDI9IjQ0NCIgeTI9IjQ1OCIgY2xhc3M9ImxpZmUiIHN0cm9rZS1kYXNoYXJyYXk9IjMgNCIvPgo8bGluZSB4MT0iODA0IiB5MT0iMzYwIiB4Mj0iODA0IiB5Mj0iNDU4IiBjbGFzcz0ibGlmZSIgc3Ryb2tlLWRhc2hhcnJheT0iMyA0Ii8+Cgo8dGV4dCB4PSI2MjQiIHk9IjM4NiIgY2xhc3M9Im11dGVkIG1vbm8iIGZvbnQtc2l6ZT0iMTMuNSIgdGV4dC1hbmNob3I9Im1pZGRsZSI+bm8gaGFuZHNoYWtlPC90ZXh0PgoKPHRleHQgeD0iNjI0IiB5PSI0MTUiIGNsYXNzPSJpbmsgbW9ubyIgZm9udC1zaXplPSIxMy41IiB0ZXh0LWFuY2hvcj0ibWlkZGxlIj50b29scy9jYWxsPC90ZXh0Pgo8bGluZSB4MT0iNDQ0IiB5MT0iNDI0IiB4Mj0iNzk4IiB5Mj0iNDI0IiBjbGFzcz0iYXJyb3ciIHN0cm9rZS13aWR0aD0iMS45IiBtYXJrZXItZW5kPSJ1cmwoI2FoKSIvPgoKPHRleHQgeD0iNjI0IiB5PSI0NDciIGNsYXNzPSJpbmsgbW9ubyIgZm9udC1zaXplPSIxMy41IiB0ZXh0LWFuY2hvcj0ibWlkZGxlIj5yZXN1bHQ8L3RleHQ+CjxsaW5lIHgxPSI4MDQiIHkxPSI0NTYiIHgyPSI0NTAiIHkyPSI0NTYiIGNsYXNzPSJhcnJvdyIgc3Ryb2tlLXdpZHRoPSIxLjkiIHN0cm9rZS1kYXNoYXJyYXk9IjYgNSIgbWFya2VyLWVuZD0idXJsKCNhaCkiLz4KCjwhLS0gPT09PT09PT09PT09PT09PT09PT09IFBhbmVsIDMg4oCUIHdoYXQgd2UgYWN0dWFsbHkgZ290ID09PT09PT09PT09PT09PT09PT09PSAtLT4KPHJlY3QgeD0iMTYiIHk9IjQ4MCIgd2lkdGg9Ijg2OCIgaGVpZ2h0PSIxODgiIHJ4PSIxMiIgY2xhc3M9InBhbmVsIi8+Cjx0ZXh0IHg9IjQwIiB5PSI1MTIiIGNsYXNzPSJiYWQgbW9ubyIgZm9udC1zaXplPSIxNSI+MjAyNi0wNy0yODwvdGV4dD4KPHRleHQgeD0iNDAiIHk9IjUzMiIgY2xhc3M9Im11dGVkIiBmb250LXNpemU9IjEyLjUiPuKApmFnYWluc3QgYSBkZWZhdWx0IEZhc3RNQ1Agc2VydmVyPC90ZXh0Pgo8dGV4dCB4PSI0MCIgeT0iNTU4IiBjbGFzcz0iYmFkIiBmb250LXNpemU9IjE0LjUiIGZvbnQtd2VpZ2h0PSI2MDAiPnRoZSBjbGllbnQgbW92ZWQgb24uIHRoZSBzZXJ2ZXIgZGlkbid0LjwvdGV4dD4KPHRleHQgeD0iNDAiIHk9IjU4MiIgY2xhc3M9Im11dGVkIG1vbm8iIGZvbnQtc2l6ZT0iMTMiPnN0YXRlbGVzc19odHRwPVRydWU8L3RleHQ+Cjx0ZXh0IHg9IjQwIiB5PSI2MDAiIGNsYXNzPSJtdXRlZCIgZm9udC1zaXplPSIxMyI+aXMgb3B0LWluLCBhbmQgb2ZmIGJ5IGRlZmF1bHQ8L3RleHQ+Cgo8cmVjdCB4PSIzOTIiIHk9IjUwMCIgd2lkdGg9IjEwNCIgaGVpZ2h0PSIzNiIgcng9IjgiIGNsYXNzPSJib3hoaSIvPgo8dGV4dCB4PSI0NDQiIHk9IjUyNCIgY2xhc3M9ImluayBtb25vIiBmb250LXNpemU9IjE0IiB0ZXh0LWFuY2hvcj0ibWlkZGxlIj5hZ2VudDwvdGV4dD4KPHJlY3QgeD0iNzUyIiB5PSI1MDAiIHdpZHRoPSIxMDQiIGhlaWdodD0iMzYiIHJ4PSI4IiBjbGFzcz0iYm94YmFkIi8+Cjx0ZXh0IHg9IjgwNCIgeT0iNTI0IiBjbGFzcz0iaW5rIG1vbm8iIGZvbnQtc2l6ZT0iMTQiIHRleHQtYW5jaG9yPSJtaWRkbGUiPnNlcnZlcjwvdGV4dD4KCjxsaW5lIHgxPSI0NDQiIHkxPSI1MzYiIHgyPSI0NDQiIHkyPSI2NTAiIGNsYXNzPSJsaWZlIiBzdHJva2UtZGFzaGFycmF5PSIzIDQiLz4KPGxpbmUgeDE9IjgwNCIgeTE9IjUzNiIgeDI9IjgwNCIgeTI9IjY1MCIgY2xhc3M9ImxpZmUiIHN0cm9rZS1kYXNoYXJyYXk9IjMgNCIvPgoKPHRleHQgeD0iNjI0IiB5PSI1NjIiIGNsYXNzPSJtdXRlZCBtb25vIiBmb250LXNpemU9IjEzLjUiIHRleHQtYW5jaG9yPSJtaWRkbGUiPm5vIGhhbmRzaGFrZTwvdGV4dD4KCjx0ZXh0IHg9IjYyNCIgeT0iNTkxIiBjbGFzcz0iaW5rIG1vbm8iIGZvbnQtc2l6ZT0iMTMuNSIgdGV4dC1hbmNob3I9Im1pZGRsZSI+dG9vbHMvY2FsbDwvdGV4dD4KPGxpbmUgeDE9IjQ0NCIgeTE9IjYwMCIgeDI9Ijc5OCIgeTI9IjYwMCIgY2xhc3M9ImFycm93IiBzdHJva2Utd2lkdGg9IjEuOSIgbWFya2VyLWVuZD0idXJsKCNhaCkiLz4KCjx0ZXh0IHg9IjYyNCIgeT0iNjI3IiBjbGFzcz0iYmFkIG1vbm8iIGZvbnQtc2l6ZT0iMTMuNSIgdGV4dC1hbmNob3I9Im1pZGRsZSI+QmFkIFJlcXVlc3Q6IE1pc3Npbmcgc2Vzc2lvbiBJRDwvdGV4dD4KPGxpbmUgeDE9IjgwNCIgeTE9IjYzNiIgeDI9IjQ1MCIgeTI9IjYzNiIgY2xhc3M9ImFycm93YmFkIiBzdHJva2Utd2lkdGg9IjEuOSIgc3Ryb2tlLWRhc2hhcnJheT0iNiA1IiBtYXJrZXItZW5kPSJ1cmwoI2FoYmFkKSIvPgo8dGV4dCB4PSI2MjQiIHk9IjY1NiIgY2xhc3M9Im11dGVkIG1vbm8iIGZvbnQtc2l6ZT0iMTIiIHRleHQtYW5jaG9yPSJtaWRkbGUiPkhUVFAgNDAwIMK3IEpTT04tUlBDIOKIkjMyNjAwPC90ZXh0Pgo8L3N2Zz4K\" width=\"900\" height=\"690\" class=\"zoomImage \" loading=\"lazy\"><span class=\"zoomImage__badge\" aria-hidden=\"true\"><svg viewBox=\"0 0 24 24\" width=\"16\" height=\"16\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"2\" stroke-linecap=\"round\"><circle cx=\"11\" cy=\"11\" r=\"7\"></circle><path d=\"M21 21l-4.3-4.3\"></path><path d=\"M11 8v6M8 11h6\"></path></svg></span></span></p>\n<p>The first two panels are the pitch, and the pitch is real. The third is what a\nbrand-new server does on the afternoon the client ships.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"three-things-that-dont-work-yet\">Three things that don't work yet<a href=\"https://development-wec.wiline.com/docs/news/langchain-mcp-first-class/#three-things-that-dont-work-yet\" class=\"hash-link\" aria-label=\"Direct link to Three things that don't work yet\" title=\"Direct link to Three things that don't work yet\" translate=\"no\">​</a></h2>\n<p>We built an MCP server exposing a small booking-and-invoicing database, pointed a\nLangChain agent at it, and gated a refund behind a human. (The agent's own model runs\non a self-hosted gateway — that sits between the agent and the LLM, not between the\nagent and MCP.) It works — that's\n<a class=\"\" href=\"https://development-wec.wiline.com/docs/tutorials/mcp-governed-tools-agent-elicitation/\">the tutorial</a>.\nGetting there surfaced three gaps between what's written and what runs.</p>\n<h3 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"1-the-spec-is-stateless-your-server-isnt-by-default\">1. The spec is stateless. Your server isn't, by default.<a href=\"https://development-wec.wiline.com/docs/news/langchain-mcp-first-class/#1-the-spec-is-stateless-your-server-isnt-by-default\" class=\"hash-link\" aria-label=\"Direct link to 1. The spec is stateless. Your server isn't, by default.\" title=\"Direct link to 1. The spec is stateless. Your server isn't, by default.\" translate=\"no\">​</a></h3>\n<p>The first request to a brand-new FastMCP 4.0.2 server — here a plain <code>curl</code> asking\nfor <code>tools/list</code>, so nothing client-side can be blamed for it:</p>\n<p><span class=\"zoomImage__wrap\"><img alt=\"A curl POST to the MCP endpoint returning Bad Request: Missing session ID with JSON-RPC error code -32600\" src=\"https://development-wec.wiline.com/docs/assets/images/mcp-session-error-7897aebee33998b591bd27b4606a6e6b.png\" width=\"666\" height=\"75\" class=\"zoomImage \" loading=\"lazy\"><span class=\"zoomImage__badge\" aria-hidden=\"true\"><svg viewBox=\"0 0 24 24\" width=\"16\" height=\"16\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"2\" stroke-linecap=\"round\"><circle cx=\"11\" cy=\"11\" r=\"7\"></circle><path d=\"M21 21l-4.3-4.3\"></path><path d=\"M11 8v6M8 11h6\"></path></svg></span></span></p>\n<p>An error about a concept the specification deleted five weeks ago. The client is\nstateless; the server still defaults to the stateful transport.</p>\n<p>The fix isn't a header, a client option, or anything you send. It's an argument on\nthe call that <strong>starts your server</strong> — the last line of the server file, where <code>mcp</code>\nis your <code>FastMCP</code> instance:</p>\n<div class=\"language-python codeBlockContainer_Ckt0 theme-code-block\" style=\"--prism-color:#393A34;--prism-background-color:#f6f8fa\"><div class=\"codeBlockTitle_OeMC\">office_tools.py</div><div class=\"codeBlockContent_QJqH\"><pre tabindex=\"0\" class=\"prism-code language-python codeBlock_bY9V thin-scrollbar\" style=\"color:#393A34;background-color:#f6f8fa\"><code class=\"codeBlockLines_e6Vv\"><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token keyword\" style=\"color:#00009f\">from</span><span class=\"token plain\"> fastmcp </span><span class=\"token keyword\" style=\"color:#00009f\">import</span><span class=\"token plain\"> FastMCP</span><br></div><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\" style=\"display:inline-block\"></span><br></div><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\">mcp </span><span class=\"token operator\" style=\"color:#393A34\">=</span><span class=\"token plain\"> FastMCP</span><span class=\"token punctuation\" style=\"color:#393A34\">(</span><span class=\"token string\" style=\"color:#e3116c\">\"office-tools\"</span><span class=\"token punctuation\" style=\"color:#393A34\">)</span><span class=\"token plain\"></span><br></div><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\" style=\"display:inline-block\"></span><br></div><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\"></span><span class=\"token comment\" style=\"color:#999988;font-style:italic\"># ... your @mcp.tool functions ...</span><span class=\"token plain\"></span><br></div><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\" style=\"display:inline-block\"></span><br></div><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\"></span><span class=\"token keyword\" style=\"color:#00009f\">if</span><span class=\"token plain\"> __name__ </span><span class=\"token operator\" style=\"color:#393A34\">==</span><span class=\"token plain\"> </span><span class=\"token string\" style=\"color:#e3116c\">\"__main__\"</span><span class=\"token punctuation\" style=\"color:#393A34\">:</span><span class=\"token plain\"></span><br></div><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\">    </span><span class=\"token comment\" style=\"color:#999988;font-style:italic\"># Without stateless_http=True this server answers a 2026-07-28 client</span><span class=\"token plain\"></span><br></div><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\">    </span><span class=\"token comment\" style=\"color:#999988;font-style:italic\"># with \"Bad Request: Missing session ID\".</span><span class=\"token plain\"></span><br></div><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\">    mcp</span><span class=\"token punctuation\" style=\"color:#393A34\">.</span><span class=\"token plain\">run</span><span class=\"token punctuation\" style=\"color:#393A34\">(</span><span class=\"token plain\">transport</span><span class=\"token operator\" style=\"color:#393A34\">=</span><span class=\"token string\" style=\"color:#e3116c\">\"http\"</span><span class=\"token punctuation\" style=\"color:#393A34\">,</span><span class=\"token plain\"> host</span><span class=\"token operator\" style=\"color:#393A34\">=</span><span class=\"token string\" style=\"color:#e3116c\">\"127.0.0.1\"</span><span class=\"token punctuation\" style=\"color:#393A34\">,</span><span class=\"token plain\"> port</span><span class=\"token operator\" style=\"color:#393A34\">=</span><span class=\"token number\" style=\"color:#36acaa\">8770</span><span class=\"token punctuation\" style=\"color:#393A34\">,</span><span class=\"token plain\"> stateless_http</span><span class=\"token operator\" style=\"color:#393A34\">=</span><span class=\"token boolean\" style=\"color:#36acaa\">True</span><span class=\"token punctuation\" style=\"color:#393A34\">)</span><br></div></code></pre></div></div>\n<p>Nothing in the announcement says so — reasonably enough, since the post is about the\nclient. But it means the client half of the upgrade is one <code>pip install</code>, and the\nserver half is a flag you have to know exists.</p>\n<p>Two flags exist, and <code>run_http_async</code> (which <code>mcp.run</code> calls for HTTP) accepts both.\nFrom its own docstring:</p>\n<div class=\"language-text codeBlockContainer_Ckt0 theme-code-block\" style=\"--prism-color:#393A34;--prism-background-color:#f6f8fa\"><div class=\"codeBlockContent_QJqH\"><pre tabindex=\"0\" class=\"prism-code language-text codeBlock_bY9V thin-scrollbar\" style=\"color:#393A34;background-color:#f6f8fa\"><code class=\"codeBlockLines_e6Vv\"><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\">stateless_http: Whether to use stateless HTTP (defaults to settings.stateless_http)</span><br></div><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\">stateless: Alias for stateless_http for CLI consistency</span><br></div></code></pre></div></div>\n<p>One switch, two names, and the setting it defaults to is off.</p>\n<h3 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"2-ctxelicit-is-the-old-api--and-the-error-goes-to-the-model-not-to-you\">2. <code>ctx.elicit</code> is the old API — and the error goes to the model, not to you<a href=\"https://development-wec.wiline.com/docs/news/langchain-mcp-first-class/#2-ctxelicit-is-the-old-api--and-the-error-goes-to-the-model-not-to-you\" class=\"hash-link\" aria-label=\"Direct link to 2-ctxelicit-is-the-old-api--and-the-error-goes-to-the-model-not-to-you\" title=\"Direct link to 2-ctxelicit-is-the-old-api--and-the-error-goes-to-the-model-not-to-you\" translate=\"no\">​</a></h3>\n<p>Every elicitation example you can find calls <code>await ctx.elicit(message, response_type)</code>\ninside the tool body, where <code>ctx</code> is the <code>Context</code> object FastMCP passes to your tool.\nThat is the <strong>handshake-era</strong> mechanism: it blocks mid-execution and\nspeaks over the session's back-channel — the back-channel a stateless connection\ndoesn't have. FastMCP's own docs are explicit that it's for connections\n<code>≤ 2025-11-25</code>, and promise that calling it on a modern one \"raises a clear era\nerror rather than failing obscurely.\"</p>\n<p>It does raise one. That isn't the problem. We rebuilt the refund tool around\n<code>ctx.elicit</code> and ran the same agent against it:</p>\n<p><span class=\"zoomImage__wrap\"><img alt=\"The agent run printing interrupt raised? False, three tool messages where issue_refund has status=error carrying &amp;#39;elicitation via server-initiated requests is unavailable on 2026-07-28 connections&amp;#39;, the model replying that it has initiated the refund and a human must approve it, and a sqlite3 query showing invoice 2 still open\" src=\"https://development-wec.wiline.com/docs/assets/images/mcp-elicit-old-api-ec02d789aa6970b70a130f83067c235b.png\" width=\"1038\" height=\"437\" class=\"zoomImage \" loading=\"lazy\"><span class=\"zoomImage__badge\" aria-hidden=\"true\"><svg viewBox=\"0 0 24 24\" width=\"16\" height=\"16\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"2\" stroke-linecap=\"round\"><circle cx=\"11\" cy=\"11\" r=\"7\"></circle><path d=\"M21 21l-4.3-4.3\"></path><path d=\"M11 8v6M8 11h6\"></path></svg></span></span></p>\n<p>Read that in order. FastMCP raised the era error, exactly as documented. LangChain\ncaught it and turned it into a <code>ToolMessage</code> with <code>status=\"error\"</code>. That is deliberate:\n<code>langchain/mcp/tools.py</code> wires a handler whose docstring says it exists to hand the\nserver's own error detail to the model \"instead of ending the run.\" The model read\nthat error and told the user:</p>\n<blockquote>\n<p>I have located Maria Alvarez and identified her open invoice (ID: 2) for $80.00. I\nhave <strong>initiated the refund request</strong> for this invoice. Please note that <strong>a human\nmust approve</strong> the amount and provide a reason to complete the refund.</p>\n</blockquote>\n<p>No interrupt was raised. The process exited 0. Nothing is pending, nothing is\nwaiting for a human, and no one will ever be asked — and the invoice is still <code>open</code>,\nso the refund didn't happen either. What you get is not a destructive action slipping\npast a gate; it's an approval workflow that silently doesn't exist, described in\nfluent English by a model that read the error and paraphrased it as progress.</p>\n<p>The fix is that the modern pattern has the tool <strong>return</strong> an <code>InputRequiredResult</code>\ndescribing what it needs, and exit. Your agent surfaces that as the LangGraph interrupt,\na human answers, and the MCP client re-issues the same <code>tools/call</code> with the answer\nattached. But the failure mode is the story: an era error is a fine thing to raise\ninto a program, and a terrible thing to hand to a language model that is rewarded for\nsounding helpful.</p>\n<h3 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"3-cachetrue-is-necessary-not-sufficient\">3. <code>cache=True</code> is necessary, not sufficient<a href=\"https://development-wec.wiline.com/docs/news/langchain-mcp-first-class/#3-cachetrue-is-necessary-not-sufficient\" class=\"hash-link\" aria-label=\"Direct link to 3-cachetrue-is-necessary-not-sufficient\" title=\"Direct link to 3-cachetrue-is-necessary-not-sufficient\" translate=\"no\">​</a></h3>\n<p>The client cache respects the <code>ttlMs</code> and <code>cacheScope</code> hints a <strong>server</strong> attaches to\nits <code>tools/list</code> response, and only against modern-era servers that send them. A default\nFastMCP server sends neither — we dumped our own <code>tools/list</code> response and it carries no\ncache hints at all. So you turn the cache on, call <code>list_tools(cache_mode=\"use\")</code> twice\nback to back, time both, and get:</p>\n<p><span class=\"zoomImage__wrap\"><img alt=\"Two tool discoveries of five tools each, timed at 14.6 ms and 12.5 ms, showing no cache effect\" src=\"https://development-wec.wiline.com/docs/assets/images/mcp-cache-bbadf88f6b26a069ec15e3b3db635c02.png\" width=\"801\" height=\"125\" class=\"zoomImage \" loading=\"lazy\"><span class=\"zoomImage__badge\" aria-hidden=\"true\"><svg viewBox=\"0 0 24 24\" width=\"16\" height=\"16\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"2\" stroke-linecap=\"round\"><circle cx=\"11\" cy=\"11\" r=\"7\"></circle><path d=\"M21 21l-4.3-4.3\"></path><path d=\"M11 8v6M8 11h6\"></path></svg></span></span></p>\n<p>You conclude the cache is broken. It isn't — there was nothing to cache. Worth noting\nthe cache belongs to the <code>fastmcp.Client</code>, not to <code>MCPAdapter</code>, and one client per\ncaller keeps catalogs from crossing between tenants.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"and-two-smaller-ones-in-the-announcement-itself\">And two smaller ones, in the announcement itself<a href=\"https://development-wec.wiline.com/docs/news/langchain-mcp-first-class/#and-two-smaller-ones-in-the-announcement-itself\" class=\"hash-link\" aria-label=\"Direct link to And two smaller ones, in the announcement itself\" title=\"Direct link to And two smaller ones, in the announcement itself\" translate=\"no\">​</a></h2>\n<p><strong>The elicitation snippet raises <code>AttributeError</code>.</strong> The post reads the interrupt as\n<code>paused[\"__interrupt__\"][0].value.requests[0]</code> — attribute access. In\n<code>langchain/mcp/elicitation.py</code> on 1.4.0:</p>\n<div class=\"language-python codeBlockContainer_Ckt0 theme-code-block\" style=\"--prism-color:#393A34;--prism-background-color:#f6f8fa\"><div class=\"codeBlockContent_QJqH\"><pre tabindex=\"0\" class=\"prism-code language-python codeBlock_bY9V thin-scrollbar\" style=\"color:#393A34;background-color:#f6f8fa\"><code class=\"codeBlockLines_e6Vv\"><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token keyword\" style=\"color:#00009f\">class</span><span class=\"token plain\"> </span><span class=\"token class-name\">MCPElicitationInterrupt</span><span class=\"token punctuation\" style=\"color:#393A34\">(</span><span class=\"token plain\">TypedDict</span><span class=\"token punctuation\" style=\"color:#393A34\">)</span><span class=\"token punctuation\" style=\"color:#393A34\">:</span><span class=\"token plain\"></span><br></div><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\">    </span><span class=\"token builtin\">type</span><span class=\"token punctuation\" style=\"color:#393A34\">:</span><span class=\"token plain\"> Literal</span><span class=\"token punctuation\" style=\"color:#393A34\">[</span><span class=\"token string\" style=\"color:#e3116c\">\"mcp_elicitation\"</span><span class=\"token punctuation\" style=\"color:#393A34\">]</span><span class=\"token plain\"></span><br></div><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\">    tool_name</span><span class=\"token punctuation\" style=\"color:#393A34\">:</span><span class=\"token plain\"> </span><span class=\"token builtin\">str</span><span class=\"token plain\"></span><br></div><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\">    requests</span><span class=\"token punctuation\" style=\"color:#393A34\">:</span><span class=\"token plain\"> </span><span class=\"token builtin\">list</span><span class=\"token punctuation\" style=\"color:#393A34\">[</span><span class=\"token plain\">MCPElicitationRequest</span><span class=\"token punctuation\" style=\"color:#393A34\">]</span><br></div></code></pre></div></div>\n<p>A <code>TypedDict</code> is a dict at runtime, so <code>.requests</code> doesn't resolve. It's\n<code>value[\"requests\"]</code> — which the same snippet gets right a few lines later, reading\n<code>question[\"key\"]</code> by subscript.</p>\n<p><strong>The elicitation docs link 404s.</strong> The post closes the section by pointing at\n<code>docs.langchain.com/oss/python/langchain/mcp/elicitation</code> for \"declining a question, and\ngating destructive tools behind the same approval flow.\" That page returns 404 as of\npublication, while the parent page and its other children — <code>mcp</code>,\n<code>mcp/connections</code>, <code>mcp/tools</code>, <code>mcp/auth</code> — all resolve. It is, inconveniently, the\none page that would have documented the approval flow gap 2 shows falling over.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"why-this-matters-for-you-specifically\">Why this matters for you, specifically<a href=\"https://development-wec.wiline.com/docs/news/langchain-mcp-first-class/#why-this-matters-for-you-specifically\" class=\"hash-link\" aria-label=\"Direct link to Why this matters for you, specifically\" title=\"Direct link to Why this matters for you, specifically\" translate=\"no\">​</a></h2>\n<p>If you consume MCP servers someone else runs, this is straightforwardly good news and\nyou'll notice mostly the shorter import path.</p>\n<p>If you <strong>write</strong> MCP servers, the July revision moved the ground and the tooling is\nstill settling. Three things are now yours to get right: your server does not become\nstateless because the spec did — you set a flag; a tool that needs human input has to\nbe written to be <strong>re-entered</strong> rather than resumed — <code>interrupt()</code> unwinds the whole\ncall, so the tool body runs again from the top when you answer; and any\nelicitation example predating August is teaching you an API whose failure lands in the\nmodel's context instead of your logs.</p>\n<p>That last one generalises past MCP. As frameworks get better at keeping agents alive\nthrough errors, the class of bug that ends a run is shrinking and the class that gets\nnarrated to a user is growing. A gate that fails closed is a bug you find in testing. A\ngate that was never installed, described by a model as awaiting your approval, is one\nyou find in an audit.</p>\n<p>The direction is right. The stateless core is what makes an agent's human-approval pause\nsurvive a redeploy, and that's a real capability rather than a refactor. Just don't\nexpect your first request to succeed.</p>\n<hr>\n<p>📖 <strong>Sources:</strong> <a href=\"https://www.langchain.com/blog/mcp-in-langchain-stateless-protocol-elicitation-and-more\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">LangChain — MCP in LangChain: stateless protocol, elicitation, and more</a> · <a href=\"https://docs.langchain.com/oss/python/langchain/mcp\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">MCP in LangChain docs</a> · <a href=\"https://docs.langchain.com/oss/python/migrate/langchain-mcp-adapters\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">Migrating from langchain-mcp-adapters</a> · <a href=\"https://gofastmcp.com/clients/client\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">FastMCP client documentation</a> · <a href=\"https://gofastmcp.com/servers/elicitation\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">FastMCP elicitation</a> · <a href=\"https://modelcontextprotocol.io/specification/2026-07-28\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">MCP 2026-07-28 specification</a></p>\n<p><em>Versions under test: <code>langchain</code> 1.4.0, <code>fastmcp</code> 4.0.2, <code>mcp</code> 2.1.1, model served over a self-hosted gateway. The <code>ctx.elicit</code> capture is a render of real captured output from that run, not a screen grab.</em></p>",
            "url": "https://development-wec.wiline.com/docs/news/langchain-mcp-first-class/",
            "title": "LangChain Just Made MCP First-Class — We Ran It the Same Day, and Three Things Don't Work Yet",
            "summary": "MCP support moved into the langchain package today, built on FastMCP, with elicitation as a LangGraph interrupt and a cacheable tool catalog. We installed it the same afternoon and pointed it at a server we wrote against the new spec. The announcement is accurate; the ecosystem around it hasn't caught up, and one of the gaps turns a human-approval gate into a sentence the model makes up.",
            "date_modified": "2026-09-03T00:00:00.000Z",
            "author": {
                "name": "Rafael Fernandes",
                "url": "https://www.linkedin.com/in/rafaelmacariofernandes/"
            },
            "tags": [
                "ai-news",
                "mcp",
                "langchain",
                "langgraph",
                "agents",
                "protocols"
            ]
        },
        {
            "id": "https://development-wec.wiline.com/docs/news/agent-web-search-domain-allowlist/",
            "content_html": "<div class=\"newsHero\"><div class=\"newsHero__glow\" aria-hidden=\"true\"></div><span class=\"newsHero__eyebrow\">Agents · AI News</span><h2 class=\"newsHero__title\">Narrow only, never wider</h2><div class=\"newsHero__transition\"><span class=\"newsHero__pill newsHero__pill--from\">Please use trusted sources</span><svg xmlns=\"http://www.w3.org/2000/svg\" width=\"20\" height=\"20\" viewBox=\"0 0 24 24\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"2.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\" class=\"lucide lucide-arrow-right newsHero__arrow\" aria-hidden=\"true\"><path d=\"M5 12h14\"></path><path d=\"m12 5 7 7-7 7\"></path></svg><span class=\"newsHero__pill newsHero__pill--to\">Trusted sources are all there are</span></div></div>\n<p>Your agent can search the web. You write in the prompt: <em>only use sec.gov and the big\nfinancial wires.</em> It usually listens. The times it does not are the times you learn that\na line in a prompt is a request, not a rule.</p>\n<p>On 19 August AWS added site filters to the web search tool in Bedrock AgentCore. Small\nfeature. But the way it handles a disagreement is worth knowing, because that is what\nmakes it a guardrail instead of a note.</p>\n<!-- -->\n<div class=\"theme-admonition theme-admonition-note admonition_xJq3 alert alert--secondary\"><div class=\"admonitionHeading_Gvgb\"><span class=\"admonitionIcon_Rf37\"><svg viewBox=\"0 0 14 16\"><path fill-rule=\"evenodd\" d=\"M6.3 5.69a.942.942 0 0 1-.28-.7c0-.28.09-.52.28-.7.19-.18.42-.28.7-.28.28 0 .52.09.7.28.18.19.28.42.28.7 0 .28-.09.52-.28.7a1 1 0 0 1-.7.3c-.28 0-.52-.11-.7-.3zM8 7.99c-.02-.25-.11-.48-.31-.69-.2-.19-.42-.3-.69-.31H6c-.27.02-.48.13-.69.31-.2.2-.3.44-.31.69h1v3c.02.27.11.5.31.69.2.2.42.31.69.31h1c.27 0 .48-.11.69-.31.2-.19.3-.42.31-.69H8V7.98v.01zM7 2.3c-3.14 0-5.7 2.54-5.7 5.68 0 3.14 2.56 5.7 5.7 5.7s5.7-2.55 5.7-5.7c0-3.15-2.56-5.69-5.7-5.69v.01zM7 .98c3.86 0 7 3.14 7 7s-3.14 7-7 7-7-3.12-7-7 3.14-7 7-7z\"></path></svg></span>Whose product this is</div><div class=\"admonitionContent_BuS1\"><p>This is AWS's own documentation of an AWS service, quoted here. It does not run on our\ninfrastructure and we have not tested it. The design is what interests us.</p></div></div>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"what-you-get\">What you get<a href=\"https://development-wec.wiline.com/docs/news/agent-web-search-domain-allowlist/#what-you-get\" class=\"hash-link\" aria-label=\"Direct link to What you get\" title=\"Direct link to What you get\" translate=\"no\">​</a></h2>\n<p>Two filters. A list of sites to allow and a list to block, plus a date range for when a\npage was published.</p>\n<p>The agent can pass them in the search call itself:</p>\n<div class=\"language-json codeBlockContainer_Ckt0 theme-code-block\" style=\"--prism-color:#393A34;--prism-background-color:#f6f8fa\"><div class=\"codeBlockContent_QJqH\"><pre tabindex=\"0\" class=\"prism-code language-json codeBlock_bY9V thin-scrollbar\" style=\"color:#393A34;background-color:#f6f8fa\"><code class=\"codeBlockLines_e6Vv\"><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token punctuation\" style=\"color:#393A34\">{</span><span class=\"token plain\"></span><br></div><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\">  </span><span class=\"token property\" style=\"color:#36acaa\">\"method\"</span><span class=\"token operator\" style=\"color:#393A34\">:</span><span class=\"token plain\"> </span><span class=\"token string\" style=\"color:#e3116c\">\"tools/call\"</span><span class=\"token punctuation\" style=\"color:#393A34\">,</span><span class=\"token plain\"></span><br></div><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\">  </span><span class=\"token property\" style=\"color:#36acaa\">\"params\"</span><span class=\"token operator\" style=\"color:#393A34\">:</span><span class=\"token plain\"> </span><span class=\"token punctuation\" style=\"color:#393A34\">{</span><span class=\"token plain\"></span><br></div><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\">    </span><span class=\"token property\" style=\"color:#36acaa\">\"name\"</span><span class=\"token operator\" style=\"color:#393A34\">:</span><span class=\"token plain\"> </span><span class=\"token string\" style=\"color:#e3116c\">\"WebSearch\"</span><span class=\"token punctuation\" style=\"color:#393A34\">,</span><span class=\"token plain\"></span><br></div><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\">    </span><span class=\"token property\" style=\"color:#36acaa\">\"arguments\"</span><span class=\"token operator\" style=\"color:#393A34\">:</span><span class=\"token plain\"> </span><span class=\"token punctuation\" style=\"color:#393A34\">{</span><span class=\"token plain\"></span><br></div><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\">      </span><span class=\"token property\" style=\"color:#36acaa\">\"query\"</span><span class=\"token operator\" style=\"color:#393A34\">:</span><span class=\"token plain\"> </span><span class=\"token string\" style=\"color:#e3116c\">\"latest SEC enforcement actions 2026\"</span><span class=\"token punctuation\" style=\"color:#393A34\">,</span><span class=\"token plain\"></span><br></div><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\">      </span><span class=\"token property\" style=\"color:#36acaa\">\"filters\"</span><span class=\"token operator\" style=\"color:#393A34\">:</span><span class=\"token plain\"> </span><span class=\"token punctuation\" style=\"color:#393A34\">{</span><span class=\"token plain\"></span><br></div><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\">        </span><span class=\"token property\" style=\"color:#36acaa\">\"domainFilter\"</span><span class=\"token operator\" style=\"color:#393A34\">:</span><span class=\"token plain\"> </span><span class=\"token punctuation\" style=\"color:#393A34\">{</span><span class=\"token plain\"> </span><span class=\"token property\" style=\"color:#36acaa\">\"include\"</span><span class=\"token operator\" style=\"color:#393A34\">:</span><span class=\"token plain\"> </span><span class=\"token punctuation\" style=\"color:#393A34\">[</span><span class=\"token string\" style=\"color:#e3116c\">\"sec.gov\"</span><span class=\"token punctuation\" style=\"color:#393A34\">]</span><span class=\"token punctuation\" style=\"color:#393A34\">,</span><span class=\"token plain\"> </span><span class=\"token property\" style=\"color:#36acaa\">\"exclude\"</span><span class=\"token operator\" style=\"color:#393A34\">:</span><span class=\"token plain\"> </span><span class=\"token punctuation\" style=\"color:#393A34\">[</span><span class=\"token punctuation\" style=\"color:#393A34\">]</span><span class=\"token plain\"> </span><span class=\"token punctuation\" style=\"color:#393A34\">}</span><span class=\"token punctuation\" style=\"color:#393A34\">,</span><span class=\"token plain\"></span><br></div><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\">        </span><span class=\"token property\" style=\"color:#36acaa\">\"publishedDateFilter\"</span><span class=\"token operator\" style=\"color:#393A34\">:</span><span class=\"token plain\"> </span><span class=\"token punctuation\" style=\"color:#393A34\">{</span><span class=\"token plain\"></span><br></div><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\">          </span><span class=\"token property\" style=\"color:#36acaa\">\"from\"</span><span class=\"token operator\" style=\"color:#393A34\">:</span><span class=\"token plain\"> </span><span class=\"token string\" style=\"color:#e3116c\">\"2026-07-01T00:00:00Z\"</span><span class=\"token punctuation\" style=\"color:#393A34\">,</span><span class=\"token plain\"></span><br></div><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\">          </span><span class=\"token property\" style=\"color:#36acaa\">\"to\"</span><span class=\"token operator\" style=\"color:#393A34\">:</span><span class=\"token plain\"> </span><span class=\"token string\" style=\"color:#e3116c\">\"2026-08-04T23:59:59Z\"</span><span class=\"token plain\"></span><br></div><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\">        </span><span class=\"token punctuation\" style=\"color:#393A34\">}</span><span class=\"token plain\"></span><br></div><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\">      </span><span class=\"token punctuation\" style=\"color:#393A34\">}</span><span class=\"token plain\"></span><br></div><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\">    </span><span class=\"token punctuation\" style=\"color:#393A34\">}</span><span class=\"token plain\"></span><br></div><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\">  </span><span class=\"token punctuation\" style=\"color:#393A34\">}</span><span class=\"token plain\"></span><br></div><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\"></span><span class=\"token punctuation\" style=\"color:#393A34\">}</span><br></div></code></pre></div></div>\n<p>An admin sets the same kind of list on the gateway instead, where the agent cannot see\nor change it. Each list holds up to 100 sites.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"the-rule\">The rule<a href=\"https://development-wec.wiline.com/docs/news/agent-web-search-domain-allowlist/#the-rule\" class=\"hash-link\" aria-label=\"Direct link to The rule\" title=\"Direct link to The rule\" translate=\"no\">​</a></h2>\n<p>So the admin has a list and the agent has a list. What happens when they disagree?</p>\n<blockquote>\n<p>Allow lists are intersected. Block lists are combined.\n<em>\"Runtime filters can narrow but never expand the scope set by an administrator.\"</em></p>\n</blockquote>\n<p>The admin's list is the ceiling. The agent can ask for less, never more. And anything\neither one blocks stays blocked.</p>\n<p>Turn it around and you see why it matters. If the two allow lists were simply added\ntogether, the admin's list would be a starting suggestion. An agent that wanted some\nother site could just name it in its own call and get it. The filter would be paperwork.</p>\n<p>That is the difference between the two places you can put a rule. Asking a model to stay\ninside a boundary means asking it to remember, halfway through a job, after reading who\nknows what. A gateway comparing two lists is not remembering anything.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"the-date-filter-is-the-sneaky-one\">The date filter is the sneaky one<a href=\"https://development-wec.wiline.com/docs/news/agent-web-search-domain-allowlist/#the-date-filter-is-the-sneaky-one\" class=\"hash-link\" aria-label=\"Direct link to The date filter is the sneaky one\" title=\"Direct link to The date filter is the sneaky one\" translate=\"no\">​</a></h2>\n<p>The site list gets the attention. The date range may matter more.</p>\n<p>An old page does not look old. A four-year-old page about a tax rule or a dead API reads\nexactly like a current one — same confident tone, and now with a citation stapled to it,\nwhich makes a wrong answer more convincing rather than less. The model cannot judge how\nfresh a page is when nothing tells it. Setting a window gives it a fact it otherwise\nnever had.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"the-catch\">The catch<a href=\"https://development-wec.wiline.com/docs/news/agent-web-search-domain-allowlist/#the-catch\" class=\"hash-link\" aria-label=\"Direct link to The catch\" title=\"Direct link to The catch\" translate=\"no\">​</a></h2>\n<p>Search costs <strong>$7 per 1,000 queries</strong> and runs in three regions. AWS's selling point is\nthat the queries stay inside their network — <em>\"without sending user prompts and\nretrieval queries to external search API providers outside of AWS.\"</em> Good if you already\nlive in AWS. Mostly irrelevant if you do not.</p>\n<p>The idea travels, though, because the tool is reached over MCP: <em>\"Web Search uses a\nbuilt-in connector target on Bedrock AgentCore Gateway using the Model Context Protocol\n(MCP).\"</em> Nothing about <em>allow lists intersect, block lists combine, the agent can only\nnarrow</em> needs Amazon. Any gateway can do it, for any tool — which files something may\nread, which hosts it may reach, which tables it may query.</p>\n<p>So the question to ask about a tool is not whether you can restrict it. It is where the\nrestriction lives, and whether the agent can move it.</p>\n<p>We went through the protocol under all this in\n<a class=\"\" href=\"https://development-wec.wiline.com/docs/news/mcp-2026-07-28-spec/\">MCP's biggest update</a>, and gave a model live search without\na managed service in\n<a class=\"\" href=\"https://development-wec.wiline.com/docs/tutorials/web-search-wiline-inference/\">web search on WEC Inference</a>. The next tutorial\nin the agent series puts an agent's tools behind a gateway that decides what it may\ncall.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"sources\">Sources<a href=\"https://development-wec.wiline.com/docs/news/agent-web-search-domain-allowlist/#sources\" class=\"hash-link\" aria-label=\"Direct link to Sources\" title=\"Direct link to Sources\" translate=\"no\">​</a></h2>\n<ul>\n<li class=\"\"><a href=\"https://aws.amazon.com/blogs/machine-learning/domain-and-publish-date-filters-for-web-search-on-agentcore/\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">Domain and publish date filters for Web Search on AgentCore</a> — 19 August 2026</li>\n<li class=\"\"><a href=\"https://aws.amazon.com/about-aws/whats-new/2026/08/web-search-amazon-bedrock/\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">Web Search in Amazon Bedrock AgentCore adds domain and published date filtering, expands to Europe and Asia Pacific</a> — 19 August 2026</li>\n<li class=\"\"><a href=\"https://aws.amazon.com/blogs/aws/announcing-web-search-on-amazon-bedrock-agentcore-ground-your-ai-agents-in-current-accurate-web-knowledge/\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">Announcing Web Search on Amazon Bedrock AgentCore</a> — 17 June 2026</li>\n</ul>",
            "url": "https://development-wec.wiline.com/docs/news/agent-web-search-domain-allowlist/",
            "title": "An agent can narrow its own web search, but never widen it",
            "summary": "AWS gave agent web search a list of allowed sites on 19 August. The admin sets one list, the agent can set another, and when they disagree the agent's list can only make the search smaller. That one rule is the difference between a guardrail and a note in the prompt.",
            "date_modified": "2026-08-26T00:00:00.000Z",
            "author": {
                "name": "Rafael Fernandes",
                "url": "https://www.linkedin.com/in/rafaelmacariofernandes/"
            },
            "tags": [
                "ai-news",
                "agents",
                "mcp",
                "guardrails",
                "governance",
                "web-search"
            ]
        },
        {
            "id": "https://development-wec.wiline.com/docs/news/llm-router-cannot-classify-yes/",
            "content_html": "<div class=\"newsHero\"><div class=\"newsHero__glow\" aria-hidden=\"true\"></div><span class=\"newsHero__eyebrow\">Routing · AI News</span><h2 class=\"newsHero__title\">Classifying the word \"yes\"</h2><div class=\"newsHero__transition\"><span class=\"newsHero__pill newsHero__pill--from\">Score this message</span><svg xmlns=\"http://www.w3.org/2000/svg\" width=\"20\" height=\"20\" viewBox=\"0 0 24 24\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"2.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\" class=\"lucide lucide-arrow-right newsHero__arrow\" aria-hidden=\"true\"><path d=\"M5 12h14\"></path><path d=\"m12 5 7 7-7 7\"></path></svg><span class=\"newsHero__pill newsHero__pill--to\">Score what it approves</span></div></div>\n<p>A model router's job is to read a request and decide which model should answer it.\nCheap questions go to a small model, hard ones to a large one, and the bill comes\ndown. The whole arrangement rests on being able to tell the difference.</p>\n<p>Then a user types \"yes\".</p>\n<p>Or \"continue\". Or \"do it\". Nothing in those two or three characters says whether\nthe work being approved is a spelling fix or a database migration. A router\nscoring the current message in isolation sees a very short string with no\ntechnical vocabulary, and does the obvious thing: cheapest model.</p>\n<p><strong>Which means if you route requests to save money, your cheapest tier is probably\nabsorbing work it should never have seen — and your savings figure is partly\nfake.</strong> On 4 August LiteLLM published a benchmark that measures both halves of\nthat: how wrong the routing gets, and what it costs to fix.</p>\n<!-- -->\n<p>The question generalises past their implementation — anything classifying a turn\nin a conversation has this problem — and the reason it's worth reading is that\nthey measured the price of the fix, not only the benefit.</p>\n<div class=\"theme-admonition theme-admonition-note admonition_xJq3 alert alert--secondary\"><div class=\"admonitionHeading_Gvgb\"><span class=\"admonitionIcon_Rf37\"><svg viewBox=\"0 0 14 16\"><path fill-rule=\"evenodd\" d=\"M6.3 5.69a.942.942 0 0 1-.28-.7c0-.28.09-.52.28-.7.19-.18.42-.28.7-.28.28 0 .52.09.7.28.18.19.28.42.28.7 0 .28-.09.52-.28.7a1 1 0 0 1-.7.3c-.28 0-.52-.11-.7-.3zM8 7.99c-.02-.25-.11-.48-.31-.69-.2-.19-.42-.3-.69-.31H6c-.27.02-.48.13-.69.31-.2.2-.3.44-.31.69h1v3c.02.27.11.5.31.69.2.2.42.31.69.31h1c.27 0 .48-.11.69-.31.2-.19.3-.42.31-.69H8V7.98v.01zM7 2.3c-3.14 0-5.7 2.54-5.7 5.68 0 3.14 2.56 5.7 5.7 5.7s5.7-2.55 5.7-5.7c0-3.15-2.56-5.69-5.7-5.69v.01zM7 .98c3.86 0 7 3.14 7 7s-3.14 7-7 7-7-3.12-7-7 3.14-7 7-7z\"></path></svg></span>Whose numbers these are</div><div class=\"admonitionContent_BuS1\"><p>Everything below is LiteLLM's own measurement, published on their blog and quoted\nhere: v1.97, classifier <code>gpt-5.4-mini</code>, their three datasets, their reference\nlabels. None of it was run on our infrastructure, and we have not reproduced it.</p></div></div>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"the-measurement\">The measurement<a href=\"https://development-wec.wiline.com/docs/news/llm-router-cannot-classify-yes/#the-measurement\" class=\"hash-link\" aria-label=\"Direct link to The measurement\" title=\"Direct link to The measurement\" translate=\"no\">​</a></h2>\n<p>The sweep: <strong>5,600 live classifier calls against real providers</strong>, described as\n<em>\"two sweeps of seven configurations each (<code>classifier_context_window_size</code> of 0,\n1, 2, 3, 5, 8, 10), one with assistant turns in the window and one without, across\nthree multi-turn datasets, with two repeats per conversation.\"</em></p>\n<p>The variable is how many prior turns the classifier gets to see. Zero means it\njudges the current message alone. Ten means it reads the last ten turns first.</p>\n<p>Agreement with reference tiers, by window size:</p>\n<table><thead><tr><th>Prior turns</th><th>Short-reply follow-ups</th><th>MT-Bench 2nd turns</th><th>ShareGPT multi-turn</th></tr></thead><tbody><tr><td>0</td><td>50.0%</td><td>49.4%</td><td>83.8%</td></tr><tr><td>1</td><td>71.2%</td><td>53.1%</td><td>84.4%</td></tr><tr><td>2</td><td>87.5%</td><td>53.1%</td><td>90.6%</td></tr><tr><td><strong>3 (default)</strong></td><td><strong>85.0%</strong></td><td><strong>55.0%</strong></td><td><strong>91.9%</strong></td></tr><tr><td>5</td><td>86.2%</td><td>53.8%</td><td>91.9%</td></tr><tr><td>8</td><td>87.5%</td><td>55.6%</td><td>91.2%</td></tr><tr><td>10</td><td>90.0%</td><td>55.6%</td><td>88.8%</td></tr></tbody></table>\n<p><span class=\"zoomImage__wrap\"><img alt=\"Line chart of classifier agreement against the number of prior conversation turns, showing the history-dependent subset rising from 14% to 78% by two turns and flat thereafter\" src=\"data:image/svg+xml;base64,PHN2ZyB4bWxucz0iaHR0cDovL3d3dy53My5vcmcvMjAwMC9zdmciIHZpZXdCb3g9IjAgMCA3NjAgNDMwIiB3aWR0aD0iNzYwIiBoZWlnaHQ9IjQzMCIgZm9udC1mYW1pbHk9InN5c3RlbS11aSwtYXBwbGUtc3lzdGVtLFNlZ29lIFVJLHNhbnMtc2VyaWYiPgo8dGl0bGU+Q2xhc3NpZmllciBhZ3JlZW1lbnQgd2l0aCByZWZlcmVuY2UgdGllcnMsIGJ5IG51bWJlciBvZiBwcmlvciBjb252ZXJzYXRpb24gdHVybnM8L3RpdGxlPgo8bGluZSB4MT0iNzgiIHkxPSIzMzAuMCIgeDI9IjczMCIgeTI9IjMzMC4wIiBzdHJva2U9IiM5NGEzYjgiIHN0cm9rZS1vcGFjaXR5PSIwLjMwIi8+Cjx0ZXh0IHg9IjY2IiB5PSIzMzQuMCIgdGV4dC1hbmNob3I9ImVuZCIgZm9udC1zaXplPSIxMyIgZmlsbD0iIzY0NzQ4YiI+MCU8L3RleHQ+CjxsaW5lIHgxPSI3OCIgeTE9IjI3MC44IiB4Mj0iNzMwIiB5Mj0iMjcwLjgiIHN0cm9rZT0iIzk0YTNiOCIgc3Ryb2tlLW9wYWNpdHk9IjAuMzAiLz4KPHRleHQgeD0iNjYiIHk9IjI3NC44IiB0ZXh0LWFuY2hvcj0iZW5kIiBmb250LXNpemU9IjEzIiBmaWxsPSIjNjQ3NDhiIj4yMCU8L3RleHQ+CjxsaW5lIHgxPSI3OCIgeTE9IjIxMS42IiB4Mj0iNzMwIiB5Mj0iMjExLjYiIHN0cm9rZT0iIzk0YTNiOCIgc3Ryb2tlLW9wYWNpdHk9IjAuMzAiLz4KPHRleHQgeD0iNjYiIHk9IjIxNS42IiB0ZXh0LWFuY2hvcj0iZW5kIiBmb250LXNpemU9IjEzIiBmaWxsPSIjNjQ3NDhiIj40MCU8L3RleHQ+CjxsaW5lIHgxPSI3OCIgeTE9IjE1Mi40IiB4Mj0iNzMwIiB5Mj0iMTUyLjQiIHN0cm9rZT0iIzk0YTNiOCIgc3Ryb2tlLW9wYWNpdHk9IjAuMzAiLz4KPHRleHQgeD0iNjYiIHk9IjE1Ni40IiB0ZXh0LWFuY2hvcj0iZW5kIiBmb250LXNpemU9IjEzIiBmaWxsPSIjNjQ3NDhiIj42MCU8L3RleHQ+CjxsaW5lIHgxPSI3OCIgeTE9IjkzLjIiIHgyPSI3MzAiIHkyPSI5My4yIiBzdHJva2U9IiM5NGEzYjgiIHN0cm9rZS1vcGFjaXR5PSIwLjMwIi8+Cjx0ZXh0IHg9IjY2IiB5PSI5Ny4yIiB0ZXh0LWFuY2hvcj0iZW5kIiBmb250LXNpemU9IjEzIiBmaWxsPSIjNjQ3NDhiIj44MCU8L3RleHQ+CjxsaW5lIHgxPSI3OCIgeTE9IjM0LjAiIHgyPSI3MzAiIHkyPSIzNC4wIiBzdHJva2U9IiM5NGEzYjgiIHN0cm9rZS1vcGFjaXR5PSIwLjMwIi8+Cjx0ZXh0IHg9IjY2IiB5PSIzOC4wIiB0ZXh0LWFuY2hvcj0iZW5kIiBmb250LXNpemU9IjEzIiBmaWxsPSIjNjQ3NDhiIj4xMDAlPC90ZXh0Pgo8bGluZSB4MT0iNzgiIHkxPSIzMzAuMCIgeDI9IjczMCIgeTI9IjMzMC4wIiBzdHJva2U9IiM5NGEzYjgiLz4KPHRleHQgeD0iNzguMCIgeT0iMzU0IiB0ZXh0LWFuY2hvcj0ibWlkZGxlIiBmb250LXNpemU9IjEzIiBmaWxsPSIjNjQ3NDhiIj4wPC90ZXh0Pgo8dGV4dCB4PSIxODYuNyIgeT0iMzU0IiB0ZXh0LWFuY2hvcj0ibWlkZGxlIiBmb250LXNpemU9IjEzIiBmaWxsPSIjNjQ3NDhiIj4xPC90ZXh0Pgo8dGV4dCB4PSIyOTUuMyIgeT0iMzU0IiB0ZXh0LWFuY2hvcj0ibWlkZGxlIiBmb250LXNpemU9IjEzIiBmaWxsPSIjNjQ3NDhiIj4yPC90ZXh0Pgo8dGV4dCB4PSI0MDQuMCIgeT0iMzU0IiB0ZXh0LWFuY2hvcj0ibWlkZGxlIiBmb250LXNpemU9IjEzIiBmaWxsPSIjNjQ3NDhiIj4zPC90ZXh0Pgo8dGV4dCB4PSI1MTIuNyIgeT0iMzU0IiB0ZXh0LWFuY2hvcj0ibWlkZGxlIiBmb250LXNpemU9IjEzIiBmaWxsPSIjNjQ3NDhiIj41PC90ZXh0Pgo8dGV4dCB4PSI2MjEuMyIgeT0iMzU0IiB0ZXh0LWFuY2hvcj0ibWlkZGxlIiBmb250LXNpemU9IjEzIiBmaWxsPSIjNjQ3NDhiIj44PC90ZXh0Pgo8dGV4dCB4PSI3MzAuMCIgeT0iMzU0IiB0ZXh0LWFuY2hvcj0ibWlkZGxlIiBmb250LXNpemU9IjEzIiBmaWxsPSIjNjQ3NDhiIj4xMDwvdGV4dD4KPHRleHQgeD0iNDA0LjAiIHk9IjM4MCIgdGV4dC1hbmNob3I9Im1pZGRsZSIgZm9udC1zaXplPSIxMy41IiBmaWxsPSIjNjQ3NDhiIj5wcmlvciBjb252ZXJzYXRpb24gdHVybnMgdGhlIGNsYXNzaWZpZXIgY2FuIHNlZTwvdGV4dD4KPHBhdGggZD0iTTc4LjAsMjg4LjYgTDE4Ni43LDE5MC45IEwyOTUuMyw5OS4xIiBmaWxsPSJub25lIiBzdHJva2U9IiNkYzI2MjYiIHN0cm9rZS13aWR0aD0iMi42IiBzdHJva2UtbGluZWpvaW49InJvdW5kIi8+CjxwYXRoIGQ9Ik0yOTUuMyw5OS4xIEw3MzAuMCw5OS4xIiBmaWxsPSJub25lIiBzdHJva2U9IiNkYzI2MjYiIHN0cm9rZS13aWR0aD0iMi42IiBzdHJva2UtZGFzaGFycmF5PSI3IDYiLz4KPHRleHQgeD0iNTEyLjciIHk9Ijg4LjEiIHRleHQtYW5jaG9yPSJtaWRkbGUiIGZvbnQtc2l6ZT0iMTIuNSIgZmlsbD0iI2RjMjYyNiI+ZmxhdCB0byBOPTEwICh0aGVpciB3b3Jkcyk8L3RleHQ+CjxjaXJjbGUgY3g9Ijc4LjAiIGN5PSIyODguNiIgcj0iMy42IiBmaWxsPSIjZGMyNjI2Ii8+CjxjaXJjbGUgY3g9IjE4Ni43IiBjeT0iMTkwLjkiIHI9IjMuNiIgZmlsbD0iI2RjMjYyNiIvPgo8Y2lyY2xlIGN4PSIyOTUuMyIgY3k9Ijk5LjEiIHI9IjMuNiIgZmlsbD0iI2RjMjYyNiIvPgo8cGF0aCBkPSJNNzguMCwxODIuMCBMMTg2LjcsMTE5LjIgTDI5NS4zLDcxLjAgTDQwNC4wLDc4LjQgTDUxMi43LDc0LjggTDYyMS4zLDcxLjAgTDczMC4wLDYzLjYiIGZpbGw9Im5vbmUiIHN0cm9rZT0iIzI1NjNlYiIgc3Ryb2tlLXdpZHRoPSIyLjYiIHN0cm9rZS1saW5lam9pbj0icm91bmQiLz4KPGNpcmNsZSBjeD0iNzguMCIgY3k9IjE4Mi4wIiByPSIzLjYiIGZpbGw9IiMyNTYzZWIiLz4KPGNpcmNsZSBjeD0iMTg2LjciIGN5PSIxMTkuMiIgcj0iMy42IiBmaWxsPSIjMjU2M2ViIi8+CjxjaXJjbGUgY3g9IjI5NS4zIiBjeT0iNzEuMCIgcj0iMy42IiBmaWxsPSIjMjU2M2ViIi8+CjxjaXJjbGUgY3g9IjQwNC4wIiBjeT0iNzguNCIgcj0iMy42IiBmaWxsPSIjMjU2M2ViIi8+CjxjaXJjbGUgY3g9IjUxMi43IiBjeT0iNzQuOCIgcj0iMy42IiBmaWxsPSIjMjU2M2ViIi8+CjxjaXJjbGUgY3g9IjYyMS4zIiBjeT0iNzEuMCIgcj0iMy42IiBmaWxsPSIjMjU2M2ViIi8+CjxjaXJjbGUgY3g9IjczMC4wIiBjeT0iNjMuNiIgcj0iMy42IiBmaWxsPSIjMjU2M2ViIi8+CjxwYXRoIGQ9Ik03OC4wLDgyLjAgTDE4Ni43LDgwLjIgTDI5NS4zLDYxLjggTDQwNC4wLDU4LjAgTDUxMi43LDU4LjAgTDYyMS4zLDYwLjAgTDczMC4wLDY3LjIiIGZpbGw9Im5vbmUiIHN0cm9rZT0iIzA1OTY2OSIgc3Ryb2tlLXdpZHRoPSIyLjYiIHN0cm9rZS1saW5lam9pbj0icm91bmQiLz4KPGNpcmNsZSBjeD0iNzguMCIgY3k9IjgyLjAiIHI9IjMuNiIgZmlsbD0iIzA1OTY2OSIvPgo8Y2lyY2xlIGN4PSIxODYuNyIgY3k9IjgwLjIiIHI9IjMuNiIgZmlsbD0iIzA1OTY2OSIvPgo8Y2lyY2xlIGN4PSIyOTUuMyIgY3k9IjYxLjgiIHI9IjMuNiIgZmlsbD0iIzA1OTY2OSIvPgo8Y2lyY2xlIGN4PSI0MDQuMCIgY3k9IjU4LjAiIHI9IjMuNiIgZmlsbD0iIzA1OTY2OSIvPgo8Y2lyY2xlIGN4PSI1MTIuNyIgY3k9IjU4LjAiIHI9IjMuNiIgZmlsbD0iIzA1OTY2OSIvPgo8Y2lyY2xlIGN4PSI2MjEuMyIgY3k9IjYwLjAiIHI9IjMuNiIgZmlsbD0iIzA1OTY2OSIvPgo8Y2lyY2xlIGN4PSI3MzAuMCIgY3k9IjY3LjIiIHI9IjMuNiIgZmlsbD0iIzA1OTY2OSIvPgo8cGF0aCBkPSJNNzguMCwxODMuOCBMMTg2LjcsMTcyLjggTDI5NS4zLDE3Mi44IEw0MDQuMCwxNjcuMiBMNTEyLjcsMTcwLjggTDYyMS4zLDE2NS40IEw3MzAuMCwxNjUuNCIgZmlsbD0ibm9uZSIgc3Ryb2tlPSIjOTRhM2I4IiBzdHJva2Utd2lkdGg9IjIuNiIgc3Ryb2tlLWxpbmVqb2luPSJyb3VuZCIvPgo8Y2lyY2xlIGN4PSI3OC4wIiBjeT0iMTgzLjgiIHI9IjMuNiIgZmlsbD0iIzk0YTNiOCIvPgo8Y2lyY2xlIGN4PSIxODYuNyIgY3k9IjE3Mi44IiByPSIzLjYiIGZpbGw9IiM5NGEzYjgiLz4KPGNpcmNsZSBjeD0iMjk1LjMiIGN5PSIxNzIuOCIgcj0iMy42IiBmaWxsPSIjOTRhM2I4Ii8+CjxjaXJjbGUgY3g9IjQwNC4wIiBjeT0iMTY3LjIiIHI9IjMuNiIgZmlsbD0iIzk0YTNiOCIvPgo8Y2lyY2xlIGN4PSI1MTIuNyIgY3k9IjE3MC44IiByPSIzLjYiIGZpbGw9IiM5NGEzYjgiLz4KPGNpcmNsZSBjeD0iNjIxLjMiIGN5PSIxNjUuNCIgcj0iMy42IiBmaWxsPSIjOTRhM2I4Ii8+CjxjaXJjbGUgY3g9IjczMC4wIiBjeT0iMTY1LjQiIHI9IjMuNiIgZmlsbD0iIzk0YTNiOCIvPgo8cmVjdCB4PSI3OCIgeT0iNDA0IiB3aWR0aD0iMjAiIGhlaWdodD0iMyIgZmlsbD0iI2RjMjYyNiIvPgo8dGV4dCB4PSIxMDUiIHk9IjQxMCIgZm9udC1zaXplPSIxMyIgZmlsbD0iIzY0NzQ4YiI+MzYgaGlzdG9yeS1kZXBlbmRlbnQgZm9sbG93LXVwczwvdGV4dD4KPHJlY3QgeD0iMzUxLjEiIHk9IjQwNCIgd2lkdGg9IjIwIiBoZWlnaHQ9IjMiIGZpbGw9IiMyNTYzZWIiLz4KPHRleHQgeD0iMzc4LjEiIHk9IjQxMCIgZm9udC1zaXplPSIxMyIgZmlsbD0iIzY0NzQ4YiI+U2hvcnQtcmVwbHkgc2V0PC90ZXh0Pgo8cmVjdCB4PSI1MTAuNiIgeT0iNDA0IiB3aWR0aD0iMjAiIGhlaWdodD0iMyIgZmlsbD0iIzA1OTY2OSIvPgo8dGV4dCB4PSI1MzcuNiIgeT0iNDEwIiBmb250LXNpemU9IjEzIiBmaWxsPSIjNjQ3NDhiIj5TaGFyZUdQVDwvdGV4dD4KPHJlY3QgeD0iNjIwLjQiIHk9IjQwNCIgd2lkdGg9IjIwIiBoZWlnaHQ9IjMiIGZpbGw9IiM5NGEzYjgiLz4KPHRleHQgeD0iNjQ3LjQiIHk9IjQxMCIgZm9udC1zaXplPSIxMyIgZmlsbD0iIzY0NzQ4YiI+TVQtQmVuY2g8L3RleHQ+Cjwvc3ZnPg==\" width=\"760\" height=\"430\" class=\"zoomImage \" loading=\"lazy\"><span class=\"zoomImage__badge\" aria-hidden=\"true\"><svg viewBox=\"0 0 24 24\" width=\"16\" height=\"16\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"2\" stroke-linecap=\"round\"><circle cx=\"11\" cy=\"11\" r=\"7\"></circle><path d=\"M21 21l-4.3-4.3\"></path><path d=\"M11 8v6M8 11h6\"></path></svg></span></span></p>\n<p><strong>Figure 1.</strong> Agreement with reference tiers by context window, drawn from the figures published in LiteLLM's benchmark. The red line is the 36-follow-up subset whose difficulty only resolves against history.</p>\n<p>Three different stories in three columns, which is the first useful thing here.</p>\n<p>ShareGPT starts at 83.8% and gains eight points. MT-Bench barely moves at all —\n49.4% to 55.6% across the whole sweep — and their own caveat explains why:\n<em>\"MT-Bench's ceiling reflects its reference labels rather than router behaviour.\"</em>\nThe short-reply set is where it bites, 50% to 90%.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"the-number-worth-quoting-with-its-denominator\">The number worth quoting, with its denominator<a href=\"https://development-wec.wiline.com/docs/news/llm-router-cannot-classify-yes/#the-number-worth-quoting-with-its-denominator\" class=\"hash-link\" aria-label=\"Direct link to The number worth quoting, with its denominator\" title=\"Direct link to The number worth quoting, with its denominator\" translate=\"no\">​</a></h2>\n<p>Inside that first column sits a subset, and it is the sharpest result in the post:</p>\n<blockquote>\n<p>Agreement there is <strong>14% at N=0, 47% at N=1, and 78% at N=2</strong>, and flat from\nthere out to N=10.</p>\n</blockquote>\n<p>Fourteen per cent to seventy-eight per cent, achieved by showing the classifier\ntwo prior turns.</p>\n<p>That figure describes <strong>36 follow-ups</strong> — the ones <em>\"whose final turn only\nresolves against the history\"</em>. It is a deliberately selected hard subset, not a\ngeneral accuracy claim, and reading it as \"routers are 14% accurate\" would be\nwrong. What it does show is the shape of the failure: when a turn's difficulty\nlives entirely in what came before it, a context-free classifier is worse than\nguessing, and almost all of the recovery happens by the second turn of history.</p>\n<p>The curve going flat at N=2 is the practically useful part. This is not a\n\"more context is better\" result. It is a \"two turns is nearly all of it\" result.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"what-it-cost-and-what-it-cost-more\">What it cost, and what it cost <em>more</em><a href=\"https://development-wec.wiline.com/docs/news/llm-router-cannot-classify-yes/#what-it-cost-and-what-it-cost-more\" class=\"hash-link\" aria-label=\"Direct link to what-it-cost-and-what-it-cost-more\" title=\"Direct link to what-it-cost-and-what-it-cost-more\" translate=\"no\">​</a></h2>\n<p>This is the half that makes the finding actionable, and where the honest answer is\nmore interesting than \"it's cheap\".</p>\n<p><strong>The window itself is free.</strong> Every paired comparison against the zero-window\nconfiguration has a 95% bootstrap confidence interval straddling zero. Window 1\nwith assistant turns off comes in at −17.8 ms, interval −68.9 to +29.6. Their\nexplanation is that the window adds prefill only — the output stays a small fixed\ntier label. Over a 318-to-1,043 prompt-token range, tokens and latency correlate\nat r = 0.007.</p>\n<p><strong>The classifier is not free.</strong> Adding history costs nothing, but asking a model\nto classify at all costs <em>\"p50 sits near 600 ms in every configuration\"</em> — and\nthat sits in front of the real completion. The classifier here is\n<code>gpt-5.4-mini</code>, and it runs at <strong>$0.31–$0.61 per 1,000 requests</strong> across the whole\nsweep. Cheap in money, half a second in time.</p>\n<p><strong>And accuracy raised the bill.</strong> This is the part worth sitting with. For the\nshort-reply set, the modelled routed cost <em>rose</em> with the window — from <strong>$2.87 to\n$6.49 per 1,000 requests</strong>. Not a regression: the reason is <em>\"a tier mix of 66%\nSIMPLE at N=0 against 30% at N=2\"</em>. Two thirds of those follow-ups were being\nserved by the cheapest model. Once the classifier could see what they were\napproving, they stopped being.</p>\n<p>So the routing was cheap because it was wrong. The saving was a bill someone else\nwas paying, in answer quality, on requests nobody was auditing. On the other two\ndatasets the effect runs the other way — ShareGPT's modelled cost falls slightly,\nbecause context lets the classifier settle ambiguous prompts into MEDIUM instead\nof defaulting up to REASONING.</p>\n<p>That is the useful shape of it: better context doesn't reliably cut spend. It moves\nspend toward where the work actually was.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"two-things-this-doesnt-settle\">Two things this doesn't settle<a href=\"https://development-wec.wiline.com/docs/news/llm-router-cannot-classify-yes/#two-things-this-doesnt-settle\" class=\"hash-link\" aria-label=\"Direct link to Two things this doesn't settle\" title=\"Direct link to Two things this doesn't settle\" translate=\"no\">​</a></h2>\n<p>The default is 3, and 3 is not the best number in any column. Ten is best for\nshort-reply follow-ups at 90.0%, and simultaneously the <em>worst</em> result for\nShareGPT past N=2, dropping to 88.8% from 91.9%. More history is not monotonically\nbetter, and the shipped default is a judgement call across datasets that disagree.</p>\n<p>And they name their own limits, which is the mark of a benchmark worth trusting:\n<em>\"Reference tiers are judgement calls\"</em>, <em>\"Routed completion cost is modelled\nrather than billed\"</em>, <em>\"One classifier model was swept\"</em>, and <em>\"Latency was\nmeasured on a single VM at concurrency 10\"</em>. Concurrency 10 on one machine is not\nproduction, and a heavier classifier model carries more absolute latency than the\none they used.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"why-it-matters-if-you-route-anything\">Why it matters if you route anything<a href=\"https://development-wec.wiline.com/docs/news/llm-router-cannot-classify-yes/#why-it-matters-if-you-route-anything\" class=\"hash-link\" aria-label=\"Direct link to Why it matters if you route anything\" title=\"Direct link to Why it matters if you route anything\" translate=\"no\">​</a></h2>\n<p>The lesson survives the specific implementation. If you are choosing models per\nrequest — with a heuristic, a classifier, or a hand-rolled rule — the turns that\nwill embarrass you are not the long technical ones. They are the short ones. A\nrouter scoring \"yes\" as a trivial request routes the approval of an architecture\nmigration to the smallest model you own, and nothing in your logs will flag it,\nbecause a cheap answer to a cheap-looking question is exactly what you asked for.</p>\n<p>Their fix is a window of prior turns. The cheaper fix, if your classifier has no\nsuch option, is to notice that short replies in an established conversation are\nthe population to worry about, and to stop scoring them on their own contents.</p>\n<p>We took the same router apart from the other direction recently — the seven\nscoring dimensions, the arithmetic on a real prompt, and what a free keyword\nscorer misses that a model catches — in\n<a class=\"\" href=\"https://development-wec.wiline.com/docs/tutorials/litellm-complexity-routing/\">route by complexity</a>.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"sources\">Sources<a href=\"https://development-wec.wiline.com/docs/news/llm-router-cannot-classify-yes/#sources\" class=\"hash-link\" aria-label=\"Direct link to Sources\" title=\"Direct link to Sources\" translate=\"no\">​</a></h2>\n<ul>\n<li class=\"\"><a href=\"https://docs.litellm.ai/blog/auto-router-context-and-benchmarks\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">Auto Router v1.97: usage benchmarks and better quality for lower cost</a> — 4 August 2026</li>\n<li class=\"\"><a href=\"https://docs.litellm.ai/docs/proxy/auto_routing\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">Auto Routing configuration reference</a></li>\n</ul>",
            "url": "https://development-wec.wiline.com/docs/news/llm-router-cannot-classify-yes/",
            "title": "A router that can't see the conversation can't classify \"yes\"",
            "summary": "LiteLLM ran 5,600 live classifier calls to answer one question: how much of the conversation does a model router need to see? On the follow-ups that only make sense against history, agreement went from 14% to 78% — and the completion bill more than doubled, because two thirds of them had been going to the cheapest model.",
            "date_modified": "2026-08-25T00:00:00.000Z",
            "author": {
                "name": "Rafael Fernandes",
                "url": "https://www.linkedin.com/in/rafaelmacariofernandes/"
            },
            "tags": [
                "ai-news",
                "routing",
                "evals",
                "cost",
                "litellm"
            ]
        },
        {
            "id": "https://development-wec.wiline.com/docs/news/mask-your-logs-not-your-prompts/",
            "content_html": "<div class=\"newsHero\"><div class=\"newsHero__glow\" aria-hidden=\"true\"></div><span class=\"newsHero__eyebrow\">Privacy · AI News</span><h2 class=\"newsHero__title\">Mask your logs, not your prompts</h2><div class=\"newsHero__transition\"><span class=\"newsHero__pill newsHero__pill--from\">Redact before the model</span><svg xmlns=\"http://www.w3.org/2000/svg\" width=\"20\" height=\"20\" viewBox=\"0 0 24 24\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"2.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\" class=\"lucide lucide-arrow-right newsHero__arrow\" aria-hidden=\"true\"><path d=\"M5 12h14\"></path><path d=\"m12 5 7 7-7 7\"></path></svg><span class=\"newsHero__pill newsHero__pill--to\">Redact before the logs</span></div></div>\n<p>Almost every \"secure your LLM app\" guide gives the same advice: before a prompt reaches\nthe model, strip the personal data out of it — swap names, emails, and account numbers for\n<code>[REDACTED]</code> or <code>&lt;PERSON&gt;</code>, <em>then</em> call the model.</p>\n<p>It sounds obviously right. I assumed it was, too. Then I went and read the research on what\nmasking actually does to a model, and the papers point the other way. The short version:\n<strong>mask your logs, not your prompts.</strong></p>\n<!-- -->\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"first-what-are-we-protecting\">First, what are we protecting?<a href=\"https://development-wec.wiline.com/docs/news/mask-your-logs-not-your-prompts/#first-what-are-we-protecting\" class=\"hash-link\" aria-label=\"Direct link to First, what are we protecting?\" title=\"Direct link to First, what are we protecting?\" translate=\"no\">​</a></h2>\n<p><strong>PII</strong> is <em>personally identifiable information</em> — a name, an email, a phone number, an\naccount or government ID. Data that points at a specific person.</p>\n<p>When people reach for pre-call masking, they're usually blending two different worries into\none:</p>\n<ol>\n<li class=\"\"><em>\"I don't want to send sensitive data to whoever runs the model.\"</em></li>\n<li class=\"\"><em>\"I don't want to store sensitive data in my logs.\"</em></li>\n</ol>\n<p>Pre-call masking is aimed at <strong>#1</strong>. Hold on to that — it turns out to matter.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"problem-1-a-masked-prompt-is-a-confused-model\">Problem 1: a masked prompt is a confused model<a href=\"https://development-wec.wiline.com/docs/news/mask-your-logs-not-your-prompts/#problem-1-a-masked-prompt-is-a-confused-model\" class=\"hash-link\" aria-label=\"Direct link to Problem 1: a masked prompt is a confused model\" title=\"Direct link to Problem 1: a masked prompt is a confused model\" translate=\"no\">​</a></h2>\n<p>Here's the simplest failure. Put three people in a prompt and redact all of them to\n<code>[REDACTED]</code>, and the model can no longer tell them apart. Who signed the contract? Who\nwas cc'd? The words that carried those relationships are gone, so the answer degrades.</p>\n<p>This isn't just intuition. A 2026 benchmark called <strong>RedacBench</strong> measured the trade-off\ndirectly, across 514 texts and 187 policies. Even when a <em>capable</em> model does the redacting,\nturning the security dial up to ~81% of sensitive content removed leaves you keeping only\n<strong>37.6%</strong> of the text's non-sensitive meaning. You throw away roughly <strong>60% of what made\nthe prompt useful</strong> to buy that privacy.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"but-i-use-smart-tokenization--it-still-bites\">\"But I use smart tokenization\" — it still bites<a href=\"https://development-wec.wiline.com/docs/news/mask-your-logs-not-your-prompts/#but-i-use-smart-tokenization--it-still-bites\" class=\"hash-link\" aria-label=\"Direct link to &quot;But I use smart tokenization&quot; — it still bites\" title=\"Direct link to &quot;But I use smart tokenization&quot; — it still bites\" translate=\"no\">​</a></h2>\n<p>The obvious fix is to stop using dumb placeholders. Deterministic tokenization maps the\nsame value to the same token every time — \"John Smith\" always becomes <code>PERSON_42</code>,\n\"Jane Doe\" always <code>PERSON_17</code> — so the model can still track who did what. That's genuinely\nbetter.</p>\n<p>But it isn't free either. One engineer ran <strong>109 masking tests</strong> across healthcare, legal,\nfinancial, and developer workflows and wrote up where it broke. <em>(Full disclosure: he also\nsells a tokenization tool, so take his framing with a grain of salt — but the failures he\nlogged are concrete and easy to reproduce.)</em></p>\n<ul>\n<li class=\"\"><strong>Context-phrase refusals.</strong> Tokenize an SSN into <code>GOV_ID_8x3m</code>, but leave the words\n\"social security number\" sitting next to it, and the model's safety filter can refuse the\nwhole request — it sees a sensitive label beside an opaque token and flags it.</li>\n<li class=\"\"><strong>False positives.</strong> The word \"Will\" in <em>\"this will update the record\"</em> got caught by the\nname detector.</li>\n<li class=\"\"><strong>Misses.</strong> A short name like \"Li\" in a table row, or an SSN buried in code comments, slid\npast the detector when there wasn't enough surrounding context.</li>\n<li class=\"\"><strong>Streaming corruption.</strong> A name or SSN can be split across two or three streaming chunks;\nprocess them one at a time and you mangle the entity.</li>\n</ul>\n<p>His overall detection came out to <strong>89%</strong> — and his sharper point was that the missing 11%\nis where it hurts: you're forced to either block a legitimate request or leak. There's no\ncomfortable default.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"why-none-of-this-is-surprising-my-read-not-a-proof\">Why none of this is surprising (my read, not a proof)<a href=\"https://development-wec.wiline.com/docs/news/mask-your-logs-not-your-prompts/#why-none-of-this-is-surprising-my-read-not-a-proof\" class=\"hash-link\" aria-label=\"Direct link to Why none of this is surprising (my read, not a proof)\" title=\"Direct link to Why none of this is surprising (my read, not a proof)\" translate=\"no\">​</a></h2>\n<p>Step back and it fits a pattern the research keeps finding: <strong>models are fragile to how a\nprompt is worded.</strong></p>\n<ul>\n<li class=\"\"><em>\"On the Worst Prompt Performance of LLMs\"</em> took the same question, reworded it in\nsemantically identical ways, and watched one model's accuracy swing by <strong>45 points</strong>\n(worst case, 9.38%).</li>\n<li class=\"\">The <strong>DETAIL</strong> framework found that <em>more specific</em> prompts reason better, especially on\nsmaller models and step-by-step tasks.</li>\n</ul>\n<p>Neither of those papers tested PII masking — so this next step is <em>my inference, not their\nclaim</em> — but masking <strong>is</strong> a prompt edit. It makes the prompt less specific and changes its\nwording, which is exactly the lever these papers show models are sensitive to. You're\nrolling dice you don't need to roll.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"the-hidden-bill\">The hidden bill<a href=\"https://development-wec.wiline.com/docs/news/mask-your-logs-not-your-prompts/#the-hidden-bill\" class=\"hash-link\" aria-label=\"Direct link to The hidden bill\" title=\"Direct link to The hidden bill\" translate=\"no\">​</a></h2>\n<p>Masking also costs money in a way that's easy to miss, because it changes the prompt on\n<strong>every</strong> call. That quietly breaks <strong>prompt caching</strong> — the discount you get when a prompt's\nopening is identical to a previous one.</p>\n<p>Both major providers cache by prefix, and both say a change up front invalidates it:</p>\n<ul>\n<li class=\"\">OpenAI: <em>\"Cache hits are only possible for exact prefix matches… a change before the\nbreakpoint will prevent a cache hit.\"</em></li>\n<li class=\"\">Anthropic: <em>\"Changes at each level invalidate that level and all subsequent levels.\"</em></li>\n</ul>\n<p>A cache read costs about <strong>10% of the normal input-token price</strong>. So every masked prompt\nthat misses the cache pays close to <strong>ten times more</strong> on those tokens, and gives up the\nfaster first token too. On top of that, when a masked placeholder like <code>PERSON_42</code> leaks\ninto a tool call, the tool rejects it and the agent retries — and a study of coding agents\nfound that small prompt-wording changes can multiply token use <strong>2.4–7.4×</strong> with no gain in\nsuccess. (Again: not masking specifically, but the same mechanism.)</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"the-reframe-the-model-was-never-the-risk\">The reframe: the model was never the risk<a href=\"https://development-wec.wiline.com/docs/news/mask-your-logs-not-your-prompts/#the-reframe-the-model-was-never-the-risk\" class=\"hash-link\" aria-label=\"Direct link to The reframe: the model was never the risk\" title=\"Direct link to The reframe: the model was never the risk\" translate=\"no\">​</a></h2>\n<p>Here's the part I had backwards. <strong>The problem was never that the model <em>sees</em> the data.</strong>\nA model reading your prompt to answer it is just doing its job — that inference pass isn't\nwhere data leaks. The real question is <strong>retention</strong>: what gets <em>stored</em>, and where.</p>\n<p>Once you see it that way, the fix is obvious. Put the masking where storage actually happens\nand where you're in control: <strong>your logs.</strong></p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"mask-your-logs-not-your-prompts\">Mask your logs, not your prompts<a href=\"https://development-wec.wiline.com/docs/news/mask-your-logs-not-your-prompts/#mask-your-logs-not-your-prompts\" class=\"hash-link\" aria-label=\"Direct link to Mask your logs, not your prompts\" title=\"Direct link to Mask your logs, not your prompts\" translate=\"no\">​</a></h2>\n<p>The pattern — often called <strong>logging-only</strong> — is simple:</p>\n<ul>\n<li class=\"\">The <strong>raw</strong> prompt reaches the model, so reasoning stays intact, the cache still hits, and\nnothing gets refused.</li>\n<li class=\"\">Your PII masking runs <strong>only on the path to storage</strong> — the request/response logs and\ntraces your gateway writes.</li>\n</ul>\n<p>Think of it as three rungs, worst to best:</p>\n<ol>\n<li class=\"\"><strong>Dumb masking</strong> (<code>[REDACTED]</code>) — destroys entity relationships.</li>\n<li class=\"\"><strong>Tokenization</strong> (<code>PERSON_42</code>) — keeps identities, but still triggers refusals and misses.</li>\n<li class=\"\"><strong>Logging-only</strong> — the model reads the real text; masking happens where the data rests.</li>\n</ol>\n<p>One honest caveat: whatever provider you send prompts to, it's worth knowing its retention\npolicy — that's a separate question from what the model reads, and it's the <em>right</em> place to\nput that worry.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"if-you-run-your-own-gateway\">If you run your own gateway<a href=\"https://development-wec.wiline.com/docs/news/mask-your-logs-not-your-prompts/#if-you-run-your-own-gateway\" class=\"hash-link\" aria-label=\"Direct link to If you run your own gateway\" title=\"Direct link to If you run your own gateway\" translate=\"no\">​</a></h2>\n<p>The nice part: this is a <strong>configuration</strong>, not a rewrite. If you've already\n<a class=\"\" href=\"https://development-wec.wiline.com/docs/tutorials/deploy-llm-gateway-wec-instance/\">deployed a gateway on a WEC Instance</a>, the\nguardrail runs in log-only mode — clean prompt out to the model, masked copy into the logs.\nA follow-up tutorial will wire it up end to end.</p>\n<p>Masking a prompt before the model buys you a compliance checkbox and quietly sells your\nmodel's intelligence. Put the privacy work where the data actually rests — in the logs — and\nlet the model read the real thing.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"sources\">Sources<a href=\"https://development-wec.wiline.com/docs/news/mask-your-logs-not-your-prompts/#sources\" class=\"hash-link\" aria-label=\"Direct link to Sources\" title=\"Direct link to Sources\" translate=\"no\">​</a></h2>\n<ul>\n<li class=\"\"><a href=\"https://arxiv.org/abs/2603.20208\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">RedacBench: Can AI Erase Your Secrets? — arXiv 2603.20208</a> — 80.9% security / 37.6% utility at aggressive redaction</li>\n<li class=\"\"><a href=\"https://arxiv.org/abs/2406.10248\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">On the Worst Prompt Performance of Large Language Models — arXiv 2406.10248</a> — 45.48% swing, 9.38% worst case</li>\n<li class=\"\"><a href=\"https://arxiv.org/abs/2512.02246\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">DETAIL Matters: Prompt Specificity and Reasoning — arXiv 2512.02246</a></li>\n<li class=\"\"><a href=\"https://arxiv.org/abs/2608.01347\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">Prompt-Induced Waste in Coding Agents — arXiv 2608.01347</a> — 2.4–7.4× token multiplication</li>\n<li class=\"\"><a href=\"https://www.reddit.com/r/LLMDevs/comments/1sgvjrg/deterministic_tokenization_vs_masking_for_pii_in/\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">Deterministic tokenization vs. masking for PII: 109 tests — r/LLMDevs</a> — practitioner write-up (author sells a tokenization tool)</li>\n<li class=\"\"><a href=\"https://developers.openai.com/api/docs/guides/prompt-caching\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">OpenAI — Prompt caching</a></li>\n<li class=\"\"><a href=\"https://platform.claude.com/docs/en/build-with-claude/prompt-caching\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">Anthropic — Prompt caching</a></li>\n</ul>",
            "url": "https://development-wec.wiline.com/docs/news/mask-your-logs-not-your-prompts/",
            "title": "Mask Your Logs, Not Your Prompts",
            "summary": "The most common LLM privacy advice — scrub personal data out of the prompt before the model sees it — aims at the wrong risk and quietly makes the model dumber. Here's what the research actually shows, and where the masking really belongs.",
            "date_modified": "2026-08-14T00:00:00.000Z",
            "author": {
                "name": "Rafael Fernandes",
                "url": "https://www.linkedin.com/in/rafaelmacariofernandes/"
            },
            "tags": [
                "ai-news",
                "privacy",
                "pii",
                "guardrails",
                "gateway"
            ]
        },
        {
            "id": "https://development-wec.wiline.com/docs/news/spec-driven-development-solution-to-vibe-coding/",
            "content_html": "<div class=\"newsHero\"><div class=\"newsHero__glow\" aria-hidden=\"true\"></div><span class=\"newsHero__eyebrow\">Engineering practice · AI News</span><h2 class=\"newsHero__title\">Spec-driven development, tested</h2><div class=\"newsHero__transition\"><span class=\"newsHero__pill newsHero__pill--from\">Write the prompt</span><svg xmlns=\"http://www.w3.org/2000/svg\" width=\"20\" height=\"20\" viewBox=\"0 0 24 24\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"2.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\" class=\"lucide lucide-arrow-right newsHero__arrow\" aria-hidden=\"true\"><path d=\"M5 12h14\"></path><path d=\"m12 5 7 7-7 7\"></path></svg><span class=\"newsHero__pill newsHero__pill--to\">Write the contract</span></div></div>\n<p>Someone posted spec-driven development on LinkedIn this week as <em>the</em> answer to vibe coding —\nto prompting an agent, half-understanding what you're building, and ending up with code you\ncan't vouch for. The linked toolkit has 127,000 stars and comes from GitHub itself. The pitch\nlands.</p>\n<p>So I installed it and pointed it at a deliberately trivial task. One of the three principles it\nwrote for me was a dependency policy I never asked for — hold that thought.</p>\n<p>Twenty minutes isn't a verdict, though. Two engineers have tested this properly, on real\nproblems, long enough for the seams to show. They used different tools, on different\ncontinents, seven months apart — and both reached for the same comparison, unprompted: the last\ntime our industry tried to generate working code from documents. On the one question that\ndecides whether any of this survives contact with AI features, they flatly contradict each\nother. Neither has a measurement.</p>\n<!-- -->\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"what-spec-driven-development-is\">What spec-driven development is<a href=\"https://development-wec.wiline.com/docs/news/spec-driven-development-solution-to-vibe-coding/#what-spec-driven-development-is\" class=\"hash-link\" aria-label=\"Direct link to What spec-driven development is\" title=\"Direct link to What spec-driven development is\" translate=\"no\">​</a></h2>\n<p>If the term is new to you, the idea is simple. Instead of prompting an agent and iterating until\nthe code looks right, you write a structured specification first — what you're building, why,\nand what \"done\" means — and the agent works from that document rather than from your prompt.\nThe spec, not the code, becomes the thing everyone points at.</p>\n<p><strong>Spec Kit</strong> is GitHub's implementation. It installs into a repo once, not per task, and gives\nyour agent a set of commands: <code>constitution</code> to set project principles, then <code>specify</code>, <code>plan</code>,\n<code>tasks</code>, <code>implement</code>. Nothing runs in the background — it drops templates and command\ndefinitions into your project, and from then on you're talking to your agent. Closer to a linter\nconfig than a product.</p>\n<!-- -->\n<p>Before any of that, though, comes <code>constitution</code> — written first, before anyone has seen the\nproblem, and everything downstream gets judged against it. Hold on to that too.</p>\n<p>The setup step tells you what it's really aiming at:</p>\n<p><span class=\"zoomImage__wrap\"><img alt=\"Spec Kit&amp;#39;s setup prompt listing more than thirty coding agent integrations, including Claude Code, GitHub Copilot, Cursor, Gemini CLI, Devin, Grok and IBM Bob.\" src=\"https://development-wec.wiline.com/docs/assets/images/spec-kit-agents-f508d78c53fe54f1ad55cd522d110859.png\" width=\"2236\" height=\"1282\" class=\"zoomImage \" loading=\"lazy\"><span class=\"zoomImage__badge\" aria-hidden=\"true\"><svg viewBox=\"0 0 24 24\" width=\"16\" height=\"16\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"2\" stroke-linecap=\"round\"><circle cx=\"11\" cy=\"11\" r=\"7\"></circle><path d=\"M21 21l-4.3-4.3\"></path><path d=\"M11 8v6M8 11h6\"></path></svg></span></span></p>\n<p>Thirty-plus agents. This isn't a GitHub-only tool — it wants to be the layer every coding agent\nplugs into, which goes a long way to explaining the star count.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"what-actually-happens-when-you-use-it\">What actually happens when you use it<a href=\"https://development-wec.wiline.com/docs/news/spec-driven-development-solution-to-vibe-coding/#what-actually-happens-when-you-use-it\" class=\"hash-link\" aria-label=\"Direct link to What actually happens when you use it\" title=\"Direct link to What actually happens when you use it\" translate=\"no\">​</a></h2>\n<p><strong>Birgitta Böckeler</strong>, Distinguished Engineer at Thoughtworks, trialled three of these tools by\nhand — Kiro, Spec Kit and Tessl — and\n<a href=\"https://martinfowler.com/articles/exploring-gen-ai/sdd-3-tools.html\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">published what she found</a>\nin October 2025.</p>\n<p>She asked Kiro to fix a small bug. It produced four user stories and sixteen acceptance\ncriteria, including — verbatim — <em>\"As a developer, I want the transformation function to handle\nedge cases gracefully, so that the system remains robust when new category formats are\nintroduced.\"</em> Her summary: <em>\"like using a sledgehammer to crack a nut.\"</em></p>\n<p>On Spec Kit with a real feature — a few days' work, by her own estimate — she never\nfinished the implementation, and reckons she could have built the thing by hand in the time she\nspent reviewing artifacts. The line that will land with anyone who has done a code review:</p>\n<blockquote>\n<p>To be honest, I'd rather review code than all these markdown files.</p>\n</blockquote>\n<p>She also found the agent ignoring the documents meant to steer it. Spec Kit's research step\ncorrectly catalogued existing classes; the agent then read those descriptions as a\nspecification and generated the classes again, as duplicates. And the reverse failure — the\nagent going <em>\"way overboard because it was too eagerly following instructions (e.g. one of the\nconstitution articles).\"</em></p>\n<p><strong>Alex Punnen</strong> went narrower and deeper, and\n<a href=\"https://github.com/alexcpn/speckit_test\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">published the whole transcript</a> with line citations.\nSpec Kit v0.8.9, on a problem at real scale: querying US elevation data across 1,756 map tiles\n— 23 billion measurements, 180GB.</p>\n<p>His finding is subtler than \"the tool invents things.\" He asked for principles covering code\nquality, testing, consistency and performance, and says plainly: <em>\"The principles themselves are\nreasonable.\"</em> What went wrong, he argues, is that one of them quietly tilted every later\ndecision:</p>\n<blockquote>\n<p>New dependencies MUST be justified in writing: problem solved, alternatives considered,\nlicense verified.</p>\n</blockquote>\n<p>Sensible in isolation. But as Punnen puts it, <em>\"the stdlib option always wins ties because it\ncosts zero justification entries.\"</em> Four phases later the plan chose SQLite because — first\nreason listed — <em>\"Stdlib, zero new dependency… that's the cheapest possible answer.\"</em> Three of\nthe four rejected alternatives fell to that same rule. His verdict: <strong>\"The constitution did the\nrejecting; the agent was just the microphone.\"</strong></p>\n<p>Then comes the part I found hardest to shake. Earlier in the process the agent had written a\nfive-minute performance budget into the spec — a number it guessed, with no measurement behind\nit, filed under the label SC-008. Later, that guess came back as the reason a better design\ncouldn't be used: <em>\"The transcode cost blows past SC-008.\"</em> Only under direct pushback did it\nconcede: <em>\"You're right that I overweighted reason #2.\"</em> The better design, Punnen writes, needed\nnothing that wasn't already in the spec, <em>\"except the willingness to revise an arbitrary number\nthe spec itself produced.\"</em></p>\n<p>His summary of that phase applies to the whole category:</p>\n<blockquote>\n<p>The artefact looks done because every template slot is filled — not because the engineering\nquestion is answered.</p>\n</blockquote>\n<p>Two things make this more than a one-off. His constitution prompt was essentially <strong>GitHub's own\ndocumented example</strong>, lightly adapted — he followed the quickstart. And the resistance he hit is\npartly by design: GitHub's methodology document calls the constitution <em>\"a set of <strong>immutable\nprinciples</strong>,\"</em> with a section headed <em>\"The Power of Immutable Principles.\"</em></p>\n<p>To be precise, the five-minute budget was not a constitutional principle — it was a success\ncriterion the agent generated downstream of one. The document never claims those are immutable.\nBut the number carried that authority anyway, and it took a human to dislodge it. The\nmethodology argues for fixing principles and says nothing about what happens when guesses\nderived from them inherit the same standing.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"the-three-levels-nobody-agrees-on\">The three levels nobody agrees on<a href=\"https://development-wec.wiline.com/docs/news/spec-driven-development-solution-to-vibe-coding/#the-three-levels-nobody-agrees-on\" class=\"hash-link\" aria-label=\"Direct link to The three levels nobody agrees on\" title=\"Direct link to The three levels nobody agrees on\" translate=\"no\">​</a></h2>\n<p>Böckeler's most useful contribution is a distinction the rest of the debate skips. \"Spec-driven\ndevelopment\" covers three different practices:</p>\n<figure class=\"stageFlow\"><div class=\"stageFlow__track\"><div class=\"stageFlow__card\" style=\"background:rgba(var(--primary-rgb), 0.050);border-color:rgba(var(--primary-rgb), 0.250)\"><span class=\"stageFlow__stage\">Spec-first</span><span class=\"stageFlow__title\">Spec written first, used for the task at hand</span><span class=\"stageFlow__tag\">all tools do this</span></div><svg xmlns=\"http://www.w3.org/2000/svg\" width=\"22\" height=\"22\" viewBox=\"0 0 24 24\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"2.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\" class=\"lucide lucide-arrow-right stageFlow__arrow\" aria-hidden=\"true\"><path d=\"M5 12h14\"></path><path d=\"m12 5 7 7-7 7\"></path></svg><div class=\"stageFlow__card\" style=\"background:rgba(var(--primary-rgb), 0.160);border-color:rgba(var(--primary-rgb), 0.450)\"><span class=\"stageFlow__stage\">Spec-anchored</span><span class=\"stageFlow__title\">Spec kept and maintained after the task</span><span class=\"stageFlow__tag\">few tools reach</span></div><svg xmlns=\"http://www.w3.org/2000/svg\" width=\"22\" height=\"22\" viewBox=\"0 0 24 24\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"2.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\" class=\"lucide lucide-arrow-right stageFlow__arrow\" aria-hidden=\"true\"><path d=\"M5 12h14\"></path><path d=\"m12 5 7 7-7 7\"></path></svg><div class=\"stageFlow__card\" style=\"background:rgba(var(--primary-rgb), 0.270);border-color:rgba(var(--primary-rgb), 0.650)\"><span class=\"stageFlow__stage\">Spec-as-source</span><span class=\"stageFlow__title\">Only the spec is edited; human never touches code</span><span class=\"stageFlow__tag\">Tessl only</span></div></div><figcaption class=\"stageFlow__caption\">Böckeler's taxonomy, October 2025. IBM published the same three names seven months later.</figcaption></figure>\n<p>Her verdict on where the tools actually sit: <em>\"All SDD approaches and definitions I've found are\nspec-first, but not all strive to be spec-anchored or spec-as-source.\"</em> Including Spec Kit.\nGitHub's methodology aspires far higher — <em>\"Specifications don't serve code—code serves\nspecifications\"</em> — but Spec Kit creates a <strong>branch per spec</strong>, so a spec lives for the lifetime\nof a change request, not a feature. Her conclusion: <em>\"spec-kit is still what I would call\nspec-first only, not spec-anchored over time.\"</em></p>\n<p>Worth knowing that IBM published <a href=\"https://www.ibm.com/think/topics/spec-driven-development\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">the identical three-level\ntaxonomy</a> seven months later,\nfootnoted. If you've seen it credited to IBM, it's hers.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"what-i-got-on-a-trivial-task\">What I got on a trivial task<a href=\"https://development-wec.wiline.com/docs/news/spec-driven-development-solution-to-vibe-coding/#what-i-got-on-a-trivial-task\" class=\"hash-link\" aria-label=\"Direct link to What I got on a trivial task\" title=\"Direct link to What I got on a trivial task\" translate=\"no\">​</a></h2>\n<p>I ran Spec Kit at commit <code>83883a2</code> on a deliberately minimal prompt — <em>\"Principles for a small\nPython utility. Keep it minimal — I have no strong constraints.\"</em></p>\n<p><span class=\"zoomImage__wrap\"><img alt=\"The Spec Kit constitution phase running in a terminal, generating principles from a one-line prompt.\" src=\"https://development-wec.wiline.com/docs/assets/images/spec-kit-constitution-2e19dee68aab5c7e2a098a49a54edc1b.png\" width=\"2176\" height=\"1288\" class=\"zoomImage \" loading=\"lazy\"><span class=\"zoomImage__badge\" aria-hidden=\"true\"><svg viewBox=\"0 0 24 24\" width=\"16\" height=\"16\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"2\" stroke-linecap=\"round\"><circle cx=\"11\" cy=\"11\" r=\"7\"></circle><path d=\"M21 21l-4.3-4.3\"></path><path d=\"M11 8v6M8 11h6\"></path></svg></span></span></p>\n<p>One of the three principles it wrote:</p>\n<div class=\"language-markdown codeBlockContainer_Ckt0 theme-code-block\" style=\"--prism-color:#393A34;--prism-background-color:#f6f8fa\"><div class=\"codeBlockContent_QJqH\"><pre tabindex=\"0\" class=\"prism-code language-markdown codeBlock_bY9V thin-scrollbar\" style=\"color:#393A34;background-color:#f6f8fa\"><code class=\"codeBlockLines_e6Vv\"><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token title important punctuation\" style=\"color:#393A34\">###</span><span class=\"token title important\"> II. Minimal Dependencies</span><span class=\"token plain\"></span><br></div><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\">Prefer the Python standard library. A third-party dependency MAY be added only when it</span><br></div><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\">removes clearly more complexity than it introduces, and MUST be recorded in the project's</span><br></div><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\">dependency file (e.g. </span><span class=\"token code-snippet code keyword\" style=\"color:#00009f\">`requirements.txt`</span><span class=\"token plain\"> or </span><span class=\"token code-snippet code keyword\" style=\"color:#00009f\">`pyproject.toml`</span><span class=\"token plain\">).</span><br></div></code></pre></div></div>\n<p>A dependency policy, written in the MUST/MAY language of a formal standard, from a prompt where\nI said I had no constraints. It isn't in\nthe local template or the skill file — but Article I of GitHub's nine constitutional articles\ndoes ask for implementations <em>\"with clear boundaries and <strong>minimal dependencies</strong>.\"</em> So it's\nconsistent with the published philosophy rather than invented on the spot. I can't tell you the\nmechanism, only what went in and what came out.</p>\n<p>One run, one trivial task, and it cost me nothing because nothing was at stake. Punnen's case\nshows what this kind of bias costs when the problem is hard enough for it to be wrong. Mine only\nshows it turns up unasked — which matters because, as Böckeler notes, Spec Kit's constitution is\nits <strong>memory bank</strong>, <em>\"a very powerful rules file\"</em> applied to every change. An unrequested\npreference doesn't sit in a document you'll discard. It becomes a standing rule.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"both-of-them-reached-for-the-1990s\">Both of them reached for the 1990s<a href=\"https://development-wec.wiline.com/docs/news/spec-driven-development-solution-to-vibe-coding/#both-of-them-reached-for-the-1990s\" class=\"hash-link\" aria-label=\"Direct link to Both of them reached for the 1990s\" title=\"Direct link to Both of them reached for the 1990s\" translate=\"no\">​</a></h2>\n<p>Here's what convinced me this is worth taking seriously rather than dismissing or evangelising.</p>\n<p>Böckeler, who worked on model-driven development early in her career, sees MDD:</p>\n<blockquote>\n<p>I wonder if spec-as-source, and even spec-anchoring, might end up with the downsides of both\nMDD and LLMs: Inflexibility and non-determinism.</p>\n</blockquote>\n<p>Punnen, twenty years in telecom, reaches independently for Rational Rose and UML — <em>\"treated as\nthe silver bullet of its decade: draw boxes, arrows, and diagrams, and the tool would magically\nturn them into working code.\"</em></p>\n<p>Neither cites the other. Different tools, different problems, different countries. Both land on\nthe same era: the last time our industry believed a document could be the source and code the\noutput.</p>\n<p>Böckeler is careful about the comparison — <em>\"I'm not nostalgic about my MDD experience.\"</em> Her\npoint is that today's tools drop the parts that made MDD painful: you no longer need a special\nspec language or a purpose-built generator. What she wonders is whether the exchange is a good\none, since the old approach at least produced the same output every time.</p>\n<p>Punnen names the trap underneath:</p>\n<blockquote>\n<p>The specification has to be very rigorous in the first place, but to become rigorous it needs\nto be iteratively refined alongside the generated code.</p>\n</blockquote>\n<p>A Catch-22, in other words: on his account you can't write a rigorous spec for a problem you\ndon't yet understand, and understanding arrives while building. He traces the thought back to\nFred Brooks, forty years ago: <em>\"descriptions of a software entity that abstract away its\ncomplexity often abstract away its essence.\"</em></p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"where-they-disagree--and-why-it-matters-to-you\">Where they disagree — and why it matters to you<a href=\"https://development-wec.wiline.com/docs/news/spec-driven-development-solution-to-vibe-coding/#where-they-disagree--and-why-it-matters-to-you\" class=\"hash-link\" aria-label=\"Direct link to Where they disagree — and why it matters to you\" title=\"Direct link to Where they disagree — and why it matters to you\" translate=\"no\">​</a></h2>\n<p>On one question these two are in direct opposition.</p>\n<p>Böckeler, generating code repeatedly from one Tessl spec: <em>\"I have seen the non-determinism in\naction… an interesting exercise to iterate on the spec and make it more and more specific to\nincrease the repeatability of the code generation.\"</em></p>\n<p>Punnen: <em>\"This is not a major problem in practice. SDD frameworks act as structured prompts, and\nmodern models produce highly consistent outputs when guided by them.\"</em></p>\n<p>GitHub takes Punnen's side and goes further, claiming <em>\"Consistency Across LLMs: Different AI\nmodels produce architecturally compatible code.\"</em> Not the same model twice — different models.\nNo evidence offered.</p>\n<p>Two experienced engineers, opposite conclusions, a vendor claim stronger than either, and not a\nnumber between them. It matters more than it looks. Keeping a spec and its code in step — the\nspec-anchored idea — relies on automated tests to catch the drift. That works for deterministic\ncode, where the same input gives the same output.</p>\n<p>Point it at an LLM feature and the check stops working. <code>assert response == expected</code> means\nnothing when the response differs every run. Unless you swap tests for <strong>evals</strong>, where each\nacceptance criterion becomes a scored assertion: did it classify correctly, did it return valid\nJSON against the schema, did it refuse when it should have.</p>\n<p>That bridge survives nondeterminism, and it's the harness I've spent several tutorials building\non the WEC Inference API. It also makes the disagreement <em>measurable</em>: write the spec, turn its\ncriteria into eval assertions, and run them across a long session to see whether adherence holds\nor decays.</p>\n<p>That's the next post.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"if-youre-going-to-try-it\">If you're going to try it<a href=\"https://development-wec.wiline.com/docs/news/spec-driven-development-solution-to-vibe-coding/#if-youre-going-to-try-it\" class=\"hash-link\" aria-label=\"Direct link to If you're going to try it\" title=\"Direct link to If you're going to try it\" translate=\"no\">​</a></h2>\n<p>Punnen's practical takeaways are better than anything I'd invent, and he earned them:</p>\n<ul>\n<li class=\"\"><strong>Treat \"no new deps\" rules as biases, not neutrals.</strong> If the right answer needs a dependency,\nyou'll have to defend it — the framework won't.</li>\n<li class=\"\"><strong>Treat generated success criteria as guesses</strong> until an engineer ratifies them. The agent will\nquote them back at you later as if they were measured.</li>\n<li class=\"\"><strong>Read every clarify menu as a design proposal in disguise.</strong> If the option you want isn't\nlisted, that's the failure — not a prompt to pick the best of three.</li>\n<li class=\"\"><strong>Push back during clarify, not during plan.</strong> Plans are long, internally consistent, and\nexhausting to revise.</li>\n</ul>\n<p>One of my own: <strong>write your principles with a date and a rationale, so a later you can supersede\nthem.</strong> IBM suggests treating specs as <em>\"stackable versioned artifacts, like architecture\ndecision records\"</em> — and an ADR is something you can mark superseded. The methodology calls\nthese principles immutable. Your problem isn't.</p>\n<p>Add IBM's cost test, the sharpest practical line any of them wrote: <em>the cost of refining the\nspec should always be lower than the cost of fixing misunderstandings in implementation. When\nthat balance flips, stop polishing and start building.</em></p>\n<p>And check the org before you install. The repo doing the rounds on LinkedIn is a <strong>fork</strong> with\neleven stars; the real project is <a href=\"https://github.com/github/spec-kit\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\"><code>github/spec-kit</code></a>. That\nmatters more here than for most tools: Spec Kit's job is writing instruction files your agent\nthen obeys, and on first run you approve that folder in one keystroke without having read them.</p>\n<hr>\n<p>So does it replace vibe coding? Not the way the pitch suggests. It doesn't help when you don't\nunderstand your problem — it helps when you understand it and communicate it badly. Those are\ndifferent failures, and only one of them has a template.</p>\n<p>Both found real value in the early phases — Punnen rates the clarification step his most useful\nof all — and both found the artifacts looking most authoritative exactly where the decisions\nunderneath them were most arbitrary. Punnen: <em>\"they are at their most dangerous when they look\nthe most rigorous.\"</em></p>\n<p>Böckeler reaches for a German compound word for it — <strong>Verschlimmbesserung</strong>. Making something\nworse in the attempt of making it better.</p>\n<p>The hard part was never writing the code.</p>",
            "url": "https://development-wec.wiline.com/docs/news/spec-driven-development-solution-to-vibe-coding/",
            "title": "Spec-Driven Development: is it the solution to Vibe Coding?",
            "summary": "A GitHub toolkit with 127k stars says you should write the spec before the code, and let the agent build from it. I ran it, then read the two engineers who tested it properly — and both reached for the same comparison: the last time our industry tried generating code from documents.",
            "date_modified": "2026-08-14T00:00:00.000Z",
            "author": {
                "name": "Rafael Fernandes",
                "url": "https://www.linkedin.com/in/rafaelmacariofernandes/"
            },
            "tags": [
                "ai-news",
                "engineering-practice",
                "agents",
                "evals",
                "tooling"
            ]
        },
        {
            "id": "https://development-wec.wiline.com/docs/news/muse-glimmer-30b-local-agentic-model/",
            "content_html": "<div class=\"newsHero\"><div class=\"newsHero__glow\" aria-hidden=\"true\"></div><span class=\"newsHero__eyebrow\">Models · AI News</span><h2 class=\"newsHero__title\">A serious agent, no data center required</h2><div class=\"newsHero__transition\"><span class=\"newsHero__pill newsHero__pill--from\">Cloud-only agents</span><svg xmlns=\"http://www.w3.org/2000/svg\" width=\"20\" height=\"20\" viewBox=\"0 0 24 24\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"2.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\" class=\"lucide lucide-arrow-right newsHero__arrow\" aria-hidden=\"true\"><path d=\"M5 12h14\"></path><path d=\"m12 5 7 7-7 7\"></path></svg><span class=\"newsHero__pill newsHero__pill--to\">One GPU, fully local</span></div></div>\n<p>Most \"run it locally\" model announcements come with an asterisk — smaller, weaker, a toy\nversion of the real thing. Meta's newest release doesn't: <strong>Muse Glimmer</strong>, a 30B\nmultimodal model built specifically for agentic work, fits on a single consumer GPU and\nbeats larger models on the benchmarks that actually measure agent behavior.</p>\n<!-- -->\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"architecture\">Architecture<a href=\"https://development-wec.wiline.com/docs/news/muse-glimmer-30b-local-agentic-model/#architecture\" class=\"hash-link\" aria-label=\"Direct link to Architecture\" title=\"Direct link to Architecture\" translate=\"no\">​</a></h2>\n<p>Muse Glimmer is <strong>30B parameters total</strong>: a 2B ViT-style vision encoder bolted onto a\n28B-parameter text decoder, 52 transformer layers using a hybrid attention pattern.\nIt's distilled from Meta's larger Muse Spark model — the capability of a bigger model,\ncompressed into something a single GPU can hold. Training data spans 100+ languages,\nwith a January 4, 2026 knowledge cutoff.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"multimodal-understanding\">Multimodal understanding<a href=\"https://development-wec.wiline.com/docs/news/muse-glimmer-30b-local-agentic-model/#multimodal-understanding\" class=\"hash-link\" aria-label=\"Direct link to Multimodal understanding\" title=\"Direct link to Multimodal understanding\" translate=\"no\">​</a></h2>\n<ul>\n<li class=\"\"><strong>Text, image, and video</strong> — video comprehension up to 96 frames at 2 fps (no audio\ntrack processed).</li>\n<li class=\"\"><strong>Open-ended object detection</strong> — it can locate and identify objects in a scene without\na predefined label set, rather than only recognizing a fixed category list.</li>\n<li class=\"\"><strong>131K+ token context window</strong>, long enough for extended agent sessions or large\ndocuments without external chunking.</li>\n</ul>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"agentic-tool-use\">Agentic tool use<a href=\"https://development-wec.wiline.com/docs/news/muse-glimmer-30b-local-agentic-model/#agentic-tool-use\" class=\"hash-link\" aria-label=\"Direct link to Agentic tool use\" title=\"Direct link to Agentic tool use\" translate=\"no\">​</a></h2>\n<p>The headline feature: <strong>multimodal tool-calling with structured outputs</strong>. It can look at\nan image and decide which function to call based on what it sees, not just parse text\ninstructions — e.g. inspecting a screenshot and calling the right API based on what's\nrendered, not a text description of it. It also generates and executes code, and is\nbuilt with explicit failure-recovery behavior rather than assuming every tool call\nsucceeds on the first try.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"running-it-locally\">Running it locally<a href=\"https://development-wec.wiline.com/docs/news/muse-glimmer-30b-local-agentic-model/#running-it-locally\" class=\"hash-link\" aria-label=\"Direct link to Running it locally\" title=\"Direct link to Running it locally\" translate=\"no\">​</a></h2>\n<ul>\n<li class=\"\"><strong>Fits on one consumer GPU</strong> — quantized variants run in 24–32GB of VRAM.</li>\n<li class=\"\"><strong>Day-0 support</strong> in <code>transformers</code>, <code>llama.cpp</code>, <code>vLLM</code>, and Inference Endpoints — no\nwaiting on community ports.</li>\n<li class=\"\"><strong>DFlash speculative decoding</strong> — up to 3x faster generation on supported hardware.</li>\n<li class=\"\">No cloud round-trip required for any of the above; everything runs on the box you own.</li>\n</ul>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"licensing\">Licensing<a href=\"https://development-wec.wiline.com/docs/news/muse-glimmer-30b-local-agentic-model/#licensing\" class=\"hash-link\" aria-label=\"Direct link to Licensing\" title=\"Direct link to Licensing\" translate=\"no\">​</a></h2>\n<p><strong>Apache 2.0</strong> — no commercial-use gate, no attribution requirement, no separate license\nnegotiation to run it in a product. Worth stating plainly because it's not a given right\nnow: some recent open-weight agentic models carry commercial-use restrictions that only\nsurface once you read the license text closely. This one doesn't.</p>\n<p>Meta also ran safety evaluations for chemical/biological, cybersecurity, and\nloss-of-control risk — all rated \"moderate or lower.\"</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"how-it-compares\">How it compares<a href=\"https://development-wec.wiline.com/docs/news/muse-glimmer-30b-local-agentic-model/#how-it-compares\" class=\"hash-link\" aria-label=\"Direct link to How it compares\" title=\"Direct link to How it compares\" translate=\"no\">​</a></h2>\n<p>Meta's own published numbers, against Gemma4-31B and Qwen3.6-27B:</p>\n<table><thead><tr><th>Benchmark</th><th>Muse Glimmer</th><th>Gemma4</th><th>Qwen3.6</th></tr></thead><tbody><tr><td>MCP Atlas (agentic)</td><td><strong>75.5</strong></td><td>54.2</td><td>62.5</td></tr><tr><td>DeepSearch QA</td><td><strong>74.6</strong></td><td>61.7</td><td>71.1</td></tr><tr><td>GAIA2</td><td><strong>43.3</strong></td><td>36.4</td><td>40.0</td></tr><tr><td>SWE-Bench Pro</td><td><strong>51.2</strong></td><td>36.9</td><td>50.2</td></tr><tr><td>SWE-Bench Verified</td><td>76.0</td><td>66.6</td><td><strong>77.2</strong></td></tr></tbody></table>\n<p>It's not a clean sweep — Qwen3.6 edges it on SWE-Bench Verified — but against Gemma4 the\ngap is wide, and against Qwen3.6 it's competitive or ahead on most agentic-specific\nbenchmarks. <em>(Numbers from\n<a href=\"https://huggingface.co/blog/muse-glimmer\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">Meta's Muse Glimmer announcement</a>;\nbenchmarks are directional, not gospel.)</em></p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"why-it-matters\">Why it matters<a href=\"https://development-wec.wiline.com/docs/news/muse-glimmer-30b-local-agentic-model/#why-it-matters\" class=\"hash-link\" aria-label=\"Direct link to Why it matters\" title=\"Direct link to Why it matters\" translate=\"no\">​</a></h2>\n<p>The self-hosted AI story has always had a quiet tax: the good agentic models needed real\ninfrastructure, so \"run it yourself\" often meant \"run a worse version of it yourself.\" A\nmodel that's genuinely built agent-first, ships permissively licensed, and fits on\nhardware a single person can own is the gap closing in real time — the same trend this\nwhole tutorial series has been betting on.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"is-it-on-wec\">Is it on WEC?<a href=\"https://development-wec.wiline.com/docs/news/muse-glimmer-30b-local-agentic-model/#is-it-on-wec\" class=\"hash-link\" aria-label=\"Direct link to Is it on WEC?\" title=\"Direct link to Is it on WEC?\" translate=\"no\">​</a></h2>\n<p>Not yet — and we're not going to pretend otherwise. We're running it through our own\nmodel-evaluation suite now, the same one that benchmarks everything already on WEC\nModels, before it earns a place in the catalog. If it holds up against what's already\nthere, expect it soon.</p>\n<hr>\n<p>📖 <strong>Sources:</strong> <a href=\"https://huggingface.co/blog/muse-glimmer\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">Meta's Muse Glimmer announcement (Hugging Face)</a> · <a href=\"https://huggingface.co/meta-models/Muse-Glimmer-30B-GGUF\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">Model card</a></p>",
            "url": "https://development-wec.wiline.com/docs/news/muse-glimmer-30b-local-agentic-model/",
            "title": "Muse Glimmer: a 30B agentic model that runs on one GPU, no data center required",
            "summary": "Meta released a 30B multimodal model built for local agent workloads — beats Gemma4 and holds its own against Qwen3.6 on agentic benchmarks, and fits on a single consumer GPU. Here's what's actually new, and what it'd take for it to land on WEC.",
            "date_modified": "2026-08-10T00:00:00.000Z",
            "author": {
                "name": "Rafael Fernandes",
                "url": "https://www.linkedin.com/in/rafaelmacariofernandes/"
            },
            "tags": [
                "ai-news",
                "models",
                "open-weight",
                "agents",
                "local-inference"
            ]
        },
        {
            "id": "https://development-wec.wiline.com/docs/news/loop-engineering-graph-engineering-what-survives/",
            "content_html": "<div class=\"newsHero\"><div class=\"newsHero__glow\" aria-hidden=\"true\"></div><span class=\"newsHero__eyebrow\">Architecture · AI News</span><h2 class=\"newsHero__title\">Loops, graphs, and the six-week obituary</h2><div class=\"newsHero__transition\"><span class=\"newsHero__pill newsHero__pill--from\">Loop engineering</span><svg xmlns=\"http://www.w3.org/2000/svg\" width=\"20\" height=\"20\" viewBox=\"0 0 24 24\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"2.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\" class=\"lucide lucide-arrow-right newsHero__arrow\" aria-hidden=\"true\"><path d=\"M5 12h14\"></path><path d=\"m12 5 7 7-7 7\"></path></svg><span class=\"newsHero__pill newsHero__pill--to\">Graph engineering</span></div></div>\n<p>In June 2026, the AI world got a new buzzword: <strong>loop engineering</strong> — roughly, disciplined\ndesign of a single agent's tool-calling loop. Six weeks later it was supposedly replaced by\n<strong>graph engineering</strong> — wiring up several agents at once — killed by twelve words that 3.1\nmillion people saw:</p>\n<div class=\"xEmbed\"><blockquote class=\"twitter-tweet\" data-dnt=\"true\" data-conversation=\"none\"><p lang=\"en\" dir=\"ltr\"></p><p>Are we still talking loops or did we shift to graphs yet?</p><p></p>— <!-- -->Peter Steinberger 🦞<!-- --> (@<!-- -->steipete<!-- -->) <a href=\"https://twitter.com/steipete/status/2078277297791189132\">July 18, 2026</a></blockquote></div>\n<p>Don't know what loop engineering is? Don't worry — neither did most of the people\ndeclaring it dead. Here are both terms, how a name became an obituary in six weeks, and\nthe twist nobody checked before writing about it.</p>\n<!-- -->\n<p>The split shows up best on a job with <strong>independent work that can run at once</strong> and where a\nwrong answer is expensive — a <strong>research briefing</strong>: <em>\"Pull together what we know about X —\nthe open web, our own docs, and the repo — and give me a one-page brief with sources.\"</em></p>\n<p>The <a class=\"\" href=\"https://development-wec.wiline.com/docs/tutorials/deploy-openclaw-docker-compose/\">OpenClaw</a> agent you stand up in these\ntutorials works the <strong>loop</strong> way — one agent cycling through its tools, one turn at a time.\nGraph engineering wires up a team of specialists instead. Same job, two shapes:</p>\n<div class=\"newsThemedWrap newsThemedWrap--light\"><p><span class=\"zoomImage__wrap\"><img alt=\"Loop: one agent cycling through its tools in sequence. Graph: many specialist agents wired into a network, each owning one tool.\" src=\"https://development-wec.wiline.com/docs/assets/images/loop-vs-graph-light-d1fd8017587aa7c27d81ad36a7ec1f0b.png\" width=\"2564\" height=\"1751\" class=\"zoomImage \" loading=\"lazy\"><span class=\"zoomImage__badge\" aria-hidden=\"true\"><svg viewBox=\"0 0 24 24\" width=\"16\" height=\"16\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"2\" stroke-linecap=\"round\"><circle cx=\"11\" cy=\"11\" r=\"7\"></circle><path d=\"M21 21l-4.3-4.3\"></path><path d=\"M11 8v6M8 11h6\"></path></svg></span></span></p></div>\n<div class=\"newsThemedWrap newsThemedWrap--dark\"><p><span class=\"zoomImage__wrap\"><img alt=\"Loop: one agent cycling through its tools in sequence. Graph: many specialist agents wired into a network, each owning one tool.\" src=\"https://development-wec.wiline.com/docs/assets/images/loop-vs-graph-dark-f9ef88bb5be59e4d6a1396b3985067fd.png\" width=\"2564\" height=\"1751\" class=\"zoomImage \" loading=\"lazy\"><span class=\"zoomImage__badge\" aria-hidden=\"true\"><svg viewBox=\"0 0 24 24\" width=\"16\" height=\"16\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"2\" stroke-linecap=\"round\"><circle cx=\"11\" cy=\"11\" r=\"7\"></circle><path d=\"M21 21l-4.3-4.3\"></path><path d=\"M11 8v6M8 11h6\"></path></svg></span></span></p></div>\n<p>Read the <strong>loop</strong> on the left as a wheel: one agent visits each tool, then comes back\naround. Push it past simple tasks and two cracks appear: it's <strong>slow</strong> (three independent\nsearches still run back-to-back, one worker at a time), and it's <strong>hard to trust</strong> (the same\nagent searched, read, and wrote the brief, so a wrong claim has no address — you can't tell\nif the search was thin or the agent invented it).</p>\n<p>The <strong>graph</strong> on the right fixes both: specialists split the job, a planner routes it. The\nsearches now fire <strong>at the same time</strong> — the brief lands in the time of the slowest one, not\nthe sum — and each agent owns one source, so a bad claim has an address. A dedicated\n<strong>fact-check</strong> node, the step a rushed loop skips, becomes a real gate.</p>\n<p>None of that is free: one prompt becomes a planner, four agents, and a verifier — a new\nfailure mode where the <em>coordination itself</em> can break. Trust problem traded for a plumbing\nproblem. And every box in the graph still runs its own loop inside: <strong>a graph contains\nloops</strong> — that's the whole point.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"both-terms-in-one-picture\">Both terms, in one picture<a href=\"https://development-wec.wiline.com/docs/news/loop-engineering-graph-engineering-what-survives/#both-terms-in-one-picture\" class=\"hash-link\" aria-label=\"Direct link to Both terms, in one picture\" title=\"Direct link to Both terms, in one picture\" translate=\"no\">​</a></h2>\n<!-- -->\n<p><strong>Loop engineering</strong> is about one agent. An AI agent is just a model in a <code>while</code> loop\nwith tools: give it a goal, it picks a tool, your code runs it, the result goes back in,\nrepeat. ChatGPT answering a question is one call. An agent that edits a file, runs the\ntests, sees them fail and tries again — that's the loop, five times over. Loop engineering\nis designing that cycle deliberately: what triggers it, who verifies the work, when it\nstops, what happens on failure. The slogan is a good one — <em>the intelligence lives in the\nmodel, the reliability lives in the loop.</em></p>\n<p><strong>Graph engineering</strong> is about several agents. Nodes are agents, edges are dependencies,\nplus the shared state and permissions between them. If loop engineering is \"make one agent\nreliable,\" graph engineering is \"make ten agents not step on each other.\"</p>\n<p>Different problems, different scale. So <em>\"loop engineering is dead, graph engineering\nreplaced it\"</em> was never a sensible sentence — it's like saying wheels replaced cars.</p>\n<p>The name for the first one arrived in June 2026, from\n<a href=\"https://x.com/addyosmani/status/2064127981161959567\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">Addy Osmani</a>. It caught on fast\nbecause everyone was already doing it, badly, with no shared vocabulary. Remember that\ndate.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"the-textbook-settles-it\">The textbook settles it<a href=\"https://development-wec.wiline.com/docs/news/loop-engineering-graph-engineering-what-survives/#the-textbook-settles-it\" class=\"hash-link\" aria-label=\"Direct link to The textbook settles it\" title=\"Direct link to The textbook settles it\" translate=\"no\">​</a></h2>\n<p>Here's the thing nobody checked before writing a guide: \"graph\" isn't a 2026 coinage, and\nthe receipts are sitting in a free MIT textbook. From <em>Mathematics for Computer Science</em>\n(6.042J), chapter 6, page 189 — <strong>Definition 6.1.1</strong>:</p>\n<blockquote>\n<p>\"A directed graph G = (V, E) consists of a nonempty set of nodes V and a set of directed\nedges E… A directed graph is <strong>simple</strong> if it has no <strong>loops</strong> (that is, edges of the\nform u→u) and no multiple edges.\"</p>\n</blockquote>\n<p>The formal definition of a graph <em>already contains the word loop</em>, as an ordinary feature\nof graphs — not their opposite. Definition 6.1.2 adds that a <strong>cycle</strong> is a walk that\nreturns to where it started, which is what a loop is.</p>\n<p><strong>A loop isn't the opposite of a graph. It's a graph that comes back to an earlier node.</strong>\nMathematics has known this since Euler walked around Königsberg in 1736.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"what-actually-happened-in-july\">What actually happened in July<a href=\"https://development-wec.wiline.com/docs/news/loop-engineering-graph-engineering-what-survives/#what-actually-happened-in-july\" class=\"hash-link\" aria-label=\"Direct link to What actually happened in July\" title=\"Direct link to What actually happened in July\" translate=\"no\">​</a></h2>\n<p>The fuse wasn't an engineering discovery — it was <strong>two product launches colliding.</strong>\nDeepLearning.AI released a course on knowledge graphs with multi-agent systems, taught by\nNeo4j's Andreas Kollegger. Around the same time, Linear shipped an agent feature called —\nof all things — <strong>Loops</strong>. Two companies, two unrelated products, two words that sounded\nlike rival philosophies.</p>\n<p>Then Steinberger, who built <a class=\"\" href=\"https://development-wec.wiline.com/docs/tutorials/deploy-openclaw-docker-compose/\">OpenClaw</a>, posted\nhis twelve words at 9:34 PM on July 17. <strong>Four and a half hours later</strong> Hamel Husain\npublished an X Article titled <em>\"Loop Engineering Is Dead. Enter Graph Engineering,\"</em> and\n<a href=\"https://x.com/svpino/status/2078516761318584774\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">Santiago Valdarrama</a> picked it up the\nsame day. Within 48 hours the new term had three competing definitions and a wave of\ncopycat posts.</p>\n<h3 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"the-part-that-should-have-ended-it\">The part that should have ended it<a href=\"https://development-wec.wiline.com/docs/news/loop-engineering-graph-engineering-what-survives/#the-part-that-should-have-ended-it\" class=\"hash-link\" aria-label=\"Direct link to The part that should have ended it\" title=\"Direct link to The part that should have ended it\" translate=\"no\">​</a></h3>\n<p><strong>Neither founding post was serious</strong> — the writers who tracked the cycle say so outright.\nLouis-François Bouchard put it plainly: <em>\"my whole feed decided we have a new discipline.\nTo be honest, both tweets were jokes.\"</em> Husain's own follow-up only widened the wink,\npromising the piece was <em>\"not what you think it is.\"</em></p>\n<p>And almost nobody could check: the article sat behind X's Premium paywall. A paradigm's\nfounding text was something few of the people citing it had actually read. What spread\nwasn't an argument — it was a shape: a punchy title, a wink right behind it, and an industry\nthat answered a joke by writing documentation for it.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"the-replies-were-smarter-than-the-announcements\">The replies were smarter than the announcements<a href=\"https://development-wec.wiline.com/docs/news/loop-engineering-graph-engineering-what-survives/#the-replies-were-smarter-than-the-announcements\" class=\"hash-link\" aria-label=\"Direct link to The replies were smarter than the announcements\" title=\"Direct link to The replies were smarter than the announcements\" translate=\"no\">​</a></h2>\n<p>The sharpest response came from outside the agent crowd. <strong>David Khourshid</strong> created\n<strong>XState</strong> — a widely-used library for modeling exactly this kind of state machine in\ncode — and has been modelling these structures for a decade:</p>\n<div class=\"xEmbed\"><blockquote class=\"twitter-tweet\" data-dnt=\"true\" data-conversation=\"none\"><p lang=\"en\" dir=\"ltr\"></p><p>First it was loops. Now it's graphs. Next month it'll be something else. Here's the thing: we're constantly rediscovering decades-old software engineering patterns and repackaging them as innovations or whatever, instead of just applying what we've already known for a long time.</p><p></p>— <!-- -->David Khourshid<!-- --> (@<!-- -->DavidKPiano<!-- -->) <a href=\"https://twitter.com/DavidKPiano/status/2079209887158989231\">July 20, 2026</a></blockquote></div>\n<p>His explanation takes two minutes. A state machine answers one question — <em>given the\ncurrent state, when an event occurs, what is the next state?</em> Draw it and states become\nnodes, transitions become edges. Then he turns it on the July argument:</p>\n<blockquote>\n<p>\"Surprise… loops are graphs: directed, cyclic ones.\"</p>\n</blockquote>\n<p>A loop, he notes, is the most basic state machine there is: two states, <code>looping</code> and\n<code>done</code>. He posted a diagram of it captioned <strong>\"This is the silly thing you all hyped for\nweeks.\"</strong></p>\n<!-- -->\n<p>That's the object six weeks of discourse was about. And real agents — the ones\nthat retry, backtrack, wait for a human — were never sequential and never acyclic to begin\nwith (<em>\"sorry, DAG lovers\"</em>).</p>\n<p>The creator of LangChain got there from the opposite direction, his LangGraph being the\nimplementation everyone kept pointing at:</p>\n<div class=\"xEmbed\"><blockquote class=\"twitter-tweet\" data-dnt=\"true\" data-conversation=\"none\"><p lang=\"en\" dir=\"ltr\"></p><p>So i didn't really know what graph engineering is, and i still don't really… but it's basically just langgraph?</p><p></p>— <!-- -->Harrison Chase<!-- --> (@<!-- -->hwchase17<!-- -->) <a href=\"https://twitter.com/hwchase17/status/2079219804951683380\">July 20, 2026</a></blockquote></div>\n<p>Four days later he and Sydney Runkle published a longer answer,\n<a href=\"https://www.langchain.com/blog/3-years-of-graph-engineering-with-langgraph\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\"><em>3 Years of Graph Engineering with LangGraph</em></a>,\nwhich is the most useful thing written during the whole episode. Their opening line is\nthe best description of the phenomenon I've read:</p>\n<blockquote>\n<p>\"It's the latest term to come out of X's AI content factory, joining prompt engineering,\ncontext engineering, harness engineering, and loop engineering.\"</p>\n</blockquote>\n<p>And then, from the people whose product is literally the graph:</p>\n<blockquote>\n<p>\"Loops are simple graphs. Loop engineering isn't an alternative to graphs, so much as a\nsimple version of them.\"</p>\n</blockquote>\n<p>They also confirm the thing most July posts got backwards: <strong>production agent graphs are\nusually not DAGs.</strong> Real agents retry, ask for missing information, revise after\nvalidation, pause for a human. Cycles aren't a design flaw to be engineered out — they're\nthe job.</p>\n<p>The clearest framing of all, though, came from a reply:</p>\n<div class=\"xEmbed\"><blockquote class=\"twitter-tweet\" data-dnt=\"true\" data-conversation=\"none\"><p lang=\"en\" dir=\"ltr\"></p><p>Graph engineering is deciding where the work is allowed to go. Loop engineering is making the work get better each time it runs. Graph is the rails. Loop is the motor. Rails keep you from crashing. The motor is what actually moves.</p><p></p>— <!-- -->Eric Osiu<!-- --> (@<!-- -->ericosiu<!-- -->) <a href=\"https://twitter.com/ericosiu/status/2079991948106957131\">July 22, 2026</a></blockquote></div>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"where-it-stands-today\">Where it stands today<a href=\"https://development-wec.wiline.com/docs/news/loop-engineering-graph-engineering-what-survives/#where-it-stands-today\" class=\"hash-link\" aria-label=\"Direct link to Where it stands today\" title=\"Direct link to Where it stands today\" translate=\"no\">​</a></h2>\n<p>It's August 5, about two and a half weeks after the obituary. Nobody won. The debate\ndidn't resolve and it didn't die — <strong>it got absorbed by content marketing.</strong> Every\nagent-infrastructure vendor now has a \"definitive guide to graph engineering,\" and behind\nthem a long tail of SEO pages saying the same thing in the same order.</p>\n<figure class=\"stageFlow\"><div class=\"stageFlow__track\"><div class=\"stageFlow__card\" style=\"background:rgba(var(--primary-rgb), 0.050);border-color:rgba(var(--primary-rgb), 0.250)\"><span class=\"stageFlow__stage\">June 7</span><span class=\"stageFlow__title\">A name appears</span><span class=\"stageFlow__tag\">it describes something real</span></div><svg xmlns=\"http://www.w3.org/2000/svg\" width=\"22\" height=\"22\" viewBox=\"0 0 24 24\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"2.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\" class=\"lucide lucide-arrow-right stageFlow__arrow\" aria-hidden=\"true\"><path d=\"M5 12h14\"></path><path d=\"m12 5 7 7-7 7\"></path></svg><div class=\"stageFlow__card\" style=\"background:rgba(var(--primary-rgb), 0.123);border-color:rgba(var(--primary-rgb), 0.383)\"><span class=\"stageFlow__stage\">Weeks 2–5</span><span class=\"stageFlow__title\">Everyone adopts it</span><span class=\"stageFlow__tag\">meaning drifts</span></div><svg xmlns=\"http://www.w3.org/2000/svg\" width=\"22\" height=\"22\" viewBox=\"0 0 24 24\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"2.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\" class=\"lucide lucide-arrow-right stageFlow__arrow\" aria-hidden=\"true\"><path d=\"M5 12h14\"></path><path d=\"m12 5 7 7-7 7\"></path></svg><div class=\"stageFlow__card\" style=\"background:rgba(var(--primary-rgb), 0.197);border-color:rgba(var(--primary-rgb), 0.517)\"><span class=\"stageFlow__stage\">July 17</span><span class=\"stageFlow__title\">Declared dead</span><span class=\"stageFlow__tag\">apparently as a joke</span></div><svg xmlns=\"http://www.w3.org/2000/svg\" width=\"22\" height=\"22\" viewBox=\"0 0 24 24\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"2.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\" class=\"lucide lucide-arrow-right stageFlow__arrow\" aria-hidden=\"true\"><path d=\"M5 12h14\"></path><path d=\"m12 5 7 7-7 7\"></path></svg><div class=\"stageFlow__card\" style=\"background:rgba(var(--primary-rgb), 0.270);border-color:rgba(var(--primary-rgb), 0.650)\"><span class=\"stageFlow__stage\">August</span><span class=\"stageFlow__title\">Vendors publish guides</span><span class=\"stageFlow__tag\">the term becomes a funnel</span></div></div><figcaption class=\"stageFlow__caption\">Naming to obituary to lead generation, in six weeks. No benchmark ran at any point.</figcaption></figure>\n<p>Nothing was <em>learned</em> in that time. No benchmark ran, no production system proved one\napproach beat the other. Two companies shipped unrelated products, two well-followed\ndevelopers made a joke, and a lot of people agreed on a word — then agreed on a different\nword.</p>\n<p>That has a real cost. If you tried to keep up by adopting each term as it trended, you\nrewrote your architecture twice in July for reasons that were social, not technical.\nMeanwhile the engineer who ignored both posts and spent that month adding stop rules and\ntyped state came out ahead, because those mattered under any label.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"the-failure-nobodys-selling-a-guide-for\">The failure nobody's selling a guide for<a href=\"https://development-wec.wiline.com/docs/news/loop-engineering-graph-engineering-what-survives/#the-failure-nobodys-selling-a-guide-for\" class=\"hash-link\" aria-label=\"Direct link to The failure nobody's selling a guide for\" title=\"Direct link to The failure nobody's selling a guide for\" translate=\"no\">​</a></h2>\n<p>One idea from this month deserves more attention than the naming war, and it comes from\n<a href=\"https://www.linkedin.com/pulse/what-graph-engineering-really-towards-artificial-intelligence-e79ic/\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">Towards AI's breakdown</a>:\n<strong>more agents doesn't mean more judgement.</strong></p>\n<p>Twenty agents running the same model, reading the same flawed context and checking against\nthe same broken metric will agree with each other at industrial scale. Worse is the\ncircular version: one agent checks a report against another report, an audit agent checks\nboth against a dashboard, and the dashboard was built from the same data. Every node\nagrees. Nothing touched reality. The system <em>looks</em> well governed, because the diagram has\nreviewers everywhere.</p>\n<p>The fix is what they call <strong>reality anchors</strong> — evidence from outside the agent system.\nTests that actually ran. Money that reached the bank. Customers who stayed. Rules the\noptimiser can't quietly rewrite. Their line is the one I'd put on the wall:</p>\n<blockquote>\n<p>\"Without anchors, a graph is a larger hallucination with better project management.\"</p>\n</blockquote>\n<p>The same trap exists one level down, inside a single loop: if the agent that does the work\nalso decides whether the work is good, you've built an expensive machine for agreeing with\nitself. A verifier only counts if it can actually fail you.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"what-to-actually-do\">What to actually do<a href=\"https://development-wec.wiline.com/docs/news/loop-engineering-graph-engineering-what-survives/#what-to-actually-do\" class=\"hash-link\" aria-label=\"Direct link to What to actually do\" title=\"Direct link to What to actually do\" translate=\"no\">​</a></h2>\n<p>Khourshid's line is the filter, and it isn't cynicism: <em>next month it'll be something\nelse.</em> When the next term lands, the question is never <em>\"is this the new paradigm?\"</em> It's\n<strong>\"what specific failure does this name, and do I have that failure yet?\"</strong></p>\n<p>Diagnose where your bottleneck actually is:</p>\n<ul>\n<li class=\"\"><strong>One agent degrading over a long session</strong> — forgetting, repeating tool calls, burning\ntokens? That's a <strong>loop</strong> problem: stop rules, what tool output re-enters context,\nwhether errors are surfaced or swallowed. No graph framework fixes any of it.</li>\n<li class=\"\"><strong>Several agents duplicating work or deadlocking on shared state?</strong> That's a\n<strong>coordination</strong> problem, and it needs an explicit control plane whatever you call it.</li>\n</ul>\n<p>And a case for neither: if the task is genuinely open-ended — research, exploration —\nforcing it into fixed paths is the wrong move. The LangChain team makes this point against\ntheir own product: they built early deep research on predefined LangGraph workflows and\nthen moved to a looser agentic loop, and GPT Researcher did the same. Structure you\nhaven't earned costs you.</p>\n<p>And whichever you have, these outlive the vocabulary: every loop needs a stop rule; state\nneeds a shape and save points; acting nodes need permission boundaries and a human gate on\nanything irreversible; add complexity only when a real failure asks for it.</p>\n<p>If you want the two concepts that pay off across all of it, take Khourshid's\nrecommendation over any of July's vocabulary: <strong>state machines and the actor model</strong>. Both\npredate this argument by decades and will outlive whatever replaces it — probably in about\nsix weeks.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"sources\">Sources<a href=\"https://development-wec.wiline.com/docs/news/loop-engineering-graph-engineering-what-survives/#sources\" class=\"hash-link\" aria-label=\"Direct link to Sources\" title=\"Direct link to Sources\" translate=\"no\">​</a></h2>\n<ul>\n<li class=\"\"><a href=\"https://x.com/addyosmani/status/2064127981161959567\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">Addy Osmani, naming \"loop engineering\" (June 2026) — X</a></li>\n<li class=\"\"><a href=\"https://x.com/steipete/status/2078277297791189132\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">Peter Steinberger's post — X</a></li>\n<li class=\"\">Hamel Husain, \"Loop Engineering Is Dead. Enter Graph Engineering\" — X Article, July 18,\n2026 (behind X Premium) · <a href=\"https://x.com/HamelHusain/status/2078348097697263855\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">his public follow-up</a></li>\n<li class=\"\"><a href=\"https://x.com/svpino/status/2078516761318584774\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">Santiago Valdarrama — X</a></li>\n<li class=\"\"><a href=\"https://x.com/DavidKPiano/status/2079209887158989231\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">David Khourshid (XState), \"State machines in 2 minutes\" — X</a></li>\n<li class=\"\"><a href=\"https://x.com/hwchase17/status/2079219804951683380\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">Harrison Chase (LangChain) — X</a> · <a href=\"https://x.com/ericosiu/status/2079991948106957131\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">Eric Osiu — X</a></li>\n<li class=\"\"><a href=\"https://www.louisbouchard.ai/graph-engineering-explained/\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">Louis-François Bouchard, \"Graph Engineering Explained: What Actually Changed\"</a></li>\n<li class=\"\"><a href=\"https://ai.gopubby.com/two-engineers-made-a-joke-about-graph-engineering-six-days-later-it-had-courses-3c082a5900fd\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">\"Two Engineers Made a Joke About Graph Engineering. Six Days Later It Had Courses.\" — AI Advances</a></li>\n<li class=\"\"><a href=\"https://smartscope.blog/en/blog/graph-engineering-loop-engineering-logic-review/\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">SmartScope, on whether the \"obituary\" is true</a></li>\n<li class=\"\"><a href=\"https://ocw.mit.edu/courses/6-042j-mathematics-for-computer-science-fall-2010/e6db7638031b754f5f68012946af4763_MIT6_042JF10_chap06.pdf\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">MIT 6.042J <em>Mathematics for Computer Science</em>, Ch. 6 \"Directed Graphs\" (free PDF)</a> — Definitions 6.1.1–6.1.2, pp. 189–191</li>\n</ul>",
            "url": "https://development-wec.wiline.com/docs/news/loop-engineering-graph-engineering-what-survives/",
            "title": "'Loop Engineering Is Dead' — and the Real Story Is Weirder Than the Obituary",
            "summary": "In June, 'loop engineering' got a name. Six weeks later it was declared dead — and the industry answered with vendor guides, competing definitions, and a wave of SEO. Here's what loop and graph engineering actually mean, why a free MIT textbook settles the argument on page 189, and what a six-week hype cycle should teach you about what to learn.",
            "date_modified": "2026-08-05T00:00:00.000Z",
            "author": {
                "name": "Rafael Fernandes",
                "url": "https://www.linkedin.com/in/rafaelmacariofernandes/"
            },
            "tags": [
                "ai-news",
                "agents",
                "architecture",
                "orchestration",
                "engineering-practice"
            ]
        },
        {
            "id": "https://development-wec.wiline.com/docs/news/mcp-2026-07-28-spec/",
            "content_html": "<figure class=\"newsHero newsHero--image\"><span class=\"newsHero__chip\">Protocols · AI News</span><img src=\"https://development-wec.wiline.com/docs/img/news/mcp-drops-sessions-retires-three-core-features-16x9.webp\" alt=\"Model Context Protocol — the 2026-07-28 specification\" loading=\"eager\"></figure>\n<p>The Model Context Protocol just had its biggest revision since launch. The\n2026-07-28 specification doesn't add a feature — it rewrites how every MCP\nserver talks to every client. If you've deployed an MCP server anywhere past\n\"runs on my laptop,\" this one touches your infrastructure, not just your\nchangelog.</p>\n<!-- -->\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"what-actually-shipped\">What actually shipped<a href=\"https://development-wec.wiline.com/docs/news/mcp-2026-07-28-spec/#what-actually-shipped\" class=\"hash-link\" aria-label=\"Direct link to What actually shipped\" title=\"Direct link to What actually shipped\" translate=\"no\">​</a></h2>\n<p>The headline change: <strong>MCP is now stateless at the protocol layer.</strong> Every\nrequest carries its own protocol version, client identity, and capabilities —\nthere's no more <code>initialize</code>/<code>initialized</code> handshake, no session ID, no\nrequirement that request N+1 lands on the same server instance that handled\nrequest N. Lead maintainer David Soria Parra called it\n<a href=\"https://blog.modelcontextprotocol.io/posts/2026-07-28/\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">\"a leap in serving scalable MCP servers\"</a>,\nbuilt on 18 months of running MCP past the local-tool stage.</p>\n<p>That one change unlocks a chain of practical ones:</p>\n<ul>\n<li class=\"\"><strong>Plain load balancing.</strong> A server that used to need sticky sessions and a\nshared session store can sit behind an ordinary round-robin balancer.</li>\n<li class=\"\"><strong>Header-based routing.</strong> Requests now carry <code>Mcp-Method</code> and <code>Mcp-Name</code>\nHTTP headers, so gateways and firewalls can route and rate-limit by\ninspecting headers instead of parsing every JSON body.</li>\n<li class=\"\"><strong>Cacheable list results.</strong> <code>tools/list</code>, <code>prompts/list</code>, <code>resources/list</code>,\nand <code>resources/read</code> now return <code>ttlMs</code> and <code>cacheScope</code>, so clients know\nhow long they're allowed to skip re-fetching.</li>\n<li class=\"\"><strong>Multi Round-Trip Requests (MRTR).</strong> The old server-initiated,\nheld-open-stream pattern for mid-call input (confirmations, missing\nparameters) is gone. A server now returns <code>resultType: \"input_required\"</code>;\nthe client retries the same call with <code>inputResponses</code> filled in — no\npersistent connection required.</li>\n</ul>\n<p>State didn't disappear, it just became explicit: if your tool genuinely needs\nit, it mints a handle and hands it back to the client to pass in on the next\ncall, instead of hiding it in the transport layer.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"authorization-got-a-real-rewrite-not-a-patch\">Authorization got a real rewrite, not a patch<a href=\"https://development-wec.wiline.com/docs/news/mcp-2026-07-28-spec/#authorization-got-a-real-rewrite-not-a-patch\" class=\"hash-link\" aria-label=\"Direct link to Authorization got a real rewrite, not a patch\" title=\"Direct link to Authorization got a real rewrite, not a patch\" translate=\"no\">​</a></h2>\n<p>This is the part that should get an AI engineer's attention before the\nprotocol change does. MCP authorization now aligns with OAuth 2.1 and OpenID\nConnect instead of leaving it to each implementer to wire up their own\nversion, <a href=\"https://workos.com/blog/mcp-2026-spec-agent-authentication\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">as WorkOS breaks down</a>:</p>\n<ul>\n<li class=\"\">Servers must implement <strong>OAuth 2.0 Protected Resource Metadata</strong> (RFC 9728)\nfor discovery and <strong>Resource Indicators</strong> (RFC 8707) so a token minted for\none MCP server can't be replayed against another.</li>\n<li class=\"\"><strong>Issuer verification is now mandatory</strong> — clients must check the <code>iss</code>\nparameter before redeeming a code, closing the authorization-server\nmix-up class of bugs.</li>\n<li class=\"\"><strong>Client ID Metadata Documents (CIMD)</strong> replace Dynamic Client Registration\nas the preferred path (DCR still works, for now).</li>\n</ul>\n<p>The practical upshot: the \"confused deputy\" problem — a tool call executing\nwith credentials meant for a different server — gets closed at the protocol\nlevel instead of depending on every server author to remember to check.\nIf you're running multiple MCP servers behind one gateway (which, per our own\n<a class=\"\" href=\"https://development-wec.wiline.com/docs/news/litellm-rust-gateway/\">gateway coverage</a>, is where this is heading for\neveryone), this is the part of the update that actually reduces your risk\nsurface, not just your ops burden.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"the-honest-part-this-breaks-things-and-the-maintainers-say-so\">The honest part: this breaks things, and the maintainers say so<a href=\"https://development-wec.wiline.com/docs/news/mcp-2026-07-28-spec/#the-honest-part-this-breaks-things-and-the-maintainers-say-so\" class=\"hash-link\" aria-label=\"Direct link to The honest part: this breaks things, and the maintainers say so\" title=\"Direct link to The honest part: this breaks things, and the maintainers say so\" translate=\"no\">​</a></h2>\n<p>Nothing here is backward compatible by accident. Soria Parra, in\n<a href=\"https://www.theregister.com/devops/2026/07/23/model-context-protocol-prepares-to-break-with-its-stateful-past/5276722\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">The Register's reporting</a>,\ndidn't sugarcoat it: <em>\"If you built your own implementation, it's going to be\na lot of uplift to make this correct,\"</em> and the stateless redesign, by his own\nadmission, <em>\"makes things on the wire a bit more complicated than they used\nto be\"</em> even as it removes session state. Roots, Sampling, and Logging are\nnow formally deprecated — Sampling for confusing semantics, Roots as \"a very\nniche thing,\" Logging for being excessively verbose — with a <strong>twelve-month\nminimum window</strong> before they're actually removed, alongside the legacy\nHTTP+SSE transport.</p>\n<p>Read charitably, this is a protocol growing up: a formal deprecation policy\nmeans you get a year of notice instead of a surprise break. Read skeptically\n— and The Register does — this is fixing problems that only showed up once\nMCP left local dev tooling for cloud deployment, which says something about\nhow much load-bearing infrastructure got built on the earlier design before\nanyone stress-tested it at scale. Both readings are true at once. That's not\na reason to panic; it's a reason to actually read the migration guide before\nyour integration tests do it for you.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"why-this-matters-for-you-specifically\">Why this matters for you, specifically<a href=\"https://development-wec.wiline.com/docs/news/mcp-2026-07-28-spec/#why-this-matters-for-you-specifically\" class=\"hash-link\" aria-label=\"Direct link to Why this matters for you, specifically\" title=\"Direct link to Why this matters for you, specifically\" translate=\"no\">​</a></h2>\n<p>If you're building agents against MCP servers you don't control, you likely\nnotice nothing immediately — Tier 1 SDKs (TypeScript, Python, Go, C#, with\nRust in beta) handle the negotiation. If you <strong>run</strong> an MCP server — for a\nRAG pipeline, an internal tool bridge, anything past a demo — this changes\nthree things you own directly: how it scales (stateless means your ops story\ngets simpler), how it's secured (OAuth 2.1 alignment means less of your own\nauth code to get wrong), and your clock (twelve months to move off Roots,\nSampling, Logging, and SSE transport before they're gone). None of that is\noptional just because you didn't ask for the rewrite.</p>\n<hr>\n<p>📖 <strong>Sources:</strong> <a href=\"https://blog.modelcontextprotocol.io/posts/2026-07-28/\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">Model Context Protocol Blog — the 2026-07-28 specification</a> · <a href=\"https://workos.com/blog/mcp-2026-spec-agent-authentication\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">WorkOS — authentication changes in the MCP 2026-07-28 spec</a> · <a href=\"https://www.theregister.com/devops/2026/07/23/model-context-protocol-prepares-to-break-with-its-stateful-past/5276722\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">The Register — MCP breaks with its stateful past</a> · <a href=\"https://venturebeat.com/infrastructure/mcp-just-got-its-biggest-update-ever-heres-what-changes-for-ai-agents\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">VentureBeat — MCP's biggest update, what changes for AI agents</a></p>",
            "url": "https://development-wec.wiline.com/docs/news/mcp-2026-07-28-spec/",
            "title": "MCP Just Shipped Its Biggest Update Ever — Here's What Actually Changes for AI Agent Engineers",
            "summary": "The 2026-07-28 MCP specification rips out sessions and rewrites authorization. If you build or run MCP servers, this changes your infrastructure, your auth flow, and your deprecation clock — whether you asked for it or not.",
            "date_modified": "2026-07-28T00:00:00.000Z",
            "author": {
                "name": "Rafael Fernandes",
                "url": "https://www.linkedin.com/in/rafaelmacariofernandes/"
            },
            "tags": [
                "ai-news",
                "mcp",
                "agents",
                "protocols",
                "infrastructure"
            ]
        },
        {
            "id": "https://development-wec.wiline.com/docs/news/gpt-5-6-token-economics/",
            "content_html": "<div class=\"newsHero newsHero--bg newsHero--split\"><div class=\"newsHero__bgSplit\" aria-hidden=\"true\"><div class=\"newsHero__panel newsHero__panel--a\" style=\"background-image:url(/docs/img/news/gpt-5.6.png)\"></div><div class=\"newsHero__panel newsHero__panel--b\" style=\"background-image:url(/docs/img/news/chatgpt.avif)\"></div><span class=\"newsHero__seam\"></span><div class=\"newsHero__scrim\"></div></div><span class=\"newsHero__eyebrow\">Models · AI News</span><h2 class=\"newsHero__title\">The frontier race just changed lanes: from smarter to cheaper per task</h2><div class=\"newsHero__transition\"><span class=\"newsHero__pill newsHero__pill--from\">Benchmark points</span><svg xmlns=\"http://www.w3.org/2000/svg\" width=\"20\" height=\"20\" viewBox=\"0 0 24 24\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"2.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\" class=\"lucide lucide-arrow-right newsHero__arrow\" aria-hidden=\"true\"><path d=\"M5 12h14\"></path><path d=\"m12 5 7 7-7 7\"></path></svg><span class=\"newsHero__pill newsHero__pill--to\">Token economics</span></div></div>\n<p>OpenAI shipped <strong>GPT-5.6</strong> on July 9 — a family of three models (Luna, Terra, Sol) — and the\nheadline claim isn't a leaderboard score. It's an efficiency number: frontier coding performance\non <strong>less than half the output tokens</strong>. If you build agents, that's a claim about your bill,\nnot about bragging rights. It's also exactly the kind of claim you should measure yourself.</p>\n<!-- -->\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"what-shipped\">What shipped<a href=\"https://development-wec.wiline.com/docs/news/gpt-5-6-token-economics/#what-shipped\" class=\"hash-link\" aria-label=\"Direct link to What shipped\" title=\"Direct link to What shipped\" translate=\"no\">​</a></h2>\n<p>Three tiers, cheapest to strongest, available in ChatGPT, Codex, and the OpenAI API:</p>\n<table><thead><tr><th>Model</th><th>Positioning</th><th>Input / Output (per 1M tokens)</th></tr></thead><tbody><tr><td>Luna</td><td>fastest, budget tier</td><td>$1 / $6</td></tr><tr><td>Terra</td><td>mid tier — \"performance competitive with GPT-5.5\"</td><td>$2.50 / $15</td></tr><tr><td>Sol</td><td>flagship, \"best coding model yet\"</td><td>$5 / $30</td></tr></tbody></table>\n<p>All three variants ship native tool use and multimodal reasoning, with long-context evals run\nout to 1M tokens; Sol also powers the new ChatGPT Work professional tier. The claims worth\nknowing (from <a href=\"https://openai.com/index/gpt-5-6/\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">OpenAI's announcement</a> — directional, not\ngospel): Sol scores <strong>80 on the Artificial Analysis Coding Agent Index</strong>, 2.8 points above\nAnthropic's Fable 5 — <em>while using less than half the output tokens, in less than half the time,\nat about a third of the estimated cost</em> — and <strong>88.8% on Terminal-Bench 2.1</strong>. On <em>Agents' Last\nExam</em>, an eval of long-running professional workflows across 55 fields, they report Terra and\nLuna beating Fable 5 at around <strong>one-sixteenth the estimated cost</strong>. Sam Altman's framing: Sol\nis \"<strong>54% more token-efficient</strong>\" on coding tasks. The release also leans hard on cybersecurity\n— \"frontier performance with significantly fewer tokens.\"</p>\n<p>Two API features ship under the same efficiency banner, and they're the most developer-relevant\npart of the launch: <strong>Programmatic Tool Calling</strong>, where the model writes and runs small\nin-memory programs that coordinate tools and process intermediate results — instead of passing\nevery tool response back through the model, so tool-heavy tasks burn fewer tokens and fewer\nround trips — and a <strong>multi-agent beta</strong> in the Responses API (the new Sol Ultra setting runs\nfour agents in parallel by default).</p>\n<p>One more first, buried in the timeline: GPT-5.6 was planned for June and shipped three weeks\nlate because a <strong>US government review gated the release</strong> — Commerce's Center for AI Standards\nand Innovation ran additional testing before OpenAI got permission for a public rollout. More on\nwhy that matters below.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"why-token-efficiency-is-the-real-story\">Why token efficiency is the real story<a href=\"https://development-wec.wiline.com/docs/news/gpt-5-6-token-economics/#why-token-efficiency-is-the-real-story\" class=\"hash-link\" aria-label=\"Direct link to Why token efficiency is the real story\" title=\"Direct link to Why token efficiency is the real story\" translate=\"no\">​</a></h2>\n<p>For chat, output tokens are a rounding error. For <strong>agents</strong>, they <em>are</em> the bill: an agentic\ncoding session burns tokens on every step — plans, diffs, retries, tool calls — and output\ntokens cost 5–6× input tokens on every pricing card above. A model that solves the same task on\nhalf the output tokens would be effectively <strong>half price and twice as fast at equal quality</strong>,\neven if it's only a couple of benchmark points better. <em>If</em> it holds on your tasks — and that\n\"if\" is the whole subject of the next section — that's a real and welcome shift in what vendors\ncompete on.</p>\n<p>That's why \"54% more token-efficient\" is a more aggressive competitive move than any leaderboard\njump. And OpenAI wasn't alone — the whole frontier spent this week competing on price per task:\n<strong>Meta launched Muse Spark 1.1</strong> for agentic coding at <strong>$1.25/$4.25</strong> per 1M tokens (with $20\nfree credits per account), and <strong>SpaceXAI released Grok 4.5</strong> — co-trained with Cursor — at\n<strong>$2/$6</strong>, marketed as roughly 6× cheaper than comparable frontier models. Even OpenAI's\ninfrastructure news pointed the same direction: engineers reportedly <strong>halved inference costs\nthrough software optimization alone</strong>. The race has visibly changed lanes, from benchmark points\nto cost per task.</p>\n<p>The three-tier ladder matters for the same reason. The emerging pattern is <strong>routing by task\ndifficulty</strong>: cheap tier for classification and extraction, mid tier for everyday generation,\nflagship only for the hard multi-step work. If you run a gateway (we've\n<a class=\"\" href=\"https://development-wec.wiline.com/docs/news/litellm-rust-gateway/\">written about why that layer matters</a>), this is what it's for.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"vendor-claims-are-eval-questions\">Vendor claims are eval questions<a href=\"https://development-wec.wiline.com/docs/news/gpt-5-6-token-economics/#vendor-claims-are-eval-questions\" class=\"hash-link\" aria-label=\"Direct link to Vendor claims are eval questions\" title=\"Direct link to Vendor claims are eval questions\" translate=\"no\">​</a></h2>\n<p>Here's the thing about \"54% more efficient\" and \"2.8 points above\": those numbers come from the\nvendor, measured on the vendor's chosen benchmark, with the vendor's harness. That's not an\naccusation — the numbers are very likely real <em>on that benchmark</em>, and we apply the same\ndiscount to everyone's, including the open-weight models we like (we said exactly this about\nGLM-5.2's tables). It's a structural point: <strong>no vendor benchmark can know your workload.</strong></p>\n<p>You don't even have to leave OpenAI's own coding table to see it — and credit to them for\npublishing the mixed rows rather than only the flattering ones:</p>\n<table><thead><tr><th>Coding eval (one vendor, one table)</th><th>GPT-5.6 Sol</th><th>Claude Fable 5</th><th>Who leads</th></tr></thead><tbody><tr><td>Artificial Analysis Coding Agent Index v1.1</td><td><strong>80</strong></td><td>77.2</td><td>Sol, +2.8</td></tr><tr><td>Terminal-Bench 2.1</td><td><strong>88.8%</strong></td><td>83.1%</td><td>Sol, +5.7</td></tr><tr><td>DeepSWE v1.1</td><td><strong>72.7%</strong></td><td>69.7%</td><td>Sol, +3.0</td></tr><tr><td>SWE-Bench Pro</td><td>64.6%</td><td><strong>80%</strong></td><td>Fable 5, +15.4</td></tr></tbody></table>\n<p><em>(Numbers from <a href=\"https://openai.com/index/gpt-5-6/\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">OpenAI's own announcement</a>.)</em></p>\n<p>Four coding benchmarks, and \"which is the better coding model\" flips depending on the row — with\nthe single largest gap pointing the <em>other</em> way. That's not a knock on either model or on the\ntable; it's what benchmarks are. If one vendor's own page can't produce a single answer, a\nlaunch-day headline certainly can't tell you what happens on <strong>your</strong> codebase. Token efficiency\ncompounds the problem: it varies wildly by task shape — a model that's terse on Python refactors\ncan be verbose on SQL or long-form answers, and an agent harness different from the vendor's can\nerase (or amplify) the whole advantage.</p>\n<p>The good news: this is a solved problem, and you already have the tooling if you followed our\nevals series. The same Promptfoo setup from\n<a class=\"\" href=\"https://development-wec.wiline.com/docs/tutorials/eval-models-promptfoo-wiline-inference/\">part 1</a> compares any two OpenAI-compatible\nendpoints on <em>your</em> prompts — with <strong>cost and latency assertions</strong>, not just quality grades.\nAnd the head-to-head worth running this week isn't GPT-5.6 against its own launch table — it's\nGPT-5.6 against the strongest open-weight model you can serve yourself:</p>\n<div class=\"language-yaml codeBlockContainer_Ckt0 theme-code-block\" style=\"--prism-color:#393A34;--prism-background-color:#f6f8fa\"><div class=\"codeBlockTitle_OeMC\">the experiment worth an afternoon (part-1 skill)</div><div class=\"codeBlockContent_QJqH\"><pre tabindex=\"0\" class=\"prism-code language-yaml codeBlock_bY9V thin-scrollbar\" style=\"color:#393A34;background-color:#f6f8fa\"><code class=\"codeBlockLines_e6Vv\"><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token key atrule\" style=\"color:#00a4db\">providers</span><span class=\"token punctuation\" style=\"color:#393A34\">:</span><span class=\"token plain\"></span><br></div><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\">  </span><span class=\"token punctuation\" style=\"color:#393A34\">-</span><span class=\"token plain\"> openai</span><span class=\"token punctuation\" style=\"color:#393A34\">:</span><span class=\"token plain\">chat</span><span class=\"token punctuation\" style=\"color:#393A34\">:</span><span class=\"token plain\">&lt;new</span><span class=\"token punctuation\" style=\"color:#393A34\">-</span><span class=\"token plain\">model</span><span class=\"token punctuation\" style=\"color:#393A34\">-</span><span class=\"token plain\">id</span><span class=\"token punctuation\" style=\"color:#393A34\">&gt;</span><span class=\"token plain\">          </span><span class=\"token comment\" style=\"color:#999988;font-style:italic\"># the challenger everyone's talking about</span><span class=\"token plain\"></span><br></div><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\">  </span><span class=\"token punctuation\" style=\"color:#393A34\">-</span><span class=\"token plain\"> </span><span class=\"token key atrule\" style=\"color:#00a4db\">id</span><span class=\"token punctuation\" style=\"color:#393A34\">:</span><span class=\"token plain\"> openai</span><span class=\"token punctuation\" style=\"color:#393A34\">:</span><span class=\"token plain\">chat</span><span class=\"token punctuation\" style=\"color:#393A34\">:</span><span class=\"token plain\">gemma4              </span><span class=\"token comment\" style=\"color:#999988;font-style:italic\"># the open-weight model already on WEC</span><span class=\"token plain\"></span><br></div><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\">    </span><span class=\"token key atrule\" style=\"color:#00a4db\">config</span><span class=\"token punctuation\" style=\"color:#393A34\">:</span><span class=\"token plain\"></span><br></div><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\">      </span><span class=\"token key atrule\" style=\"color:#00a4db\">apiBaseUrl</span><span class=\"token punctuation\" style=\"color:#393A34\">:</span><span class=\"token plain\"> https</span><span class=\"token punctuation\" style=\"color:#393A34\">:</span><span class=\"token plain\">//inference.wiline.com/v1</span><br></div><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\">      </span><span class=\"token key atrule\" style=\"color:#00a4db\">apiKeyEnvar</span><span class=\"token punctuation\" style=\"color:#393A34\">:</span><span class=\"token plain\"> WEC_API_KEY</span><br></div><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\"></span><span class=\"token key atrule\" style=\"color:#00a4db\">defaultTest</span><span class=\"token punctuation\" style=\"color:#393A34\">:</span><span class=\"token plain\"></span><br></div><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\">  </span><span class=\"token key atrule\" style=\"color:#00a4db\">assert</span><span class=\"token punctuation\" style=\"color:#393A34\">:</span><span class=\"token plain\"></span><br></div><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\">    </span><span class=\"token punctuation\" style=\"color:#393A34\">-</span><span class=\"token plain\"> </span><span class=\"token key atrule\" style=\"color:#00a4db\">type</span><span class=\"token punctuation\" style=\"color:#393A34\">:</span><span class=\"token plain\"> latency</span><br></div><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\">      </span><span class=\"token key atrule\" style=\"color:#00a4db\">threshold</span><span class=\"token punctuation\" style=\"color:#393A34\">:</span><span class=\"token plain\"> </span><span class=\"token number\" style=\"color:#36acaa\">5000</span><span class=\"token plain\"></span><br></div><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\">    </span><span class=\"token punctuation\" style=\"color:#393A34\">-</span><span class=\"token plain\"> </span><span class=\"token key atrule\" style=\"color:#00a4db\">type</span><span class=\"token punctuation\" style=\"color:#393A34\">:</span><span class=\"token plain\"> cost</span><br></div><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\">      </span><span class=\"token key atrule\" style=\"color:#00a4db\">threshold</span><span class=\"token punctuation\" style=\"color:#393A34\">:</span><span class=\"token plain\"> </span><span class=\"token number\" style=\"color:#36acaa\">0.002</span><br></div></code></pre></div></div>\n<p>Same prompts, both models, and the eval reports quality, tokens, latency, and cost side by side.\nAn afternoon of this tells you what no launch post can: whether the efficiency claim survives\ncontact with <em>your</em> traffic. (And if the model backs a RAG service or agent, gate it like we\ngated ours in <a class=\"\" href=\"https://development-wec.wiline.com/docs/tutorials/rag-docs-assistant-wiline-inference/\">the capstone</a> — judge for\nsemantics, deterministic regressions for known bugs.)</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"the-subtext-portability-just-got-more-valuable\">The subtext: portability just got more valuable<a href=\"https://development-wec.wiline.com/docs/news/gpt-5-6-token-economics/#the-subtext-portability-just-got-more-valuable\" class=\"hash-link\" aria-label=\"Direct link to The subtext: portability just got more valuable\" title=\"Direct link to The subtext: portability just got more valuable\" translate=\"no\">​</a></h2>\n<p>Two things happened around this launch that matter more together than apart.</p>\n<p>First, the <strong>government gate</strong>: for the first time, a frontier model's public release needed a\nfederal review to proceed — and the concern wasn't abstract. The same week, Sysdig documented\n<strong>JadePuffer</strong>, the first fully autonomous AI-agent ransomware operation, which exploited a\nLangflow CVE and then performed reconnaissance, lateral movement, and extortion on its own,\nrecovering from failed attempts within seconds. Whatever your politics, the engineering fact is\nthat access to closed frontier models now has one more valve you don't control — alongside\npricing, deprecations, and rate limits. We made this argument when <a class=\"\" href=\"https://development-wec.wiline.com/docs/news/glm-5-2-open-weight-top-10/\">an open-weight model cracked\nthe proprietary top 10</a>; this release made it for us.</p>\n<p>Second, the <strong>open-weight world had a loud week too</strong>: Google released <strong>Gemma 4</strong>, an\nopen-weight, natively multimodal family from 2.3B to 31B parameters (dense and MoE variants,\nvision and audio input, a thinking mode); Mistral opened early access on a new open-weight\nMixture-of-Experts family aimed squarely at the frontier gap; and Together AI closed an\n<strong>$800M Series C</strong> at an $8.3B valuation on the back of open-model inference — citing over\n$1B a year in bookings. The money is saying the same thing the GPT-5.6 delay is saying: models\nyou can download and run are a hedge worth paying for.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"the-takeaway--and-something-you-can-try-today\">The takeaway — and something you can try today<a href=\"https://development-wec.wiline.com/docs/news/gpt-5-6-token-economics/#the-takeaway--and-something-you-can-try-today\" class=\"hash-link\" aria-label=\"Direct link to The takeaway — and something you can try today\" title=\"Direct link to The takeaway — and something you can try today\" translate=\"no\">​</a></h2>\n<p>The developer takeaway is the same one that's held all series: <strong>keep model choice a config\nchange</strong>. Code against the OpenAI-compatible API, put a gateway in front, and swapping models —\nclosed, open, or whatever lands on the <a class=\"\" href=\"https://development-wec.wiline.com/docs/cloud_portal/platform/inference/models_hub/\">WEC Models catalog</a>\nnext — is a base-URL and model-ID edit, not a rewrite. We don't offer GPT-5.6 on WEC today;\nthe point is that if you build portable and eval before you migrate, that fact constrains you\nexactly zero.</p>\n<p>Because here's the trap in every launch week: <em>new</em> starts feeling like a reason. It isn't —\nit's a hypothesis. The question your eval should answer is whether the shiny closed model beats\nan open-weight model you control — on your prompts, at your latency budget, per dollar — by\nenough to be worth the valve someone else's hand is on. Sometimes it will, and then you switch\nwith evidence instead of hype. And sometimes the open model holds the line on the tasks you\nactually run, and you just saved yourself a migration and a dependency in one afternoon.</p>\n<p>You can run that experiment today, because\n<a href=\"https://deepmind.google/models/gemma/gemma-4/\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\"><strong>Gemma 4</strong></a> — the open-weight, natively\nmultimodal family Google shipped this same week — <strong>is already live on WEC inference</strong>. And it's\na serious counterpart, not a consolation prize. What Gemma 4 brings to the table:</p>\n<ul>\n<li class=\"\"><strong>Natively multimodal</strong> — image <em>and</em> audio understanding in an open model, so document\nscreenshots, scanned forms, and call recordings go through the same endpoint as your text\nprompts.</li>\n<li class=\"\"><strong>Thinking variants</strong> for the harder multi-step chains — the same \"spend tokens to reason\"\ntrade GPT-5.6 is selling, on weights you control.</li>\n<li class=\"\"><strong>A size ladder of its own</strong> (edge-sized E2B/E4B up to 31B) — the same route-by-difficulty\npattern from above, without a per-tier vendor contract.</li>\n<li class=\"\"><strong>140+ languages</strong>, and vendor benchmarks that put the 31B in frontier company — Google claims\nperformance comparable to models 10–30× larger. Same discount applies as to OpenAI's numbers,\nand the same tool settles it: put it in the eval.</li>\n</ul>\n<p>One model-ID swap and you're on it:</p>\n<div class=\"language-bash codeBlockContainer_Ckt0 theme-code-block\" style=\"--prism-color:#393A34;--prism-background-color:#f6f8fa\"><div class=\"codeBlockContent_QJqH\"><pre tabindex=\"0\" class=\"prism-code language-bash codeBlock_bY9V thin-scrollbar\" style=\"color:#393A34;background-color:#f6f8fa\"><code class=\"codeBlockLines_e6Vv\"><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token function\" style=\"color:#d73a49\">curl</span><span class=\"token plain\"> https://inference.wiline.com/v1/chat/completions </span><span class=\"token punctuation\" style=\"color:#393A34\">\\</span><span class=\"token plain\"></span><br></div><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\">  </span><span class=\"token parameter variable\" style=\"color:#36acaa\">-H</span><span class=\"token plain\"> </span><span class=\"token string\" style=\"color:#e3116c\">\"Authorization: Bearer </span><span class=\"token string variable\" style=\"color:#36acaa\">$WEC_API_KEY</span><span class=\"token string\" style=\"color:#e3116c\">\"</span><span class=\"token plain\"> </span><span class=\"token punctuation\" style=\"color:#393A34\">\\</span><span class=\"token plain\"></span><br></div><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\">  </span><span class=\"token parameter variable\" style=\"color:#36acaa\">-H</span><span class=\"token plain\"> </span><span class=\"token string\" style=\"color:#e3116c\">\"Content-Type: application/json\"</span><span class=\"token plain\"> </span><span class=\"token punctuation\" style=\"color:#393A34\">\\</span><span class=\"token plain\"></span><br></div><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\">  </span><span class=\"token parameter variable\" style=\"color:#36acaa\">-d</span><span class=\"token plain\"> </span><span class=\"token string\" style=\"color:#e3116c\">'{ \"model\": \"gemma4\", \"messages\": [{\"role\":\"user\",\"content\":\"Summarize this incident report…\"}] }'</span><br></div></code></pre></div></div>\n<p>GPT-5.6's token-efficiency push is good news for everyone — even if you never send OpenAI a\nsingle request, it drags the whole market toward pricing honesty. Take the positive at face\nvalue, take the numbers as hypotheses, and let your own eval — GPT-5.6 on one side, Gemma 4 on\nWEC on the other — make the call.</p>\n<hr>\n<p>📖 <strong>Sources:</strong> <a href=\"https://openai.com/index/gpt-5-6/\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">OpenAI — GPT-5.6 announcement</a> · <a href=\"https://deepmind.google/models/gemma/gemma-4/\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">Google DeepMind — Gemma 4</a> · <a href=\"https://techcrunch.com/2026/07/09/openai-launches-its-new-family-of-models-with-gpt-5-6/\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">TechCrunch — OpenAI launches GPT-5.6</a> · <a href=\"https://www.cnbc.com/2026/07/08/openai-expanding-gpt-5point6-ai-model-release-ending-government-limits.html\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">CNBC — public release after government limits</a> · <a href=\"https://www.axios.com/2026/07/09/ai-openai-gpt-release\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">Axios — GPT-5.6 and ChatGPT Work</a> · <a href=\"https://www.engadget.com/2210308/openai-rolls-out-gpt5-6-july-9/\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">Engadget — rollout timing</a> · <a href=\"https://www.techtimes.com/articles/319798/20260706/mistral-ai-targets-frontier-gap-open-weight-model-entering-july-early-access.htm\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">TechTimes — Mistral's open-weight MoE early access</a> · <a href=\"https://asanify.com/blog/news/open-weight-model-funding-july-7-2026/\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">Asanify — Together AI's $800M round</a> · <a href=\"https://medium.com/nlplanet/gpt-5-6-is-out-weekly-ai-newsletter-july-13th-2026-4502e4c324a7\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">NLPlanet — weekly AI newsletter, July 13</a> <em>(Muse Spark 1.1, Grok 4.5, Gemma 4, JadePuffer, inference-cost items)</em></p>",
            "url": "https://development-wec.wiline.com/docs/news/gpt-5-6-token-economics/",
            "title": "GPT-5.6: OpenAI's new pitch is cheaper per task, not just smarter — verify it on your workload before you switch",
            "summary": "GPT-5.6 (Luna, Terra, Sol) ships with a claim aimed straight at your bill: frontier coding scores on half the output tokens. What that means for agent economics, why you should verify it on your own workload — and how to build so model choice stays a config change.",
            "date_modified": "2026-07-13T00:00:00.000Z",
            "author": {
                "name": "Rafael Fernandes",
                "url": "https://www.linkedin.com/in/rafaelmacariofernandes/"
            },
            "tags": [
                "ai-news",
                "models",
                "openai",
                "token-efficiency",
                "evals"
            ]
        },
        {
            "id": "https://development-wec.wiline.com/docs/news/glm-5-2-open-weight-top-10/",
            "content_html": "<div class=\"newsHero newsHero--bg\" style=\"background-image:linear-gradient(rgba(2,12,31,0.62), rgba(2,12,31,0.80)), url(/docs/img/news/glm-cover-a.webp)\"><span class=\"newsHero__eyebrow\">Models · AI News</span><h2 class=\"newsHero__title\">An open-weight model just cracked the proprietary top 10</h2><div class=\"newsHero__transition\"><span class=\"newsHero__pill newsHero__pill--from\">Closed frontier</span><svg xmlns=\"http://www.w3.org/2000/svg\" width=\"20\" height=\"20\" viewBox=\"0 0 24 24\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"2.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\" class=\"lucide lucide-arrow-right newsHero__arrow\" aria-hidden=\"true\"><path d=\"M5 12h14\"></path><path d=\"m12 5 7 7-7 7\"></path></svg><span class=\"newsHero__pill newsHero__pill--to\">Open weights</span></div></div>\n<p>Look at almost any current model leaderboard and the top is a wall of Anthropic and\nOpenAI. Then, sitting in the top 10, there's one outlier that isn't proprietary at\nall: <strong>GLM-5.2</strong> from Z.ai — open weights, MIT-licensed. That's the story worth\npaying attention to.</p>\n<!-- -->\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"the-standing\">The standing<a href=\"https://development-wec.wiline.com/docs/news/glm-5-2-open-weight-top-10/#the-standing\" class=\"hash-link\" aria-label=\"Direct link to The standing\" title=\"Direct link to The standing\" translate=\"no\">​</a></h2>\n<p>On the <a href=\"https://arena.ai/leaderboard/agent\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">Arena.ai agent leaderboard</a>, GLM-5.2\n(Max) lands at <strong>#10</strong> — the <strong>only open-weight model in the top 10</strong>, surrounded\nentirely by closed frontier models from Anthropic and OpenAI. (Leaderboards move;\nthis is a snapshot — <a href=\"https://arena.ai/leaderboard/agent\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">check the live ranking</a>.)</p>\n<p><span class=\"zoomImage__wrap\"><img alt=\"GLM-5.2 (Max) on the Arena.ai agent leaderboard — the only open-weight model in the top 10\" src=\"https://development-wec.wiline.com/docs/assets/images/glm-5-2-leaderboard-3f2c9fe86e103693aa80fdbcda6b054b.png\" width=\"1400\" height=\"837\" class=\"zoomImage \" loading=\"lazy\"><span class=\"zoomImage__badge\" aria-hidden=\"true\"><svg viewBox=\"0 0 24 24\" width=\"16\" height=\"16\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"2\" stroke-linecap=\"round\"><circle cx=\"11\" cy=\"11\" r=\"7\"></circle><path d=\"M21 21l-4.3-4.3\"></path><path d=\"M11 8v6M8 11h6\"></path></svg></span></span></p>\n<p>That's the headline: not that it tops the chart, but that an <strong>MIT-licensed model you\ncan download, self-host, and ship commercially</strong> is now trading blows with models you\ncan only rent.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"what-glm-52-actually-is\">What GLM-5.2 actually is<a href=\"https://development-wec.wiline.com/docs/news/glm-5-2-open-weight-top-10/#what-glm-52-actually-is\" class=\"hash-link\" aria-label=\"Direct link to What GLM-5.2 actually is\" title=\"Direct link to What GLM-5.2 actually is\" translate=\"no\">​</a></h2>\n<ul>\n<li class=\"\"><strong>Open weights, MIT-licensed</strong> — no regional limits; download, self-host, fine-tune, and ship it commercially (<a href=\"https://huggingface.co/zai-org/GLM-5.2\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">weights on Hugging Face</a>).</li>\n<li class=\"\"><strong>A solid 1M-token context</strong> (~750k words), built for long-horizon agent work. Its new <strong>IndexShare</strong> attention reuses one indexer across every four sparse layers — Z.ai reports <strong>~2.9× fewer per-token FLOPs at 1M context</strong>, which is what keeps that window affordable to run.</li>\n<li class=\"\"><strong>Two thinking-effort levels (High / Max)</strong> to trade latency for depth — <code>Max</code> for hard multi-step coding, <code>High</code> for lighter, faster work.</li>\n<li class=\"\"><strong>Anthropic/OpenAI-compatible API</strong> — drop it into Claude Code, OpenClaw, Cline, and others with a base-URL + model-ID swap; your harness and prompts stay put.</li>\n</ul>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"how-it-compares\">How it compares<a href=\"https://development-wec.wiline.com/docs/news/glm-5-2-open-weight-top-10/#how-it-compares\" class=\"hash-link\" aria-label=\"Direct link to How it compares\" title=\"Direct link to How it compares\" translate=\"no\">​</a></h2>\n<p>Z.ai's published benchmarks put GLM-5.2 shoulder-to-shoulder with the closed frontier on coding, and ahead on some reasoning:</p>\n<table><thead><tr><th>Benchmark</th><th>GLM-5.2</th><th>Claude Opus 4.8</th><th>GPT-5.5</th></tr></thead><tbody><tr><td>SWE-bench Pro</td><td>62.1</td><td>69.2</td><td>58.6</td></tr><tr><td>Terminal-Bench 2.1 (best harness)</td><td>82.7</td><td>78.9</td><td>83.4</td></tr><tr><td>FrontierSWE (dominance)</td><td>74.4</td><td>75.1</td><td>72.6</td></tr><tr><td>AIME 2026</td><td>99.2</td><td>95.7</td><td>98.3</td></tr></tbody></table>\n<p>It edges Opus 4.8 on Terminal-Bench, beats GPT-5.5 on FrontierSWE, tops both on AIME, and trails Opus on SWE-bench Pro — remarkably close for a model you can simply download. <em>(Numbers from <a href=\"https://docs.z.ai/guides/llm/glm-5.2\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">Z.ai's GLM-5.2 benchmarks</a>; benchmarks are directional, not gospel.)</em></p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"why-it-matters\">Why it matters<a href=\"https://development-wec.wiline.com/docs/news/glm-5-2-open-weight-top-10/#why-it-matters\" class=\"hash-link\" aria-label=\"Direct link to Why it matters\" title=\"Direct link to Why it matters\" translate=\"no\">​</a></h2>\n<p>The gap between open-weight and proprietary frontier models has been closing all\nyear. What's changed is the <strong>terms</strong>: with an MIT license and a clean API, GLM-5.2\nis something you can <em>own and deploy</em>, not just call. When access to closed models\ncan shift with export controls or pricing overnight, an open-weight model that holds\ntop-10 quality is a foundation that stays put.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"run-it-on-wiline-edge-cloud\">Run it on WiLine Edge Cloud<a href=\"https://development-wec.wiline.com/docs/news/glm-5-2-open-weight-top-10/#run-it-on-wiline-edge-cloud\" class=\"hash-link\" aria-label=\"Direct link to Run it on WiLine Edge Cloud\" title=\"Direct link to Run it on WiLine Edge Cloud\" translate=\"no\">​</a></h2>\n<p>You don't need a third-party account to try it — <strong>GLM-5.2 is available on WiLine\nEdge Cloud through WEC Models</strong>, our OpenAI-compatible inference. Point any compatible\ntool at the WEC inference endpoint and use GLM-5.2 as the model:</p>\n<div class=\"language-bash codeBlockContainer_Ckt0 theme-code-block\" style=\"--prism-color:#393A34;--prism-background-color:#f6f8fa\"><div class=\"codeBlockContent_QJqH\"><pre tabindex=\"0\" class=\"prism-code language-bash codeBlock_bY9V thin-scrollbar\" style=\"color:#393A34;background-color:#f6f8fa\"><code class=\"codeBlockLines_e6Vv\"><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token function\" style=\"color:#d73a49\">curl</span><span class=\"token plain\"> https://inference.wiline.com/v1/chat/completions </span><span class=\"token punctuation\" style=\"color:#393A34\">\\</span><span class=\"token plain\"></span><br></div><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\">  </span><span class=\"token parameter variable\" style=\"color:#36acaa\">-H</span><span class=\"token plain\"> </span><span class=\"token string\" style=\"color:#e3116c\">\"Authorization: Bearer </span><span class=\"token string variable\" style=\"color:#36acaa\">$WEC_API_KEY</span><span class=\"token string\" style=\"color:#e3116c\">\"</span><span class=\"token plain\"> </span><span class=\"token punctuation\" style=\"color:#393A34\">\\</span><span class=\"token plain\"></span><br></div><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\">  </span><span class=\"token parameter variable\" style=\"color:#36acaa\">-H</span><span class=\"token plain\"> </span><span class=\"token string\" style=\"color:#e3116c\">\"Content-Type: application/json\"</span><span class=\"token plain\"> </span><span class=\"token punctuation\" style=\"color:#393A34\">\\</span><span class=\"token plain\"></span><br></div><div class=\"token-line\" style=\"color:#393A34\"><span class=\"token plain\">  </span><span class=\"token parameter variable\" style=\"color:#36acaa\">-d</span><span class=\"token plain\"> </span><span class=\"token string\" style=\"color:#e3116c\">'{ \"model\": \"glm-5.2\", \"messages\": [{\"role\":\"user\",\"content\":\"Refactor this for performance…\"}] }'</span><br></div></code></pre></div></div>\n<p>If you followed the <a class=\"\" href=\"https://development-wec.wiline.com/docs/tutorials/\">Self-hosting OpenClaw series</a>, this is the natural\nnext move: keep your agent, swap the model — point OpenClaw at GLM-5.2 on WEC instead\nof a closed provider, and you're running a top-10 model you fully control.</p>\n<hr>\n<p>📖 <strong>Sources:</strong> <a href=\"https://arena.ai/leaderboard/agent\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">Arena.ai agent leaderboard</a> · <a href=\"https://huggingface.co/zai-org/GLM-5.2\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">GLM-5.2 on Hugging Face</a> · <a href=\"https://docs.z.ai/guides/llm/glm-5.2\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">Z.ai model docs</a> · <a href=\"https://arxiv.org/abs/2602.15763\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">GLM-5 technical report (arXiv)</a></p>",
            "url": "https://development-wec.wiline.com/docs/news/glm-5-2-open-weight-top-10/",
            "title": "GLM-5.2: the only open-weight model in the top 10 — and you can run it on WEC",
            "summary": "GLM-5.2 is the lone open-weight, MIT-licensed model holding its own against the proprietary frontier — a 1M-token context and top open-source coding scores. And it's available on WiLine Edge Cloud.",
            "date_modified": "2026-06-24T00:00:00.000Z",
            "author": {
                "name": "Rafael Fernandes",
                "url": "https://www.linkedin.com/in/rafaelmacariofernandes/"
            },
            "tags": [
                "ai-news",
                "models",
                "open-weight",
                "glm",
                "inference"
            ]
        },
        {
            "id": "https://development-wec.wiline.com/docs/news/litellm-rust-gateway/",
            "content_html": "<figure class=\"newsHero newsHero--image\"><span class=\"newsHero__chip\">Infrastructure · AI News</span><img src=\"https://development-wec.wiline.com/docs/img/news/litellm-rust.webp\" alt=\"LiteLLM — migrating the AI gateway to Rust\" loading=\"eager\"></figure>\n<p>The AI ecosystem is quietly going through the same transition web infrastructure went through years ago: the performance-critical pieces are moving off interpreted runtimes onto systems languages like Rust. <a href=\"https://docs.litellm.ai/blog/litellm-rust-launch\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">LiteLLM rewriting its AI gateway in Rust</a> is the clearest evidence yet — and a sign the AI stack is maturing from experiment into production infrastructure.</p>\n<!-- -->\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"the-gateway-is-becoming-critical-infrastructure\">The gateway is becoming critical infrastructure<a href=\"https://development-wec.wiline.com/docs/news/litellm-rust-gateway/#the-gateway-is-becoming-critical-infrastructure\" class=\"hash-link\" aria-label=\"Direct link to The gateway is becoming critical infrastructure\" title=\"Direct link to The gateway is becoming critical infrastructure\" translate=\"no\">​</a></h2>\n<p>LiteLLM is the open-source proxy a lot of teams put in front of their models to get one OpenAI-compatible endpoint across 100+ providers. If you've run one in production, this line from the announcement will feel familiar:</p>\n<blockquote>\n<p>Under real load, CPU and memory climb with concurrency, and pods get OOM-killed at the worst time.</p>\n</blockquote>\n<p>That's the quiet tax of a gateway: it sits on the hot path of <em>every</em> request — every completion, embedding, moderation call, and agent action flows through it — so its own overhead and memory footprint multiply across pods and regions. For years AI conversations were about model quality. As teams ship agents, RAG, and multi-model routing to production, the layer <em>in front</em> of the model is turning into a first-class infrastructure concern. Moving it to Rust is what that realization looks like in code.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"the-numbers\">The numbers<a href=\"https://development-wec.wiline.com/docs/news/litellm-rust-gateway/#the-numbers\" class=\"hash-link\" aria-label=\"Direct link to The numbers\" title=\"Direct link to The numbers\" translate=\"no\">​</a></h2>\n<p>From LiteLLM's published benchmarks (reproducible — the harness ships with the post):</p>\n<div class=\"metricCompare\"><div class=\"metricCard\"><span class=\"metricCard__label\">Per-request overhead</span><span class=\"metricCard__factor\">~150× lower</span><div class=\"metricCard__rows\"><div class=\"metricCard__row metricCard__row--a\"><span class=\"metricCard__name\">LiteLLM (Python)</span><span class=\"metricCard__val\">~7.5 ms</span></div><div class=\"metricCard__row metricCard__row--b\"><span class=\"metricCard__name\">Rust gateway</span><span class=\"metricCard__val\">~0.05 ms</span></div></div></div><div class=\"metricCard\"><span class=\"metricCard__label\">Throughput under load</span><span class=\"metricCard__factor\">~15× higher</span><div class=\"metricCard__rows\"><div class=\"metricCard__row metricCard__row--a\"><span class=\"metricCard__name\">LiteLLM (Python)</span><span class=\"metricCard__val\">453 req/s</span></div><div class=\"metricCard__row metricCard__row--b\"><span class=\"metricCard__name\">Rust gateway</span><span class=\"metricCard__val\">6,782 req/s</span></div></div></div><div class=\"metricCard\"><span class=\"metricCard__label\">Peak memory under load</span><span class=\"metricCard__factor\">~11× lighter</span><div class=\"metricCard__rows\"><div class=\"metricCard__row metricCard__row--a\"><span class=\"metricCard__name\">LiteLLM (Python)</span><span class=\"metricCard__val\">358.9 MB</span></div><div class=\"metricCard__row metricCard__row--b\"><span class=\"metricCard__name\">Rust gateway</span><span class=\"metricCard__val\">31.7 MB</span></div></div></div></div>\n<p>This measures the gateway <em>forwarding path</em> (transform → forward → handle response), not a full production workload — but that's exactly the layer you don't want eating CPU and memory under concurrency.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"the-real-win-is-memory-not-latency\">The real win is memory, not latency<a href=\"https://development-wec.wiline.com/docs/news/litellm-rust-gateway/#the-real-win-is-memory-not-latency\" class=\"hash-link\" aria-label=\"Direct link to The real win is memory, not latency\" title=\"Direct link to The real win is memory, not latency\" translate=\"no\">​</a></h2>\n<p>Most readers will fixate on <strong>150× lower overhead</strong>. But for anyone <em>operating</em> a gateway, the more consequential number is <strong>11× less memory</strong>: 359 MB → ~32 MB. Latency is a per-request improvement; memory is what drives your bill and your reliability.</p>\n<p>A gateway that holds ~32 MB instead of ~359 MB changes the operational math across the board:</p>\n<ul>\n<li class=\"\"><strong>Kubernetes sizing</strong> — smaller pods, higher density per node.</li>\n<li class=\"\"><strong>Cloud cost</strong> — that footprint multiplies across every pod, region, and replica you run.</li>\n<li class=\"\"><strong>Autoscaling</strong> — lower, more predictable memory means less scaling churn.</li>\n<li class=\"\"><strong>OOM crashes</strong> — the failure mode that takes you down at peak largely goes away.</li>\n</ul>\n<p>When a component sits on the hot path of every request, shaving an order of magnitude off its memory compounds at scale far more than the headline latency figure.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"a-low-risk-rollout\">A low-risk rollout<a href=\"https://development-wec.wiline.com/docs/news/litellm-rust-gateway/#a-low-risk-rollout\" class=\"hash-link\" aria-label=\"Direct link to A low-risk rollout\" title=\"Direct link to A low-risk rollout\" translate=\"no\">​</a></h2>\n<p>This is <strong>not a v2 and not a rewrite you have to migrate to</strong>. Config files, database schema, client APIs, and provider coverage stay the same. They're moving it in careful stages — a pure-Rust core via PyO3 bindings first (data transformation, no I/O), then the full server on axum/hyper — each route shipped to production behind passing parity tests before the next one starts:</p>\n<figure class=\"stageFlow\"><div class=\"stageFlow__track\"><div class=\"stageFlow__card\" style=\"background:rgba(var(--primary-rgb), 0.050);border-color:rgba(var(--primary-rgb), 0.250)\"><span class=\"stageFlow__stage\">Stage 0 · Today</span><span class=\"stageFlow__title\">Python proxy</span><span class=\"stageFlow__tag\">0% Rust</span></div><svg xmlns=\"http://www.w3.org/2000/svg\" width=\"22\" height=\"22\" viewBox=\"0 0 24 24\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"2.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\" class=\"lucide lucide-arrow-right stageFlow__arrow\" aria-hidden=\"true\"><path d=\"M5 12h14\"></path><path d=\"m12 5 7 7-7 7\"></path></svg><div class=\"stageFlow__card\" style=\"background:rgba(var(--primary-rgb), 0.123);border-color:rgba(var(--primary-rgb), 0.383)\"><span class=\"stageFlow__stage\">Stage 1 · Core in Rust</span><span class=\"stageFlow__title\">Python drives transforms via PyO3</span><span class=\"stageFlow__tag\">transforms + router</span></div><svg xmlns=\"http://www.w3.org/2000/svg\" width=\"22\" height=\"22\" viewBox=\"0 0 24 24\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"2.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\" class=\"lucide lucide-arrow-right stageFlow__arrow\" aria-hidden=\"true\"><path d=\"M5 12h14\"></path><path d=\"m12 5 7 7-7 7\"></path></svg><div class=\"stageFlow__card\" style=\"background:rgba(var(--primary-rgb), 0.197);border-color:rgba(var(--primary-rgb), 0.517)\"><span class=\"stageFlow__stage\">Stage 2 · Thin shell</span><span class=\"stageFlow__title\">FastAPI shell, hot path in Rust</span><span class=\"stageFlow__tag\">~full forwarding path</span></div><svg xmlns=\"http://www.w3.org/2000/svg\" width=\"22\" height=\"22\" viewBox=\"0 0 24 24\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"2.5\" stroke-linecap=\"round\" stroke-linejoin=\"round\" class=\"lucide lucide-arrow-right stageFlow__arrow\" aria-hidden=\"true\"><path d=\"M5 12h14\"></path><path d=\"m12 5 7 7-7 7\"></path></svg><div class=\"stageFlow__card\" style=\"background:rgba(var(--primary-rgb), 0.270);border-color:rgba(var(--primary-rgb), 0.650)\"><span class=\"stageFlow__stage\">Stage 3 · Pure Rust</span><span class=\"stageFlow__title\">axum server, Python in a sidecar</span><span class=\"stageFlow__tag\">100% Rust</span></div></div><figcaption class=\"stageFlow__caption\">Four stages — each shipped to production behind passing parity tests before the next begins.</figcaption></figure>\n<p>Beta signup is open now; the roadmap targets OCR routes by mid-August 2026, <code>/chat/completions</code> and <code>/messages</code> by September, and the full server by <strong>December 1, 2026</strong>.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"what-this-means-for-ai-builders\">What this means for AI builders<a href=\"https://development-wec.wiline.com/docs/news/litellm-rust-gateway/#what-this-means-for-ai-builders\" class=\"hash-link\" aria-label=\"Direct link to What this means for AI builders\" title=\"Direct link to What this means for AI builders\" translate=\"no\">​</a></h2>\n<p>For developers building on <strong>WiLine Edge Cloud</strong>, the gateway sits directly between your applications and your models — so a leaner, faster gateway flows straight through to the apps you ship:</p>\n<ul>\n<li class=\"\"><strong>Faster AI APIs.</strong> Less proxy overhead means faster responses where model latency is already low — embeddings, reranking, moderation, classification. On those workloads the gateway <em>was</em> the tax; now it nearly isn't.</li>\n<li class=\"\"><strong>Better reliability.</strong> Lower memory pressure reduces OOM kills, request failures, and autoscaling churn — the things that quietly erode an AI product's uptime in production.</li>\n<li class=\"\"><strong>More efficient multi-model deployments.</strong> If you route traffic across many providers, gateway cost stops scaling as aggressively with traffic — you serve more without your proxy fleet ballooning.</li>\n<li class=\"\"><strong>Stronger infrastructure foundations.</strong> As AI apps become production systems rather than experiments, the infra layers underneath them matter as much as model quality.</li>\n</ul>\n<p>If you followed the <a class=\"\" href=\"https://development-wec.wiline.com/docs/tutorials/\">OpenClaw series</a>, you already put a gateway-shaped thing on the critical path — a reverse proxy, a model router, an agent runtime. The lesson generalizes.</p>\n<h2 class=\"anchor anchorTargetStickyNavbar_Vzrq\" id=\"the-bigger-lesson\">The bigger lesson<a href=\"https://development-wec.wiline.com/docs/news/litellm-rust-gateway/#the-bigger-lesson\" class=\"hash-link\" aria-label=\"Direct link to The bigger lesson\" title=\"Direct link to The bigger lesson\" translate=\"no\">​</a></h2>\n<p>For years, most AI discussion focused on model quality. But as organizations deploy agents, retrieval systems, and multi-model workflows in production, <strong>the layer in front of your models is infrastructure</strong> — and every millisecond and megabyte on the hot path compounds at scale. LiteLLM's move to Rust reflects a broader industry realization: infrastructure efficiency is no longer a footnote to model performance, it's part of it.</p>\n<p><strong>Worth watching, not yet worth switching:</strong> it's beta, and the Python proxy isn't going anywhere. But the direction of travel is clear.</p>\n<hr>\n<p>📖 <strong>Read the full announcement</strong> — the benchmarks, the route-by-route migration plan, and the architecture diagrams are all worth your time: <a href=\"https://docs.litellm.ai/blog/litellm-rust-launch\" target=\"_blank\" rel=\"noopener noreferrer\" class=\"\">LiteLLM — Building the fastest AI gateway in Rust</a>.</p>",
            "url": "https://development-wec.wiline.com/docs/news/litellm-rust-gateway/",
            "title": "Why LiteLLM Is Rewriting Its Gateway in Rust — and Why AI Developers Should Care",
            "summary": "LiteLLM is moving its AI gateway from Python to Rust. It's a signal that AI gateways are becoming critical infrastructure — with real consequences for latency, cost, and reliability on WiLine Edge Cloud.",
            "date_modified": "2026-06-24T00:00:00.000Z",
            "author": {
                "name": "Rafael Fernandes",
                "url": "https://www.linkedin.com/in/rafaelmacariofernandes/"
            },
            "tags": [
                "ai-news",
                "gateways",
                "performance",
                "self-hosting"
            ]
        }
    ]
}