AI Tutorials
Hands-on, tested guides for self-hosting AI infrastructure on a WEC Instance — deploy agents, connect your own models, add channels, evals, and observability. Every guide is run from scratch on a real box, with the actual commands, versions, and fixes.
Hardening self-hosted AI infra
Your firewall is lying to you: hardening Docker networks for multi-agent systems
A box running five agent stacks, a firewall set to deny everything, and services still answering from the public internet. We probe a real deployment, find a database with no password, prove why UFW never sees Docker traffic, and fix it four ways — every command and result from a live run.
Read more →Root by default: hardening container privilege on a self-hosted AI stack
Two of the six containers behind a live Langfuse deployment were running as root, for no reason anyone chose. Fixing it took one line each — and broke a service that had nothing to do with the fix. A real conversion, a real coordination failure, and how to catch both.
Read more →One login for everything: putting authentik in front of a self-hosted app
Part 2 fixed how containers run. This fixes who gets to log in. Deploy authentik as your own identity provider, connect Langfuse to it over OIDC, and understand the pieces — provider, application, redirect URI, scopes — instead of copying a config. Includes the two failures a real run produced: a compose file that silently swallows your settings, and OAuthAccountNotLinked.
Read more →The binding that wasn't there: group access control, and what SSO doesn't protect
Part 3 shipped SSO with empty bindings. Closing that gap on the LiteLLM gateway turned up something worse: authentik's wizard let us configure the bindings, showed them in its own table, and saved none of them — no error, no warning. Then, once access control actually worked, the API answered a request from a shell with no account, no session and no membership. Both of those are the post.
Read more →The key that expires: giving an agent its own identity at the gateway
Part 4 ended with a shell, a static key and a model that answered. This replaces that key with a token your identity provider issues, that carries a scope, and that dies after five minutes. authentik issues it over client credentials, the gateway verifies it in twenty-five lines, and three status codes prove the boundary. Includes the requirement authentik doesn't document, and the licence wall you hit if you follow LiteLLM's own guide.
Read more →The wrong valid token: authenticating an MCP tools server with authentik
Agent orchestration part 4 scoped tools in the client and admitted the server still trusted whoever reached the port. This puts a JWT verifier in front of it, with its own authentik client, so the tools server refuses anonymous callers — and refuses the gateway's own valid token, because a separate provider means a separate issuer. Includes the documented client-credentials helper that cannot work here, and the four silent 404s it produces instead of saying so.
Read more →Firewall an agent container so it can reach one API and nothing else
Every part of this series so far has controlled what an agent may call. None of it controls where an agent may go. OWASP's mitigations for Excessive Agency are three rules about tools and say nothing about the network. This puts a default-deny egress policy in front of one container, allows exactly one API, and proves the boundary with a lookup that succeeds and a packet that dies. Includes the conntrack rule everyone tells you to add and this one doesn't need, and a measurement that was wrong by three orders of magnitude.
Read more →Move the Docker daemon off root so a container escape lands on an ordinary user
If you can run docker without sudo, you can already read every file on the machine. Not through a bug — that is how Docker is built. One command proves it. For a host running agents that execute generated code, that is the whole attack, and Part 2's non-root containers do not touch it. This moves the daemon itself to an unprivileged account, with the UID arithmetic you can check, four places the install walks you into a wall on Ubuntu, and what stops working afterwards.
Read more →Sandbox the code your agent writes, and prove every limit actually applied
OWASP names running model-generated code through exec or eval as its own vulnerability, tells you to treat the model as untrusted — and says nothing about how to contain it. So we built the containment: no network, 256 MB, 64 processes, read-only disk, no capabilities. Every one of those held. The CPU limit was refused outright, and eight busy loops took 605% of the host. The cause is one line of systemd configuration nobody mentions, and the fix needed no reboot despite the documentation insisting otherwise.
Read more →Give each MCP tool its own scope, and return a refusal the client can act on
Part 6 gated the whole server with one scope, which meant the token that let an agent look up a customer also let it issue refunds. Splitting that per tool is easy. Making the refusal useful is not: the obvious implementation returns HTTP 200 with the error buried in the body, so a client has nothing to trigger step-up authorization on. Getting the 403 and WWW-Authenticate challenge the spec defines means leaving the framework's abstraction entirely — and the version that does it decodes the token a second time, unverified.
Read more →Agent orchestration with LangGraph
Checkpoint a LangGraph agent on a WEC Instance so crashes cost you nothing
Persist agent state to Postgres on a WEC Instance, call a model on the WEC Inference API, then kill the process mid-run twice and resume both times from exactly where it stopped — plus the human-approval pause, and the interrupt that silently fires your side effects twice.
Read more →Hand work between LangGraph agents without corrupting shared state
A supervisor routes work to specialists, an agent hands off mid-task, and two agents write the same state key in the same step. One of those raises an error rather than picking a winner — and the routing decision costs 412 output tokens to say one word.
Read more →Book appointments and refund invoices from a LangGraph agent over MCP
The scheduler and billing agents from Parts 1 and 2 never actually scheduled or billed anything. This gives them a calendar, a customer list and a refund that no model can issue alone — over MCP, on the stateless spec, with the model on WEC Inference. Includes three things nobody documents: your server still defaults to sessions, ctx.elicit is the old API and its era error goes to the model, not to you, and cache=True does nothing unless the server advertises a TTL.
Read more →Split one MCP toolbox between two agents so the scheduler cannot issue refunds
Part 3 handed one agent all five tools and let a human pause guard the dangerous one. That guard depends on the model asking. This splits the same MCP server between a scheduler and a billing agent, so the scheduler cannot refund an invoice for the simplest reason available — the tool was never in its list. Then the supervisor goes back on top, and routing turns out not to be the guarantee.
Read more →Self-hosting an LLM gateway
One endpoint, many models: deploy an LLM gateway on a WEC Instance
Every app on a box holding the same API key is a problem waiting to happen. A gateway fixes that — one endpoint, scoped keys per app, per-key budgets, and a log of who spent what. Deployed for real on a host already running five other stacks, including the parts that surprised us.
Read more →LiteLLM complexity routing: the right model for each request, and what it costs in latency
How LiteLLM's complexity router decides which model answers a request — the seven scoring dimensions, the arithmetic on a real prompt, and a measured comparison of the free keyword scorer against an LLM classifier.
Read more →Mask PII at the gateway: set up Presidio, plus the one line the docs leave out
Set up PII masking on a self-hosted LLM gateway with Presidio. Masking can sit in three places — before the model, on the response, or only on the path to your logs — and only one keeps your app working. Follow the documented config and your model's answers come back redacted; here is why, and the single setting that fixes it.
Read more →Clean traces, untouched answers: masking PII in LiteLLM's logs without corrupting the response
Send a gateway's traffic to Langfuse and the model's own reply carries the PII straight back into your traces. Widen Presidio's scope and it masks the answer your users receive instead. Neither setting gives you both, so here is a forty-line callback that does — measured, and with the mistake that quietly makes it slower.
Read more →Load test an LLM gateway and find the stall the median hides
Part 4 measured a blocking call in a callback at 200-350ms and said the real damage only shows under concurrency. This runs it: sixteen requests at once, with a control against the upstream so you can tell your gateway's fault from the model's. The medians of the two versions are almost identical. One version drops requests and silently stops masking.
Read more →Self-hosting Hermes
Self-host the Hermes Agent with persistent memory
Deploy Nous Research's open-source Hermes Agent on your WEC Instance — with SQLite-backed persistent memory that survives a full reboot. Real install, model config, and a memory-survives-restart test.
Read more →Migrate OpenClaw's Telegram Bot to Hermes
Move the same Telegram bot from the OpenClaw series onto Hermes — same chat, same users, a different agent answering underneath. Real setup wizard, a real allowlist gotcha, and proof it answers.
Read more →AI evals & observability
Evaluate your models with Promptfoo on the WEC Inference API
Stop eyeballing LLM output. Build a real evaluation harness with Promptfoo pointed at the WEC Inference API — assertions, latency guardrails, JSON-schema checks, model-graded rubrics, an all-WEC model comparison, and a CI gate. Every command and result is real.
Read more →Trustworthy JSON: schema-validate your model's structured output
LLMs promise JSON and deliver markdown fences and reasoning. Build a Promptfoo eval that classifies support tickets into schema-validated JSON on the WEC Inference API — with transforms to recover messy output, a model-reliability matrix, and a CI gate. Every command and result is real.
Read more →Stop hand-writing test cases: generate an eval dataset with the WEC API
Two hand-typed test cases aren't an eval. Use the WEC Inference API to generate a labeled evaluation dataset — then validate and curate it, because generated labels aren't automatically correct. The result is real test data at scale; feed it to your harness and coverage surfaces the misclassifications and debatable labels two cases would hide. Every command and result is real.
Read more →Catch what your tests miss: observe and score your WEC app in production with Langfuse
CI evals pass on a fixed test set — but production sends inputs you never tested. Self-host Langfuse on a WEC Instance, trace every real call, auto-score live traffic with an LLM judge, drill into the exact step that broke, and feed failures back to make your evals stronger. Every command — and every dead end — is real.
Read more →The capstone: build, evaluate, and observe a RAG docs assistant on the WEC API
Build a production-shaped RAG service in Docker: scrape a real docs site, embed locally, generate on the WEC Inference API — then catch a real hallucination, root-cause it to your own scraper, fix it, and pin it with a regression test. Every command, number, and error is real.
Read more →Regression-test your RAG service with DeepEval — and settle a model debate with data
Wrap the RAG assistant from part 5 in a containerized DeepEval suite judged by gemma4 on the WEC Inference API — no OpenAI key anywhere. Then use it to answer a real question: should we swap our generation model? Same quality, 5.5× the tokens: the eval says no. Every command, number, and error is real.
Read more →Component-level tracing: debugging agent tool calls
Your agent's thinking and its actions are two different layers. Build a tool-calling agent on WEC Inference, trace it with Langfuse, then debug two real failures from the trace — including the confident, wrong answer that never throws a stack trace.
Read more →Add web search to your WEC Inference calls
Your RAG assistant answers great from your docs — but not about anything recent. WEC Inference now has a built-in web-search tool: add it to a chat call and the model pulls in current info. A quick before/after test.
Read more →WhatsApp automation on WEC
Self-hosting OpenClaw
Deploy OpenClaw on a WEC Instance via Docker Compose
From a fresh WEC Instance to a self-hosted OpenClaw agent that actually answers — Docker Compose, your own model key, and every real error and fix from a live deploy.
Read more →Secure OpenClaw with a Caddy reverse proxy + HTTPS
Put Caddy in front of OpenClaw for real HTTPS and device-paired auth, then close the gateway's ports so the proxy is the only way in — gotchas and all.
Read more →Add a Telegram channel to OpenClaw
Talk to your self-hosted OpenClaw agent from your phone. Create a Telegram bot, connect it, clear the pairing gate, and chat — captured from a real run.
Read more →Make OpenClaw private with a NetBird mesh VPN
Put OpenClaw on a private mesh with NetBird, repoint your hostname at the mesh IP, and close the public ports — so the same chat URL works only for your devices. Real commands and gotchas from a live run.
Read more →Run OpenClaw on WEC Models
Swap the closed provider for WiLine's own inference: point your self-hosted OpenClaw agent at WEC Models — same box, a base-URL/key/model change, no per-token lock-in.
Read more →Add a WhatsApp channel to OpenClaw
Put your self-hosted agent on WhatsApp. Add the channel, trust the plugin, scan a QR — and understand why a companion link makes the agent act as your account. Every command and error from a real run.
Read more →