Skip to main content

AI Tutorials

Hands-on, tested guides for self-hosting AI infrastructure on a WEC Instance — deploy agents, connect your own models, add channels, evals, and observability. Every guide is run from scratch on a real box, with the actual commands, versions, and fixes.

Level
Topic

Hardening self-hosted AI infra

IntermediatePart 1

Your firewall is lying to you: hardening Docker networks for multi-agent systems

· 15 min read

A box running five agent stacks, a firewall set to deny everything, and services still answering from the public internet. We probe a real deployment, find a database with no password, prove why UFW never sees Docker traffic, and fix it four ways — every command and result from a live run.

dockersecuritynetworkingself-hostingagentshardening
Read more →
IntermediatePart 2

Root by default: hardening container privilege on a self-hosted AI stack

· 8 min read

Two of the six containers behind a live Langfuse deployment were running as root, for no reason anyone chose. Fixing it took one line each — and broke a service that had nothing to do with the fix. A real conversion, a real coordination failure, and how to catch both.

dockersecuritycontainersself-hostingagentshardening
Read more →
IntermediatePart 3

One login for everything: putting authentik in front of a self-hosted app

· 20 min read

Part 2 fixed how containers run. This fixes who gets to log in. Deploy authentik as your own identity provider, connect Langfuse to it over OIDC, and understand the pieces — provider, application, redirect URI, scopes — instead of copying a config. Includes the two failures a real run produced: a compose file that silently swallows your settings, and OAuthAccountNotLinked.

securityssooidcauthentikidentitylangfuseself-hostinghardening
Read more →
IntermediatePart 4

The binding that wasn't there: group access control, and what SSO doesn't protect

· 16 min read

Part 3 shipped SSO with empty bindings. Closing that gap on the LiteLLM gateway turned up something worse: authentik's wizard let us configure the bindings, showed them in its own table, and saved none of them — no error, no warning. Then, once access control actually worked, the API answered a request from a shell with no account, no session and no membership. Both of those are the post.

securityssooidcauthentikidentityrbaclitellmgatewayhardening
Read more →
IntermediatePart 5

The key that expires: giving an agent its own identity at the gateway

· 19 min read

Part 4 ended with a shell, a static key and a model that answered. This replaces that key with a token your identity provider issues, that carries a scope, and that dies after five minutes. authentik issues it over client credentials, the gateway verifies it in twenty-five lines, and three status codes prove the boundary. Includes the requirement authentik doesn't document, and the licence wall you hit if you follow LiteLLM's own guide.

securityoauth2oidcauthentikidentitylitellmgatewayagentshardening
Read more →
IntermediatePart 6

The wrong valid token: authenticating an MCP tools server with authentik

· 16 min read

Agent orchestration part 4 scoped tools in the client and admitted the server still trusted whoever reached the port. This puts a JWT verifier in front of it, with its own authentik client, so the tools server refuses anonymous callers — and refuses the gateway's own valid token, because a separate provider means a separate issuer. Includes the documented client-credentials helper that cannot work here, and the four silent 404s it produces instead of saying so.

securityoauth2oidcauthentikidentitymcpagentsself-hostinghardening
Read more →
IntermediatePart 7

Firewall an agent container so it can reach one API and nothing else

· 21 min read

Every part of this series so far has controlled what an agent may call. None of it controls where an agent may go. OWASP's mitigations for Excessive Agency are three rules about tools and say nothing about the network. This puts a default-deny egress policy in front of one container, allows exactly one API, and proves the boundary with a lookup that succeeds and a packet that dies. Includes the conntrack rule everyone tells you to add and this one doesn't need, and a measurement that was wrong by three orders of magnitude.

securitydockernetworkingiptablesfirewallagentsself-hostinghardening
Read more →
IntermediatePart 8

Move the Docker daemon off root so a container escape lands on an ordinary user

· 14 min read

If you can run docker without sudo, you can already read every file on the machine. Not through a bug — that is how Docker is built. One command proves it. For a host running agents that execute generated code, that is the whole attack, and Part 2's non-root containers do not touch it. This moves the daemon itself to an unprivileged account, with the UID arithmetic you can check, four places the install walks you into a wall on Ubuntu, and what stops working afterwards.

securitydockerrootlesscontainersprivilegeagentsself-hostinghardening
Read more →
IntermediatePart 9

Sandbox the code your agent writes, and prove every limit actually applied

· 11 min read

OWASP names running model-generated code through exec or eval as its own vulnerability, tells you to treat the model as untrusted — and says nothing about how to contain it. So we built the containment: no network, 256 MB, 64 processes, read-only disk, no capabilities. Every one of those held. The CPU limit was refused outright, and eight busy loops took 605% of the host. The cause is one line of systemd configuration nobody mentions, and the fix needed no reboot despite the documentation insisting otherwise.

securitydockersandboxagentscode-executioncgroupsself-hostinghardening
Read more →
IntermediatePart 10

Give each MCP tool its own scope, and return a refusal the client can act on

· 12 min read

Part 6 gated the whole server with one scope, which meant the token that let an agent look up a customer also let it issue refunds. Splitting that per tool is easy. Making the refusal useful is not: the obvious implementation returns HTTP 200 with the error buried in the body, so a client has nothing to trigger step-up authorization on. Getting the 403 and WWW-Authenticate challenge the spec defines means leaving the framework's abstraction entirely — and the version that does it decodes the token a second time, unverified.

securitymcpoauth2authentikjwtscopesagentsself-hostinghardening
Read more →

Agent orchestration with LangGraph

IntermediatePart 1

Checkpoint a LangGraph agent on a WEC Instance so crashes cost you nothing

· 20 min read

Persist agent state to Postgres on a WEC Instance, call a model on the WEC Inference API, then kill the process mid-run twice and resume both times from exactly where it stopped — plus the human-approval pause, and the interrupt that silently fires your side effects twice.

aiagentslanggraphstatepostgreslangfuseself-hostingwecwec-inference
Read more →
IntermediatePart 2

Hand work between LangGraph agents without corrupting shared state

· 16 min read

A supervisor routes work to specialists, an agent hands off mid-task, and two agents write the same state key in the same step. One of those raises an error rather than picking a winner — and the routing decision costs 412 output tokens to say one word.

aiagentslanggraphmulti-agentstatepostgreslangfuseself-hostingwecwec-inference
Read more →
IntermediatePart 3

Book appointments and refund invoices from a LangGraph agent over MCP

· 20 min read

The scheduler and billing agents from Parts 1 and 2 never actually scheduled or billed anything. This gives them a calendar, a customer list and a refund that no model can issue alone — over MCP, on the stateless spec, with the model on WEC Inference. Includes three things nobody documents: your server still defaults to sessions, ctx.elicit is the old API and its era error goes to the model, not to you, and cache=True does nothing unless the server advertises a TTL.

aiagentslanggraphlangchainmcptoolsself-hostingwecwec-inference
Read more →
IntermediatePart 4

Split one MCP toolbox between two agents so the scheduler cannot issue refunds

· 13 min read

Part 3 handed one agent all five tools and let a human pause guard the dangerous one. That guard depends on the model asking. This splits the same MCP server between a scheduler and a billing agent, so the scheduler cannot refund an invoice for the simplest reason available — the tool was never in its list. Then the supervisor goes back on top, and routing turns out not to be the guarantee.

aiagentslanggraphlangchainmcptoolssecurityself-hostingwecwec-inference
Read more →

Self-hosting an LLM gateway

IntermediatePart 1

One endpoint, many models: deploy an LLM gateway on a WEC Instance

· 16 min read

Every app on a box holding the same API key is a problem waiting to happen. A gateway fixes that — one endpoint, scoped keys per app, per-key budgets, and a log of who spent what. Deployed for real on a host already running five other stacks, including the parts that surprised us.

llmgatewaydockerself-hostingapilitellmwec
Read more →
IntermediatePart 2

LiteLLM complexity routing: the right model for each request, and what it costs in latency

· 12 min read

How LiteLLM's complexity router decides which model answers a request — the seven scoring dimensions, the arithmetic on a real prompt, and a measured comparison of the free keyword scorer against an LLM classifier.

llmgatewayroutingcostself-hostinglitellmwec
Read more →
IntermediatePart 3

Mask PII at the gateway: set up Presidio, plus the one line the docs leave out

· 21 min read

Set up PII masking on a self-hosted LLM gateway with Presidio. Masking can sit in three places — before the model, on the response, or only on the path to your logs — and only one keeps your app working. Follow the documented config and your model's answers come back redacted; here is why, and the single setting that fixes it.

llmgatewaypiiprivacyguardrailsself-hostinglitellmpresidiowec
Read more →
AdvancedPart 4

Clean traces, untouched answers: masking PII in LiteLLM's logs without corrupting the response

· 15 min read

Send a gateway's traffic to Langfuse and the model's own reply carries the PII straight back into your traces. Widen Presidio's scope and it masks the answer your users receive instead. Neither setting gives you both, so here is a forty-line callback that does — measured, and with the mistake that quietly makes it slower.

llmgatewaypiiprivacyguardrailsobservabilitylangfuselitellmpresidiowec
Read more →
AdvancedPart 5

Load test an LLM gateway and find the stall the median hides

· 19 min read

Part 4 measured a blocking call in a callback at 200-350ms and said the real damage only shows under concurrency. This runs it: sixteen requests at once, with a control against the upstream so you can tell your gateway's fault from the model's. The medians of the two versions are almost identical. One version drops requests and silently stops masking.

llmgatewaylitellmperformanceconcurrencypresidioobservabilitywecwec-inference
Read more →

Self-hosting Hermes

AI evals & observability

BeginnerPart 1

Evaluate your models with Promptfoo on the WEC Inference API

· 14 min read

Stop eyeballing LLM output. Build a real evaluation harness with Promptfoo pointed at the WEC Inference API — assertions, latency guardrails, JSON-schema checks, model-graded rubrics, an all-WEC model comparison, and a CI gate. Every command and result is real.

aievalspromptfooinferencetestingobservability
Read more →
IntermediatePart 2

Trustworthy JSON: schema-validate your model's structured output

· 14 min read

LLMs promise JSON and deliver markdown fences and reasoning. Build a Promptfoo eval that classifies support tickets into schema-validated JSON on the WEC Inference API — with transforms to recover messy output, a model-reliability matrix, and a CI gate. Every command and result is real.

aievalspromptfooinferencejsonstructured-outputobservability
Read more →
IntermediatePart 3

Stop hand-writing test cases: generate an eval dataset with the WEC API

· 14 min read

Two hand-typed test cases aren't an eval. Use the WEC Inference API to generate a labeled evaluation dataset — then validate and curate it, because generated labels aren't automatically correct. The result is real test data at scale; feed it to your harness and coverage surfaces the misclassifications and debatable labels two cases would hide. Every command and result is real.

aievalspromptfooinferencedatasetssynthetic-dataobservability
Read more →
AdvancedPart 4

Catch what your tests miss: observe and score your WEC app in production with Langfuse

· 18 min read

CI evals pass on a fixed test set — but production sends inputs you never tested. Self-host Langfuse on a WEC Instance, trace every real call, auto-score live traffic with an LLM judge, drill into the exact step that broke, and feed failures back to make your evals stronger. Every command — and every dead end — is real.

aiobservabilitylangfuseevalsinferenceproductionself-hosting
Read more →
AdvancedPart 5

The capstone: build, evaluate, and observe a RAG docs assistant on the WEC API

· 28 min read

Build a production-shaped RAG service in Docker: scrape a real docs site, embed locally, generate on the WEC Inference API — then catch a real hallucination, root-cause it to your own scraper, fix it, and pin it with a regression test. Every command, number, and error is real.

airagembeddingschromadbdockerevalspromptfoolangfuseinference
Read more →
AdvancedPart 6

Regression-test your RAG service with DeepEval — and settle a model debate with data

· 18 min read

Wrap the RAG assistant from part 5 in a containerized DeepEval suite judged by gemma4 on the WEC Inference API — no OpenAI key anywhere. Then use it to answer a real question: should we swap our generation model? Same quality, 5.5× the tokens: the eval says no. Every command, number, and error is real.

aievalsdeepevalragdockerciinference
Read more →
IntermediatePart 7

Component-level tracing: debugging agent tool calls

· 12 min read

Your agent's thinking and its actions are two different layers. Build a tool-calling agent on WEC Inference, trace it with Langfuse, then debug two real failures from the trace — including the confident, wrong answer that never throws a stack trace.

aievalsobservabilitylangfuseagentstracinginference
Read more →
BeginnerPart 8

Add web search to your WEC Inference calls

· 3 min read

Your RAG assistant answers great from your docs — but not about anything recent. WEC Inference now has a built-in web-search tool: add it to a chat call and the model pulls in current info. A quick before/after test.

aiinferenceweb-searchtool-use
Read more →

WhatsApp automation on WEC

Self-hosting OpenClaw

BeginnerPart 1

Deploy OpenClaw on a WEC Instance via Docker Compose

· 12 min read

From a fresh WEC Instance to a self-hosted OpenClaw agent that actually answers — Docker Compose, your own model key, and every real error and fix from a live deploy.

aiself-hostingdockeropenclawvps
Read more →
IntermediatePart 2

Secure OpenClaw with a Caddy reverse proxy + HTTPS

· 9 min read

Put Caddy in front of OpenClaw for real HTTPS and device-paired auth, then close the gateway's ports so the proxy is the only way in — gotchas and all.

aiself-hostingdockeropenclawcaddyhttpssecurity
Read more →
BeginnerPart 3

Add a Telegram channel to OpenClaw

· 6 min read

Talk to your self-hosted OpenClaw agent from your phone. Create a Telegram bot, connect it, clear the pairing gate, and chat — captured from a real run.

aiself-hostingopenclawtelegramchatbot
Read more →
IntermediatePart 4

Make OpenClaw private with a NetBird mesh VPN

· 18 min read

Put OpenClaw on a private mesh with NetBird, repoint your hostname at the mesh IP, and close the public ports — so the same chat URL works only for your devices. Real commands and gotchas from a live run.

aiself-hostingopenclawnetbirdvpnmeshsecurity
Read more →
BeginnerPart 5

Run OpenClaw on WEC Models

· 7 min read

Swap the closed provider for WiLine's own inference: point your self-hosted OpenClaw agent at WEC Models — same box, a base-URL/key/model change, no per-token lock-in.

aiself-hostingopenclawwec-modelsinference
Read more →
BeginnerPart 6

Add a WhatsApp channel to OpenClaw

· 6 min read

Put your self-hosted agent on WhatsApp. Add the channel, trust the plugin, scan a QR — and understand why a companion link makes the agent act as your account. Every command and error from a real run.

aiself-hostingopenclawwhatsappchatbot
Read more →