Skip to main content

7 posts tagged with "langfuse"

View all tags
intermediatePart 3

One login for everything: putting authentik in front of a self-hosted app

· 20 min read
Rafael Fernandes
NLP Engineer & Tech Writer at WiLine
Share:
authentik+

Part 1 closed the ports nobody meant to open. Part 2 stopped two containers from running as root. Both were about the machine. Neither touched the question a self-hosted AI stack answers worst: who is allowed to log in, and where is that decided?

Same box as before. The Langfuse we've been hardening since Part 1 — the one whose Postgres, ClickHouse and Redis Part 1 found correctly bound to localhost, and whose worker Part 2 dropped off root — was deployed back in Catch what your tests miss. Two parts have now secured the machine underneath it without once touching the application's own front door. This is the first part that changes Langfuse itself.

Right now that decision is made in each app, separately. Langfuse has its own email-and- password table. So does every other tool on the box. Each one is a place where an account can outlive the person who owned it, where a password can be reused, and where "remove this person's access" means remembering that the app exists at all.

This post moves that decision to one place. That place is authentik — an open-source identity provider you run yourself, the self-hosted counterpart to Okta or Auth0. It holds the accounts, runs the login screen, and vouches for who someone is to any app that asks. Apps stop storing passwords and start asking authentik.

It's the same job Keycloak does, and Keycloak is the better-known name. authentik earns the pick here on setup cost: a compose file and a wizard against Keycloak's realms, clients and JVM tuning. For one box and one app, that difference is the whole decision.

We deploy it, connect Langfuse to it over OIDC, and end with a login that goes through authentik and back.

intermediatePart 2

Hand work between LangGraph agents without corrupting shared state

· 16 min read
Rafael Fernandes
NLP Engineer & Tech Writer at WiLine
Share:
+PostgreSQL+
0/4
🎯 Skill path0/4 earned
Agent orchestration with LangGraph

You have two agents. One books appointments. One handles billing.

A ticket arrives that needs both: reschedule my install, and my bill looks wrong. So you send it to both.

The first one books Tuesday. The second one sees the account is past due and freezes it. Both finish at almost the same moment, and both save what they decided.

Only one of them is saved. Which one? Whichever finished first — which comes down to how slow an API call was that day. So you book an appointment on a frozen account, or freeze an account you just promised an engineer to. Afterwards it looks like one clean decision was made.

Part 1 built one agent that runs its steps in a fixed order. This post has several: a supervisor that picks who works on what, an agent that passes the job to another one halfway through, and two agents put in each other's way on purpose — to find out what LangGraph does when they disagree.

intermediatePart 1

Checkpoint a LangGraph agent on a WEC Instance so crashes cost you nothing

· 20 min read
Rafael Fernandes
NLP Engineer & Tech Writer at WiLine
Share:
+PostgreSQL+
0/4
🎯 Skill path0/4 earned
Agent orchestration with LangGraph

Your agent has been working on a customer's request for forty seconds. It has read the ticket, pulled the account, and checked three engineers' calendars. It is about to book the appointment.

Then the process dies. A deploy going out, the host running out of memory, someone restarting the container — it does not matter which.

The agent does not pick up where it left off, because there is nothing to pick up from. Everything it learned lived in variables inside a process that no longer exists. The customer is still waiting. Run it again and you pay for all that work a second time. And the part that should worry you most: nobody can say whether the appointment was booked in the last second before it died.

An agent is a model in a loop. As an ordinary Python script, that loop is exactly as fragile as the process holding it.

Here we build one that saves its state to Postgres after every step, so a crash costs nothing.

advancedPart 4

Clean traces, untouched answers: masking PII in LiteLLM's logs without corrupting the response

· 15 min read
Rafael Fernandes
NLP Engineer & Tech Writer at WiLine
Share:
LiteLLM+Presidio+
0/5
🎯 Skill path0/5 earned
Self-hosting an LLM gateway

Part 3 configured Presidio to mask prompts at logging time and deliberately stopped short of proving it. Nothing had been wired to a logging destination yet, so there was no trace to inspect.

This post wires one up, finds <PERSON> where it should be — and then finds the customer's real email address sitting a few lines below it, in the model's reply.

intermediatePart 7

Component-level tracing: debugging agent tool calls

· 12 min read
Rafael Fernandes
NLP Engineer & Tech Writer at WiLine
Share:
+
0/8
🎯 Skill path0/8 earned
AI evals & observability

An agent that calls tools does two very different things: it reasons ("I should search the docs") and it acts (actually calls the tool). When something goes wrong, the first question is always which layer failed — did it think wrong, or did the doing break? A flat log can't answer that. A trace can.

In this tutorial you build a small tool-calling agent on WEC Inference, instrument it with Langfuse so every reasoning step and every tool call becomes an inspectable node, then debug two real failures from the trace tree — including the worst kind: a confident, wrong answer that never throws an error. Every command, error, and screenshot below is from a real run.

advancedPart 5

The capstone: build, evaluate, and observe a RAG docs assistant on the WEC API

· 28 min read
Rafael Fernandes
NLP Engineer & Tech Writer at WiLine
Share:
+Promptfoo+
0/8
🎯 Skill path0/8 earned
AI evals & observability

Everything in this series was building to this. You can prove a model works (part 1), make its output machine-reliable (part 2), generate real test data (part 3), and observe production (part 4). Now we spend all four skills at once on the pattern behind almost every serious LLM product: RAG — retrieval-augmented generation.

We'll build a docs assistant: a containerized HTTP service that answers customer questions from WEC's own documentation. Not a notebook — a service, Docker-first, the shape you'd actually deploy. And because this series doesn't do happy-path demos: along the way our RAG hallucinates a GPU price, we root-cause it to our own scraper, fix it, and pin the fix with a regression test. Every command, number, and error below is from a real run.

advancedPart 4

Catch what your tests miss: observe and score your WEC app in production with Langfuse

· 18 min read
Rafael Fernandes
NLP Engineer & Tech Writer at WiLine
Share:
+
0/8
🎯 Skill path0/8 earned
AI evals & observability

A customer says your support bot promised them a refund policy that doesn't exist. Your feature made two LLM calls — classify, then reply. Which one invented it? If you can't answer that, your app is a black box — and this guide fixes exactly that.

You can now prove a model works (part 1), make its output machine-reliable (part 2), and generate a real test set to check it against (part 3). But all of that runs offline, in CI, on inputs you chose. Production doesn't play along.

This guide closes the gap. We'll self-host Langfuse — the open-source, self-hostable alternative to LangSmith — trace every real call, auto-score live traffic with an LLM judge, drill into the exact step that breaks, and feed failures back so your part-3 dataset gets stronger. Offline eval tells you it worked on your test set; this tells you it works in the wild.