DEV Community

Cover image for Self-hosted Langfuse: tracing 7% of my AI agents, and ClickHouse logging itself
Christian Anderson
Christian Anderson

Posted on

Self-hosted Langfuse: tracing 7% of my AI agents, and ClickHouse logging itself

On Discord I told my agent post followed by a comment id. That's the approval step for a dev.to reply it had drafted: read the draft, say "post", done.

Nothing got posted. The turn took 25 minutes and came back with a confident, detailed summary of the article. It was made up.

I self-host Langfuse to trace my agents, so I went and looked at the trace. Then I checked what else it had seen. It had been watching about 7% of my agents' traffic.

The 25-minute turn

Here is the trace of that one turn:

Hermes turn                      1526.3 s
  LLM call 1   qwen3.5-64k        569.4 s
  Tool: web_search                  0.8 s
  Tool: web_search                  0.9 s
  LLM call 2   qwen3.5-64k        207.4 s
  Tool: web_extract                 2.9 s
  LLM call 3   qwen3.5-64k        740.2 s
Enter fullscreen mode Exit fullscreen mode

Tools: 4.6 seconds in total. Model: 1,517 seconds. The local model spent 25 minutes thinking its way to the wrong answer, with three short trips to the web in between.

Without the trace I'd have blamed the web tools. A turn that searches the web and takes 25 minutes feels like a network problem. The trace says the network was the fastest thing in the room.

The actual cause had nothing to do with speed. The Discord conversation was one long-lived session, started days before the reply skill existed. The agent reads its list of skills once, at session start. So as far as that session knew, there was no skill for posting a reply. Asked to "post" something it had no tool for, it did what models do and improvised: searched the web for the article, extracted the page, and wrote a plausible summary instead.

The fix was a fresh session. But it made me curious about what else the traces could tell me.

What the traces said about models

Pulling latency per model from the same Langfuse data (the default profile, all history):

Model Calls p50 Slowest
qwen/qwen3.7-flash 92 4.3 s 180.5 s
qwen3.5-64k (local, 6 GB GPU) 91 125.2 s 900.3 s
deepseek-v4-flash 65 6.7 s
mistral-nemo (local) 24 106.6 s 1,802.9 s
hermes3-64k 15 2.5 s

On the tool side, terminal ran 135 times with a median of 0.6 s, and the slowest single tool call was a search_files at 120.6 s.

None of this is shocking. A local model on a 6 GB card is slow, and cheap cloud models are fast. But "slow" and "125 seconds median, about 30 times slower than the fastest cloud option" are different sentences, and only one of them helps decide which job runs where.

Then I noticed the call counts. 91 calls for the local model. That seemed low for a system I lean on all day.

The coverage check

Langfuse held 618 events across 79 traces in four weeks. That's about three traces a day.

The agent keeps its own session database, so I counted from the other side. The last 14 days alone:

  • 188 sessions
  • 1,387 model calls
  • 10 agent profiles

Tracing was enabled on exactly one of ten profiles: the default chat one, with 91 calls. That's roughly 7% of the traffic.

The untraced ones were the ones doing the work:

Profile Model calls (14 days) Traced
coder 522 no
sec 390 no
itadmin 172 no
writer 120 no
default (chat) 91 yes
reviewer 55 no
scout 29 no
others small no

The coding agent alone made more than five times as many calls as the only profile I was watching.

Why it happened

The tracing plugin is enabled per profile, and it reads its Langfuse API keys through that profile's own secret scope. That's deliberate: profile B's traces shouldn't be shipped with profile A's keys.

When I set Langfuse up, I enabled the plugin and added the keys on the default profile, and moved on. Every specialist profile then ran without tracing. No error. No warning in a log. The plugin fails open by design (no keys means its hooks quietly do nothing), which is right for a plugin, and exactly why nobody noticed.

A monitoring tool that's switched off looks identical to one that has nothing to report.

There was a second, smaller gap. One-shot jobs launched in "safe mode" skip plugins entirely, and the reply drafter ran that way because it's quicker. That one is closed now: the drafter sends its own traces through the Langfuse SDK, so it no longer depends on the plugin.

The second surprise: the disk

Langfuse v3 onwards stores traces in ClickHouse. Postgres holds users and projects, Redis handles the queue, MinIO stores blobs. So no, you can't drop ClickHouse to save resources; it is the trace store.

ClickHouse was using 6.03 GiB of disk. The actual Langfuse trace data in it was 2.2 MiB.

The rest was ClickHouse's own diagnostic logging:

System table Size Rows
system.trace_log 3.38 GiB 174.6 million
system.text_log 1.19 GiB
system.part_log 492 MiB
system.asynchronous_metric_log 485 MiB 949 million
system.metric_log 461 MiB

That's roughly 2,700 times more bytes about ClickHouse than about my agents.

Those tables aren't written once and forgotten. They're written continuously. The node this runs on keeps its guests on a single 5400 rpm HDD, and it already has a roughly 15-minute IO storm after every reboot. The Langfuse container is the one I stop to relieve it. So the trace store I'd barely been using was grinding that disk all day, recording its own profiler samples.

The fix

Tracing on every profile

For each of the ten profiles:

  1. Copied the Langfuse keys into that profile's own env file (keeping the per-profile secret scoping).
  2. Enabled the plugin.
  3. Set the Langfuse environment to the profile name.

The third step is the one that matters for later. Traces carry no profile tag of their own, so without it every trace lands in one undifferentiated pile. With it, "cost and latency per agent role" is a filter in the UI rather than a spreadsheet exercise.

Then I restarted the gateway and sent a one-word probe on two profiles. Traces arrived tagged legal (on a cloud flash model) and writer (on the local qwen3.5-64k). That's the test I should have run the first time.

ClickHouse logging itself

ClickHouse lets you override server config with a file dropped into /etc/clickhouse-server/config.d/. This one removes twelve system log tables, keeps query_log with a 7-day TTL (it's genuinely useful when a query is slow), and turns the server logger down to warnings:

<clickhouse>
  <trace_log remove="1"/>
  <text_log remove="1"/>
  <metric_log remove="1"/>
  <asynchronous_metric_log remove="1"/>
  <part_log remove="1"/>
  <processors_profile_log remove="1"/>
  <query_thread_log remove="1"/>
  <query_views_log remove="1"/>
  <query_metric_log remove="1"/>
  <opentelemetry_span_log remove="1"/>
  <session_log remove="1"/>
  <latency_log remove="1"/>
  <query_log>
    <ttl>event_date + INTERVAL 7 DAY DELETE</ttl>
  </query_log>
  <logger><level>warning</level></logger>
</clickhouse>
Enter fullscreen mode Exit fullscreen mode

It's mounted read-only into that directory through the compose file, and the container was recreated. Checks afterwards:

  • container healthy
  • trace_log row count unchanged over 30 seconds, so it has stopped writing
  • trace data intact
  • Langfuse web health check returns 200

One honest caveat: removing a log table from config stops ClickHouse writing to it, but it doesn't delete what's already there. The old 5.9 GiB of log tables was still on disk. I dropped those deliberately, as a separate step, rather than bundling a destructive delete into a config change. The Langfuse container went from 15 GB of disk used to 7.7 GB, with the trace data intact.

Why I want this at all

It would have been easy to conclude Langfuse was dead weight: a few gigabytes of disk for 79 traces.

But the estate runs around ten agent profiles: a chief of staff that I chat with, a coder, security, IT admin, writer, scout, reviewer, legal, analyst and a few more. They run on a mix of local models on a 6 GB GPU and cheap paid cloud models with a fallback chain. The questions I actually have about it are ones only traces answer:

  • Which role burns the calls? Coder, at 522 of 1,387.
  • Where does a slow turn's time go? Model or tool. See the 25-minute turn, which looked like a web problem and was a thinking problem.
  • Did the fallback fire? A cloud model failing over to another is invisible from the outside if the answer still arrives.
  • Local or cloud for this job? 125 seconds median against 4 to 7 seconds is a real trade, and it should be made per job, not by habit.
  • What does each role cost? Now that traces are tagged per profile, that's a query.

The rule I run the estate by is "scripts gather facts, models never do". Models get to summarise what scripts measured; they never go and measure it themselves. Langfuse is the fact-gatherer for the agents themselves. It just can't gather facts from agents it isn't connected to.

What the weekly report tracks now

A small script now reads the agents' own session records and Langfuse side by side, and every Monday it posts a report: per role, the turns, model calls, tokens, cost, errors and the share that was traced; per model, p50 and p95 latency and cost; how turn time splits between model and tools; the five slowest turns; and tool errors. It's built to answer these questions:

  • Does the coder profile's call volume hold up, and is it doing many short turns or a few enormous ones?
  • How often do the local models hit their multi-minute tail, and on which profiles?
  • How often does the cloud fallback chain actually fire, and for which model?
  • Which one or two jobs account for most of the cost, and would they be fine on a local model?
  • Does ClickHouse stay small now that it's only storing what I asked it to?

If you self-host Langfuse, check these three things

  1. Is tracing actually enabled everywhere you think it is? Per service, per profile, per worker. Count calls from your application's side and compare with Langfuse's trace count. If one is an order of magnitude bigger than the other, like mine, you'll know.
  2. How big are ClickHouse's system log tables? Check system.parts grouped by table. If trace_log and asynchronous_metric_log dwarf your actual data, a config.d override fixes it in a few lines.
  3. Send a probe trace from every agent. One word, one turn, per profile, and confirm each one lands with the right tag. A plugin that fails open will never tell you it isn't there.

Postscript: it happened again

On 22 September the tracing server's address changed, and all 12 specialist agents' configs were repointed at the new one. When it moved back on 23 September, only the main config was updated. Every specialist silently stopped tracing. I found and fixed it on 24 September.

No error, no warning, nothing in a log. It's the same fail-open silence this whole post is about, and the reason point 1 above isn't a one-off check.


πŸ€– Drafted with AI assistance from my own homelab notes, logs and repos, then reviewed and edited before publishing.

Top comments (6)

Collapse
 
aifrontierpost profile image
AI Frontier Post •

Tracing 7% of your agent traffic while the coder profile's 522 calls ran completely dark means the dashboard was measuring the quietest part of your system. Partial observability that reads as complete coverage is arguably worse than none, because it calibrates your confidence wrong.

Collapse
 
c1-anderson profile image
Christian Anderson •

Yeah, that's exactly what tripped me up, and it happened again on 23 September when the server moved back to its old IP. Only the main config updated, so every specialist agent stopped tracing silently for a whole day. After that I made sure to send a probe turn on each profile, rather than trusting the plugin to tell me it was working. How do you keep your traces honest without that kind of manual spot-check?

Collapse
 
reidmarlow profile image
Reid Marlow •

The fail-open behavior in agent telemetry plugins is such a quiet trap. When a hook fails open to protect agent availability, silence looks identical to a healthy idle queue.

The only way I kept my local agent tracing honest after a similar outage was moving the verification to the gateway rather than the client config. I wrap the upstream router to reject model turns if the active profile lacks a valid trace session header, and run a scheduled synthetic heartbeat every two hours to confirm the trace store actually wrote records. Catching broken endpoints at the proxy layer keeps you from finding out three weeks later that your heaviest background workers ran completely dark.

Also, seeing ClickHouse system tables chew through six gigabytes on a 5400 rpm drive brought back painful homelab memories. Trimming the trace log overrides before that container churns IO into the ground is mandatory hygiene.

Collapse
 
c1-anderson profile image
Christian Anderson •

Wrapping the upstream router is a level cleaner than what I did. I went profile by profile copying keys in, and my verification was sending a one-word probe on exactly two of them. A synthetic heartbeat at the proxy layer would have caught the September address change in hours instead of me finding it the next day.

Collapse
 
prpatel05 profile image
Pratik Patel •

The fail-open-per-profile bit is the part that bites. We had the same shape: keys on the default surface, specialists running dark, and the dashboard looking healthy because it only ever saw the traced slice. What finally caught us was a coverage check β€” traced calls over total model calls by profile β€” not another Langfuse panel. Curious whether you gate deploy on that ratio now, or only spot-check.

Collapse
 
c1-anderson profile image
Christian Anderson •

Nah, not something I covered, it's still a weekly spot check, on purpose. A deploy gate only shifts when you find out, and the gap that's bitten twice is the kind no panel would catch: tracing address changed, every specialist went dark, nothing so much as a warning in any log.