I’m launching Datool, an open-source alternative to Langfuse for tracing and evaluating AI applications.
The workflow I want to make easier is the one after you find a bad answer: keep the example, turn it into a test, try a change, and check whether the result improved.
From a Bad Answer to Report showing how much it was improved
Imagine a support agent telling a customer that returns are allowed for 60 days when the supplied policy says 30. A trace helps you see the input, output, and steps that produced that answer. The next step is keeping that case so the same mistake can be tested again.
In Datool, you can:
- Inspect a trace and its nested spans.
- Save real examples as dataset cases with expected outputs.
- Define scorers in JavaScript or Python and test them against recorded outputs.
- Run evaluations and compare a candidate with a baseline.
- Inspect the individual cases behind the scores.
ALL VIA MCP.
Prompt versions keep the change identifiable. Dataset snapshots help keep the cases consistent between runs. The CLI and MCP interface let these workflows fit into development and agent tooling.
What’s included
- Traces and spans for AI workflows
- Versioned prompts
- Datasets and evaluation runs
- Custom scorers and saved result views
- Self-hosting under Apache 2.0
The launch video uses an illustrative support-agent example to show how these pieces connect.
Watch the 46-second launch video.
Try it or inspect the source
I’d like feedback from people building AI apps: how do you turn production failures into regression tests today, and where does that process get stuck?
Top comments (0)