I’ve been building an open-source project called TraceMotive.
It started from a problem I kept running into with AI agents:
When an agent run fails, the place where the error appears isn’t always where the execution first started going wrong.
That makes debugging agent workflows harder than it looks.
So I built TraceMotive, a local-first tracing and debugging tool for AI agent execution.
What TraceMotive does
The current v0.1 includes:
- Python SDK
- canonical traces and spans
- a local Collector backed by SQLite
- a React UI for inspecting agent runs
- optional OpenAI Agents SDK integration
TraceMotive is local-first, and content capture is disabled by default.
I’m intentionally keeping the first version small. I’m not trying to add replay, automatic root-cause analysis, cloud sync, or support for every agent framework yet.
Why?
I’d rather get real feedback before adding a lot of features.
Right now I want people who actually build AI agents to try it and tell me:
- where setup is confusing
- what breaks
- what information is missing from traces
- what feels awkward in the API
The longer-term direction is:
“The causal debugger for AI agents.”
Eventually, I want TraceMotive to help identify where an agent execution first started going in the wrong direction, instead of only showing where the final error appeared.
But first, I want to make the basic observation and debugging layer solid.
Try it
PyPI:
pip install tracemotive
GitHub:
https://github.com/doraemonfv-glitch/tracemotive
If you build AI agents, I’d really appreciate you trying it for a few minutes and telling me what you run into.
Even small feedback is useful.
Top comments (2)
The local-first default is the important choice here. Agent traces get sensitive fast, especially when tool calls include prompts, file paths, or customer data. Are you planning to make redaction part of the trace schema, or leave it to each adapter?
Nice writeup. I have been testing a system prompt as an AGENTS.md contract instead of a prompt and it changed how the model behaves more than any instruction tweak. Curious how you handle long context though, that is where mine still drifts.