DEV Community

Cover image for The decision was not in a doc. It was in a Claude Code session on my cofounder's laptop
Sid Probstein
Sid Probstein

Posted on

The decision was not in a doc. It was in a Claude Code session on my cofounder's laptop

Three weeks ago my cofounder recommended we not do monthly license keys.

I could not have told you that this morning. I knew the decision. I could not have told you who argued for it, when, or why. It was never written in a doc. It happened inside a Claude Code session on his laptop, on a Tuesday, while he was reviewing two pull requests about something else.

So this morning I searched for it. One query, across my team's history, from my own laptop. Three results, eleven days apart, in order.

Tokenome search across a team's shared AI history. The query is

22 September, 2:02pm. John O'Neil: "The one design change I recommend is long-lived keys with cancellation-driven expiry, instead of monthly keys with a 31-day overhang." The decision came to me the same afternoon.

3 October. The settled model: "No monthly keys, no refresh, no fetch token, nothing phones home."

That is the whole chain. Nobody wrote it down, because nobody knew at the time which conversation would turn out to matter. Git told me the key format changed. It did not tell me that John had weighed it against the monthly form and said the long-lived one would annoy fewer customers and be no easier to pirate.

This is the normal state of things now. The reasoning behind your product lives in chat logs on individual machines, in one vendor's format, in one person's account. A reviewer cannot see what the agent was told. A new hire cannot find why the retry logic looks the way it does. And an agent in the next session re-derives, at full price, a conclusion somebody already reached.

We built Tokenome for that. Here is what is underneath it, because I would want to know before installing anything.

What gets captured, and when

Four sources today, each arriving differently.

  • Claude Code is read from ~/.claude/projects automatically. If you install the Claude Code plugin, a session hook also indexes the transcript when the session ends, so today's work is searchable tomorrow.
  • Claude Desktop and Cowork are pulled from the claude.ai API with your own session key, when you set one. This is the one source that makes a network call on your behalf, and it is to your own account.
  • ChatGPT and Gemini arrive as exports: drop ChatGPT's conversations.json, or Gemini's flat JSON, into their import folders.

Threads are split into question and answer documents, each stamped with time, project, platform and model. Tool actions are captured too, not only the chat around them, which matters more than it sounds: the edit is often where the reason is. Only new or changed turns get re-embedded, so a live session costs roughly one turn of work per poll.

Five stages, in order: capture, segment, embed, index, serve. Embedding is on-device ONNX, no GPU. The index is Typesense, with BM25 fused with vector similarity. Serving is a CLI, a local web UI and an MCP server.

Everything it creates sits under ~/.tokenome: the index, the journal, the drop folders and the local API token.

Local by construction, not local by policy

The Python runtime, the search engine and the embedding model are all inside the download. There is nothing to fetch on first run, no account, no sign-in. Embeddings are computed on your own hardware, so no text is sent anywhere to be turned into vectors.

The one thing that reaches the network on the free tier is the desktop app's update check against the public releases page, which carries no conversation data. The other two exceptions are the ones you turn on yourself: the Claude Desktop source above, and Team, which sends only the projects you opt in.

That is a design choice, not a measurement. Nobody audited it for us. Check it with a packet capture if you are the sort of person who would.

The reason it is not negotiable for us: a transcript corpus is the most sensitive text a developer owns. Your repo has been read by every reviewer on the team. Your transcripts have the stack trace you pasted with a token still in it, the customer name you said out loud while debugging, and the architecture you rejected and do not want quoted back at you. A searchable, embedded, cross-tool index of all that is either the best thing on your laptop or the worst thing in somebody's cloud, depending on one decision made early.

tokenome search "type hints" --since 30d
tokenome search "license keys" --project tokenome --json
Enter fullscreen mode Exit fullscreen mode

From a line of code to the conversation behind it

This is the feature I use most.

tokenome context src/ledger.py 31
Enter fullscreen mode Exit fullscreen mode

It blames the line with git blame to learn when it was last changed and by whom; if the file is not reachable now, it falls back to the blame captured when the conversation was indexed. It takes that timestamp, finds the conversations from just before the edit, and prints the committer, the time, the commit, and the conversations, each one openable in full. --window N changes the minutes considered, default 30. --symmetric looks on both sides of the edit, for when the edit landed some time after the conversation that drove it.

The boring case is a line written once and never touched. The interesting case is a reversal: the line that used to be something else. That is where the gap between what and why is widest. The commit message says "go back to direct account lookups". It does not say that the cache could serve a stale account and misroute money, and that the latency win was not worth it. That sentence exists. It was typed, by a person, at the moment it was true. It is just not in the repo. Walking back from the line to the half hour before the commit is the cheapest way I know to get it back.

A good result has three parts: a blame that points at a real commit, at least one conversation inside the window, and a reason in that conversation that reads like a decision rather than a status update.

The MCP server, which is the part I would care about

Tokenome runs an MCP server on your machine, so the agent asks your history itself instead of you pasting context into it. The endpoint is served by the running service at http://127.0.0.1:8741/mcp; tokenome mcp proxies it over stdio for clients that need that, including Claude Desktop.

In Claude Code it arrives as a plugin, from the shell or with /plugin marketplace add tokenome/releases inside a session:

claude plugin marketplace add tokenome/releases
claude plugin install tokenome@tokenome

# or, from the app, which also registers the MCP server with this machine's token
tokenome claude install
Enter fullscreen mode Exit fullscreen mode

The tools an agent gets:

Tool What it answers
why_was The stated reason behind a decision, with the quote that proves it. It searches the journal, so the answer is the reason rather than a turn that happens to mention it.
ask_faq A recurring "why" question, from the running FAQ, with its evidence and its answer history.
search_memory Why a past decision was made, what was tried, who agreed to what. The general entry point.
get_journal What was done on a day, and why, with the quoted receipts.
get_conversation A past thread, turn by turn, as it actually unfolded.
get_document One turn in detail, with the exact wording.
find_related Other times a topic came up, starting from a result already in hand.
find_conversations_near What was being discussed when a change was made. Pass a commit or a blame timestamp.
browse_by_label What a particular model or tool was used for.
get_recent What has been worked on most recently, across every tool.

why_was is the one that changes how a session feels. It returns the decision and the reversal that followed it, each with the conversation it came from, so the agent sees that the question was settled, then unsettled, and in which order. An agent that can see a reversal stops confidently re-proposing the thing you already tried.

The plugin also adds a skill that loads on demand and teaches the agent which tool answers which question and how to chain them, a prompt hook that nudges it to look in history first on retrospective prompts, the session hook mentioned earlier, and /tokenome:why and /tokenome:search. The hooks are plain HTTP calls to the local service, authenticated with the token file, so they work whether or not tokenome is on your PATH. Results come back compact, with ids the agent can expand, so a lookup costs a fraction of pasting a transcript in.

Now the number, which only means anything with its qualifier attached. In a controlled evaluation of forty sessions, ten questions asked four ways over one corpus, the recommended configuration (routing by question type) answered questions about past work with 38% fewer context tokens and no measurable change in answer quality.

Two things that are easy to get wrong about that figure. It is context tokens, not money: dollar cost fell 22 percent in the same arm, a smaller and separate number, because cached context is cheap to re-read. If you are estimating your bill rather than your context, use 22. And "no measurable change in answer quality" is the whole claim. The recall differences between configurations sat inside the measurement noise, in both directions. Better answers is not established, and I am not claiming it.

The mechanism is not what people assume. A single index lookup returns about three times more text than a single grep, roughly 3,000 tokens against 1,000. The saving comes from needing fewer operations: 241 filesystem operations in the baseline became 136 under routing, and agent turns fell from 252 to 161. Fewer lookups, not smaller answers.

The team server

Searching my own history would not have found John's recommendation. It was on his laptop.

Team is self-hosted. One container image, ghcr.io/tokenome/tokenome-server; the compose bundle on the releases page pins the version and brings up the index, the server and automatic TLS, and on Railway it is the same image as one service with a volume. We do not run one for you.

How it behaves, in the order it happens:

  1. Enrollment is per device. tokenome remote enroll generates a device keypair locally and sends only the public half. Enrollment tokens are single use and expire in 72 hours. A second device for the same person never uses a second seat.
  2. Nothing is shared until you share it, by name, per project. tokenome remote share <project> --backfill. There is no default-on, and no "share everything".
  3. Redaction happens on the laptop, before the send. The share gate replaces secrets and personal data before anything leaves the machine, and works from a names dictionary you keep. What it cannot make safe it holds back and shows you, to release or discard, instead of shipping it and hoping.
  4. Every team result carries the name of the person it came from. The three results in the screenshot above are John's, and they say so.
  5. The audit log is append-only, filterable and exportable. A search row never holds the query text, only a hash, a category and the names of the filters used.
  6. Retraction is real. tokenome remote unshare <project> stops the sharing; --purge removes what was already sent. Wiping your local index is index maintenance and is deliberately not mirrored to the server, so a reindex does not quietly delete the team's history.

The honest part: the server holds shared text and vectors in the clear, because that is what makes search possible. The protection there is organizational, not cryptographic. It is your server, inside your boundary, with your retention windows on it.

Team is in beta and free while the beta runs.

What it does not do

A post with no limits in it is an ad. Here are mine.

  • Code provenance needs a git repository and a committed line. An uncommitted line blames to your working copy and has no timestamp to search around.
  • It needs the conversation to be indexed. Work done before you installed it, or in a tool Tokenome does not read, is not there. Drop in an export and it will be.
  • It finds conversations near the edit in time. That is not proof of cause. Two things discussed in the same half hour both come back. You read them and decide; the tool does not assert that one caused the other.
  • It does not summarize. You get the turns, with their exact wording, because the detail that matters is usually a specific number, id or model name that a summary would drop.
  • A line nobody discussed returns nothing. That is the honest answer and it is worth having: it tells you the decision was never actually made in writing.
  • The evaluation is one careful measurement, not a benchmark suite. One corpus, one project, one grader. Per-query results under routing ranged from -87% to +109%, and two of ten queries cost more than the baseline; broad "what did we conclude" questions win, narrow questions with an obvious file to open do not, which is what the routing is for.
  • Agents drop facts they already hold. Roughly half of all missed facts, in every configuration we tested, were sitting in a tool result the agent had already received and never wrote down. An index does not fix that.
  • There is no Windows build. Not built, not signed, not tested. WSL2 runs the Linux wheel, but capture from Windows-side tools is untested, so I will not call it supported.
  • Several things people ask for are not built, including a pull request provenance bot, a team FAQ folded from everyone's journals, and a provenance API. They are on the roadmap and that is all they are.

Trying it

Free on one machine, no account, nothing to sign up for. The command line on its own:

uv tool install tokenome-ai
Enter fullscreen mode Exit fullscreen mode

Then the two plugin commands above, if you want Claude Code to have the tools. The desktop app is a signed and notarized DMG for Apple Silicon and an AppImage or deb for Linux on x86_64 and aarch64, both on the public releases page with a .sha256 beside each download. The app installs the CLI for you from Settings, and tokenome app opens the local web UI at localhost:8741.

tokenome.ai

Top comments (0)