Every time I run Claude Code, it quietly writes the entire session to disk. No flag, no setup, no instrumentation. The file is already there when I want it. That single fact is the reason I have been able to build small analysis tools on top of my own agent runs without ever touching the agent itself. In this tutorial I want to show you how to read those files and pull real structure out of them: the turns, the tool calls, and the results those calls produced.
I learned this the boring way, by reading a real 34MB transcript line by line, so everything below is grounded in shapes I actually saw on disk rather than shapes I hoped were there.
Where the files live
Claude Code stores one file per session under:
~/.claude/projects/<project-slug>/<session-id>.jsonl
The slug is your project path with slashes turned into dashes. The file is JSONL: one JSON object per line, appended as the session runs. That append-only nature matters later.
Finding them is a one-liner:
from pathlib import Path
def find_sessions(root: Path | None = None) -> list[Path]:
root = root or (Path.home() / ".claude" / "projects")
if not root.exists():
return []
return sorted(root.glob("**/*.jsonl"))
Reading lines defensively
The first thing you learn reading a live transcript is that the last line can be torn. If Claude Code is mid-write while you read, the final line may be half a JSON object. The right move is not to crash the whole analysis over one bad line, it is to skip it:
import json
from typing import Iterator
def records(path: Path) -> Iterator[dict]:
with path.open(errors="replace") as fh:
for line in fh:
line = line.strip()
if not line:
continue
try:
yield json.loads(line)
except json.JSONDecodeError:
# A transcript being appended to while we read it can produce a
# torn final line. Skipping it lets us still analyse a live session.
continue
What a record looks like
Not every line is a conversation turn. In one real file the type field held values like user, assistant, queue-operation, attachment, file-history-snapshot, mode, and a few others. The two you care about for turns are user and assistant. Each of those carries a message object, and inside it a content field that is a list of typed blocks.
The blocks are where the interesting stuff is. A text block looks like {"type": "text", "text": "..."}. A tool call is a tool_use block. A tool's output is a tool_result block. So the mental model is:
- lines give you turns,
- content blocks inside a turn give you text, tool calls, and results.
Extracting tool calls and results
A tool_use block has these keys: type, id, name, input, and sometimes caller. The id is the anchor. When the tool finishes, a later line carries a tool_result block whose tool_use_id matches that id. That pairing, call to result by id, is the whole game.
Here is the core extraction. I make two passes: collect every call and every result keyed by id, then join them. Two passes because results do not always appear in a tidy order relative to calls, and scanning a file twice is cheap compared to getting the pairing subtly wrong.
def tool_calls(path: Path) -> list[dict]:
uses: dict[str, dict] = {}
results: dict[str, dict] = {}
for rec in records(path):
msg = rec.get("message") or {}
content = msg.get("content")
if not isinstance(content, list):
continue
for block in content:
if not isinstance(block, dict):
continue
btype = block.get("type")
if btype == "tool_use":
uses[block["id"]] = {
"name": block.get("name"),
"input": block.get("input") or {},
"ts": rec.get("timestamp"),
}
elif btype == "tool_result":
tid = block.get("tool_use_id")
if tid:
results[tid] = {
"content": block.get("content"),
"ts": rec.get("timestamp"),
"is_error": bool(block.get("is_error")),
}
joined = []
for tid, use in uses.items():
res = results.get(tid, {})
joined.append({
"id": tid,
"tool": use["name"],
"input": use["input"],
"output": flatten(res.get("content")),
"started_at": use.get("ts"),
"ended_at": res.get("ts"),
"is_error": res.get("is_error", False),
})
return joined
The one shape that will bite you
tool_result content is not always a string. Sometimes it is a plain string, and sometimes it is a list of typed blocks like [{"type": "text", "text": "..."}]. If you assume one form, half your results come out as None or as ugly repr strings. Normalize it:
def flatten(content) -> str:
if isinstance(content, str):
return content
if isinstance(content, list):
parts = []
for c in content:
if isinstance(c, dict):
parts.append(c.get("text") or json.dumps(c))
else:
parts.append(str(c))
return "\n".join(parts)
return "" if content is None else str(content)
Timestamps and duration
Records carry an ISO-8601 timestamp with a trailing Z. Python's fromisoformat wants an explicit offset, so swap the Z before parsing. Once a call and its result both have timestamps, the difference is your tool latency, no profiler required:
from datetime import datetime
def ts(value: str | None) -> datetime | None:
if not value:
return None
try:
return datetime.fromisoformat(value.replace("Z", "+00:00"))
except ValueError:
return None
Putting it to use
Once you have this, small tools fall out of it almost for free. In agentrace I filter tool calls down to the ones named Agent (subagent delegations), pair each delegation's prompt with the report that came back, and time how long each subagent ran. In ctxlens I walk the same turns to see how context accumulates across a session. Neither tool needed any hooks or wrappers around Claude Code. The data was already on disk; I just had to read it correctly.
An honest caveat
This is an undocumented, internal format. It is not a stable public API, and it can change between Claude Code versions. New record types can appear, keys can be added, block shapes can shift. That is exactly why the code above is defensive at every step: skip lines it does not understand, guard every .get, and normalize content that comes in more than one shape. Write your parser to tolerate the unexpected rather than to assume today's shape is forever, and it will survive most format drift. When something does change, the fix is usually a print loop over type and block keys to see what is new.
If you want a worked example of this technique applied to subagent observability, my project agentrace does exactly that on top of these same files: github.com/AgentPostmortem/agentrace. Clone it, point it at your own ~/.claude/projects, and you will be reading your agent's history in a couple of minutes.
Top comments (0)