About four months ago, I started using coding agents, beginning with Codex. Small tasks went well. I could describe a change, inspect the result, and move on.
Longer projects felt different. As a conversation grew, it accumulated more than code: why I had chosen a direction, which alternatives I had rejected, what I wanted to leave alone, and the boundaries of the work. Some of that understanding lived only in the session.
I became reluctant to replace it. I worried that a fresh agent would lose the context behind the project, even if it could read every file. I didn't need a spectacular failure to feel that dependence.
Threshold grew out of that discomfort. I'm trying to make it possible for a project to continue when the agent session changes.
My first approach became too complicated
Initially, I tried to preserve comprehensive records and work out how a system could mechanically classify their correctness and authority. The design kept growing. I was adding machinery before I had a clear way to build and validate it.
Eventually, I made a major refactor and narrowed the question: what does a fresh session actually need to continue useful work?
Project files, Git state, and executable tests give it things it can observe again. A checkpoint can explain what the previous agent thought it had done, why it made a choice, and what remains. Those serve different purposes. A checkpoint saying that tests passed is still a report; the next agent can run them.
That doesn't make intent disappear into code. If a requirement or decision isn't represented in the files, I need to preserve it explicitly in a checkpoint or a new Task within the Project. Task instructions remain available to read again. They don't become a new system role just because they persist.
The shape I kept
Threshold is a small local service with three main concepts: Project, Task, and Run. A Project carries the longer-lived work. A Task describes something to do. A Run is an agent session working on it.
The Task stays in place while its Runs change. The second Run starts with a fresh session and inspects the saved state.
Runs are independent peers. They don't own the project or become a hierarchy of permanent supervisors and subagents. A coordinating Run should be replaceable too. Checkpoints, messages, and task state help the next Run find its bearings; it still needs to inspect the work.
Early in this direction, I deliberately used disposable Runs with a single input to test the idea. In some trials, fresh Runs continued the task; sometimes they corrected a previous checkpoint. Those were limited observations, but they were enough to keep exploring.
I also want to avoid adding entities unless they're necessary. Optional skills and plugins are selected for a Run, rather than becoming global baggage or automatically following every fresh session. The core doesn't need to understand every domain the workers might work in.
A small handoff with real code
For a concrete demonstration, a tiny CSV summary tool was built in a cloud environment using Threshold 0.2.0-alpha.7 and DeepSeek's deepseek-flash model.
The tool had a deliberately narrow job: read unquoted CSV rows containing category and amount, reject malformed input, sum amounts by category, and print JSON with alphabetically sorted keys. It needed tests and usage documentation, with no extra package dependencies.
The split was staged. The operator explicitly limited the first Run to parsing, summing, and unit tests, then deliberately started a fresh second Run with instructions to inspect the handoff and finish the CLI. The model wrote the implementation and checkpoints, and the tests produced real results. This was a controlled handoff, with the task and transition chosen in advance.
The first Run created the parsing and summing module and 17 unit tests. It ran the tests, saved a checkpoint describing the files and remaining work, and ended with the Task still marked in_progress.
The second Run used a different session on the same Project and Task. It read the Task and checkpoint, inspected Git and the actual files, and ran the existing tests. Its baseline was 17 passing tests. It then added the CLI, six integration tests, and usage documentation before saving another checkpoint and marking the Task done.
A separate check after that Run finished reported 23 passing tests and no failures. Running the CLI on the sample returned books: 20 and coffee: 3.25. The implementation stayed uncommitted, so this handoff also depended on inspecting the working tree, including untracked files.
The useful detail for me is the baseline check. The second Run had a previous agent's account of the work and an opportunity to compare that account with the project in front of it.
What this leaves open
This was one small, deliberately staged handoff. It gave me a concrete example of a fresh Run checking the work and continuing it. It does not yet tell me how reliably that will work on larger projects, especially when decisions change or a checkpoint leaves something important out.
The recording below replays real timestamped TUI output, with edited timing and labeled playback speeds. The Runs were started through the CLI while the TUI observed them.
Watch the two-Run CSV handoff on Reddit
I'm sharing Threshold while these questions are still open. If you've also hesitated to replace a long-running coding session, I'd be interested in what context made you hesitate, and where a fresh Run loses its footing when you try it.
The project is on GitHub: https://github.com/Key-of-door/Threshold
AI assistance disclosure: AI helped turn my account of building Threshold and the recorded demonstration into this English draft.

Top comments (0)