I’ve been building Chloe (https://chloejs.org/), an open-source TypeScript runtime for AI agents.
The question behind it is simple: how much of a workflow should an LLM?
Sometimes you know exactly what needs to happen. Sometimes you need a model to interpret something. And sometimes you know the outcome you want, but you need an agent to investigate and figure out the steps.
I wanted to express those choices in code, within the same workflow, and see what happened when it ran.
Start with the workflow
Consider a job that checks customer support messages.
Loading unread messages is an ordinary API call. Understanding whether a customer is describing a delivery problem may require a model. Investigating that problem might require an agent to look through orders and previous conversations. Issuing a large refund might require a person’s approval.
Those are four different kinds of work. Giving a model control over all of them introduces decisions that the application could already make.
Chloe lets you choose the appropriate level of autonomy for each step:
- work.step() — I know what to do. Run ordinary code.
- work.model() — I know what to ask. Get a structured answer from a model.
- work.agent() — I know what I want. Give an agent a goal, tools, and limits.
- work.ask() — A person needs to decide. Pause the job and wait for an answer.
The principle is the least autonomy that does the job.
Your loops, conditions, and business rules stay in TypeScript. You decide when to call a model, what information it receives, and what happens with its answer.
This doesn’t make the model’s judgment deterministic. It makes the boundary around that judgment explicit.
An agent is a file you can read
In Chloe, you define an agent in TypeScript: its model, instructions, memory folder, tools, jobs, and channels.
You can read that definition and understand what you configured it to do. Another agent is another file, with its own responsibilities.
Jobs can run on a schedule, start from a conversation, or be triggered manually. An agent might send a morning briefing, answer questions about your email, or investigate something using a limited set of tools.
For an autonomous step, you can specify the tools available to it, limit how many steps it takes, and set a spending budget. When a decision needs human approval, the job can wait — even across a runtime restart.
See what happened and what it cost
Writing the workflow is only part of the problem. Once it runs, I want to understand its behavior.
Chloe records the steps, model calls, and tool calls in a run, along with timing and model costs. You can see where ordinary code ended and model judgment began.
That distinction also helps with testing. Application logic gets ordinary tests; prompts get evals. A correct workflow and a good model decision are related, but they need different checks.
The dashboard gives you a place to inspect runs, chat with agents, and browse their memory.
Run locally, access remotely
Chloe runs on your own machine or server as a single Node.js process. Runtime state lives in SQLite, and agent memory lives in files you own.
The dashboard can run alongside it. An optional cloud connection lets you access that dashboard remotely while the runtime continues executing on your machine.
Local execution doesn’t mean every model runs locally: model calls go to the provider you configure.
I’ve kept the architecture small deliberately. Chloe currently targets workloads that fit in one process on one machine. It isn’t a distributed workflow engine, and an interrupted active run is marked failed rather than automatically continuing from any arbitrary point.
Why I’m building it
I’m interested in AI systems that fit into real processes: systems with clear inputs, useful outputs, visible costs, and decisions that someone can review.
Chloe is my way of exploring that through a runtime where the workflow is readable TypeScript and autonomy is a choice you make at each step.
There’s room for ordinary code, model judgment, autonomous investigation, and human decisions in the same application. I want the developer to be able to say clearly which one belongs where.
Chloe is open source under the MIT license. If this approach matches something you’re building, you can explore the examples (https://chloejs.org/examples), read the documentation, or look through the source on GitHub (https://github.com/carlosmartinezt/chloejs).
I’d be interested to hear where you would draw the boundary between a workflow you write and one an agent figures out.
Top comments (0)