Most agent tools give you two choices when the output is wrong. Start over, or edit the prompt and hope. We wanted something in between for Bees, the open source desktop app we build that runs a team of AI agents on your own computer.
Every goal gets a reviewer
A goal in Bees runs through three stages: Work, Review, Done. A work agent does the job. Then a fresh reviewer agent checks the result. It gets the goal, the deliverables and the evidence Bees recorded, but not the worker's private reasoning, so it judges the result on its own.
If the result falls short, the reviewer sends the same item back to Work with specific feedback, and the next attempt sees both the old output and the notes. Retries are capped. A goal that keeps failing stops under Needs your attention instead of looping all night.
You can also put yourself in that loop. A goal or a stage can ask for your review, and then you get Approve or Reject. Reject needs a reason.
The reason has a scope
This is the part I think matters most. On a one-off goal, your reason goes back to the agent for a revision and nothing else changes.
On a scheduled job, say a weekday support briefing, Bees asks where the reason should apply:
- This run only: the agent revises this run, and future runs behave the same as before.
- Future schedule runs: the reason is also added to a playbook for that job.
That playbook belongs to a specialist. The first time an agent handles a named schedule, Bees creates a specialist for that pair. It inherits the base agent and only adds the job-specific notes. So a lesson like this one:
Do not count auto-replies as customer conversations.
sticks to the support briefing. It doesn't leak into everything else that agent does.
What it doesn't do
- One-off goals never change future behavior.
- The reviewer can't edit the worker.
- Nothing changes because of a score.
- No rejection changes the base agent for the whole organization.
Saying no to a tool action is kept apart from rejecting the work, too. If you deny an agent sending an email, that's a permission call for this moment. It doesn't turn into a rule.
You can see it and undo it
Under Process Runs → Schedules, each specialist shows its playbook, a revision number and the history of where each change came from: your feedback, a manual edit, an undo or a reset. You can edit it, undo the last change, or reset it to empty and go back to the base agent.
We picked this over automatic learning for a simple reason. "Bad result" makes a bad rule, and a rule you can't see is worse. A specific reason written by a person, scoped to one job, is something you can check later.
Bees is free and open source (MIT or Apache 2.0). Version 0.2.0 runs on Mac, Windows and Linux, with your own Codex or Claude Code subscription, a hosted model, or a local model.
- Download: https://bees.bot/download/
- Code: https://github.com/Bees-bot/bees-desktop
- How the review loop works: https://bees.bot/help/loop-engineering
- Specialists and playbooks: https://bees.bot/help/self-improving-agents
Top comments (0)