AI coding agents are getting very good at taking real actions: merging pull requests, deploying applications, and changing production systems.
But there is a problem I kept running into:
An agent saying "done" is not proof that the intended change actually happened.
So I built TookEffect.
TookEffect independently checks the real external system after an AI agent performs an action, verifies the expected outcome, and keeps evidence of what actually happened.
A simple example
An AI agent says:
"The pull request was merged."
TookEffect doesn't trust that response.
It reads GitHub independently, checks the expected repository, PR, branch and resulting state, and produces a verdict:
-
APPLIED— the expected effect is proven -
NOT_APPLIED— the expected effect did not happen -
AMBIGUOUS— there isn't enough evidence to prove either outcome
The same idea applies to deployments.
An agent says:
"Production was deployed."
TookEffect checks the real deployment platform instead of trusting the agent's own tool response.
What works today
Right now TookEffect supports verified actions across:
- GitHub
- Vercel
- Cloudflare
It can be used through an API or MCP, so the verification layer is independent of the AI model or coding agent you're using.
The principle behind the project is:
AI says done. TookEffect proves it.
I built this because I think AI agents will increasingly be allowed to perform consequential actions on real systems, and we need something outside the agent itself to verify the resulting state.
It's still early, and I'm especially interested in feedback from developers already using AI agents in real repositories or deployment workflows.
Would independent verification like this be useful in your workflow?
Top comments (2)
The best verification tools separate agent confidence from user-visible reality. A log can say the task completed while the product surface still proves otherwise.
Exactly — that gap is precisely what TookEffect is built to solve.
If you're open to it, I'd love for you to try it and tell me where it holds up or falls short in a real agent workflow. I'm looking for feedback from people who already understand this problem.
tookeffect.com/