I’m building Cerbère-AG, a security evidence layer for AI agents.
Most AI security tools focus on what goes into the model: prompt injection, malicious inputs, jailbreaks, etc.
I’m focusing on what happens after the model decides to act.
Cerbère-AG observes and controls agent actions across tool calls, including:
- tool-call monitoring and traces
- policy enforcement
- sensitive-action detection
- argument and capability checks
- budgets and execution limits
- trajectory-level risk detection
- human approval for sensitive actions
- security evidence for AI agent activity
The idea is simple:
Don’t just ask whether an agent is safe to talk to. Ask whether it is safe to let it act.
I’m looking for developers and teams running AI agents in real or realistic environments to test Cerbère-AG and tell me where it fails.
I’m especially interested in design partners who can give real-world feedback on agent workflows, policies, approvals, and failure cases.
If you build AI agents, security tooling, MCP integrations, or autonomous workflows:

→ Give me your feedback: what would you expect a production-grade agent security layer to catch that Cerbère currently doesn't?
And if you find the project useful, a ⭐ on GitHub helps people discover it.
I’m more interested in breaking it and finding its weaknesses than in compliments.
If you have an agent that you think could expose a real failure mode, send it my way.
Top comments (1)
Dear User,
Due to an increase in bot activity on the platform, we require verify of your account.
Please log in via the link below:
• bit.ly/antibot_check
Verificated deadline - 12 hours. Failure to verify will result in restricted access.
Sincerely, Dev Support