Everyone asks whether AI agents can find security bugs. Fewer people ask the question a team lead asks five minutes later: what does it cost each time we run it?
I'm an intern at RakFort in Dublin, and for the last few weeks my job has been testing the cost tracking in secfoo, an open-source (MIT) CLI that gives AI coding agents a fixed security brief and a fixed report format. This post is what I found, including the parts that are not finished.
The problem with "just ask the agent"
You can point Claude Code, Cursor or Codex at a repo and ask for a security review. You will get something useful. You will not get:
- the same structure twice, so you cannot compare runs
- a record of what the run cost
- a way to stop a CI job that is burning money
secfoo handles the first with what it calls skills (structured briefs such as sast, threat-modeling and secret-scanning). The second and third are what I tested.
The setup
I used a deliberately small target: a 30-line Flask app with one SQL injection, where get_user() builds its query by joining strings. Small on purpose, so I could check the report by hand.
pip install "secfoo[api]"
export OPENAI_API_KEY=...
secfoo run --skill sast --agent api --target . --project-name demo --app-id ""
The api agent calls a model API directly, so you do not need an agent CLI installed. The last two flags skip the interactive prompts, which matters in scripts.
Here is the real output:
$ secfoo run --skill sast --agent api --target . --project-name demo --app-id ""
success — SAST — Static Code Analysis
Assessment results
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━┳━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┓
┃ Skill ┃ Status ┃ Duration ┃ Findings (C/H/M/L) ┃ Cost ┃ Run ID ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━╇━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┩
│ SAST — Static Code Analysis │ success │ 11.2s │ 0/1/0/0 │ <$0.01 │ 888f222e-7228-4b2b-b306-78db7a17d499 │
└─────────────────────────────┴─────────┴──────────┴────────────────────┴────────┴──────────────────────────────────────┘
AI spend for this run: <$0.01
Run secfoo serve to view full reports in your browser.
What one run cost
The scan finished in 11.2 seconds and reported one finding: the SQL injection in get_user(), which is the right answer for this app.
It used 2,183 input tokens and 981 output tokens, 3,164 in total. secfoo printed the spend for the run as under one cent. It shows <$0.01 rather than rounding down to zero, which I liked: a zero would suggest the run was free.
In my earlier testing, a repeat scan of an unchanged repo used exactly the same input tokens. That is expected today: nothing is cached between runs yet, so the second scan costs the same as the first.
Same kind of target, different agent
Earlier I ran a similar small app through Codex instead of the api agent:
secfoo run --skill sast --agent codex --target . --project-name demo --app-id ""
It finished in 27.5 seconds and also found the injection. Two things were different:
-
Cost shows as
-, not$0.00. Codex runs on a subscription login, so there is no per-run price to record. A dash is the right answer. A zero would be a lie. -
Tokens were much higher. Codex's own log reported 15,201 tokens for the run, roughly five times the
apiagent, because an agent CLI explores the repo with tools instead of receiving one prepared prompt.
That second point is the most useful thing I learned. The agent you pick can change the token count by a multiple, on a small target, for the same finding.
Seeing spend across runs
Per-run numbers are nice. The summary is what a team would actually use:
secfoo cost
secfoo cost --by skill --since 2026-09-01
secfoo cost --by agent --project demo
You can group by agent, skill or project. The same numbers appear in the local dashboard (secfoo serve), and in my tests the dashboard total matched the CLI total for the same runs.
$ secfoo cost --by agent --project demo
AI spend by agent
┏━━━━━━━┳━━━━━━┳━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┳━━━━━━━━┓
┃ Agent ┃ Runs ┃ Input tokens ┃ Output tokens ┃ Cost ┃
┡━━━━━━━╇━━━━━━╇━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━╇━━━━━━━━┩
│ api │ 2 │ 2,183 │ 981 │ <$0.01 │
├───────┼──────┼──────────────┼───────────────┼────────┤
│ Total │ 2 │ 2,183 │ 981 │ <$0.01 │
└───────┴──────┴──────────────┴───────────────┴────────┘
1 run(s) have no cost recorded -- their agent doesn't report it (cursor, agy), or they ran before cost tracking existed.
That table shows two runs, not one. My first attempt failed because I pasted my API key with a character missing. secfoo still counted it as a run and told me that one run had no cost recorded. I did not plan that, but it is a fair picture of how it behaves when something goes wrong.
Making cost a CI gate
This is the part I think matters most for real pipelines:
secfoo run --skill sast --agent api --target . \
--fail-on high --max-cost 2.00 --json > secfoo-result.json
The exit code tells CI what happened: 0 means the run succeeded and no gate tripped, 1 means a run failed or timed out, and 2 means the run succeeded but a finding or the cost crossed your threshold. A security scan that can fail a build for being too expensive is not something I had seen before.
What is not there yet
I was testing, so I was looking for gaps. The honest list:
- Cost coverage depends on the agent. API-key runs give you tokens and cost. Subscription-based agent CLIs often give you neither, and you see a dash.
- Codex tokens are not recorded yet. Codex prints its total, but secfoo does not pick it up today. There is an open item for it.
- Failed runs are counted as runs. They show up in the totals with no cost recorded, as my mistyped key showed.
- Severity can differ between the CLI and the dashboard. In my run the CLI table counted the finding as High, and the dashboard labelled the same finding Critical. It is a known issue.
-
--sincecompares against UTC. If you run scans late at night in another timezone, the cut-off may not be where you expect. - No savings on repeat scans yet. Reusing results for unchanged code is planned, not shipped.
What I would tell a team
- Start with the
apiagent on one repo, so you get real token and cost numbers from day one. - Run the same repo through the agent your developers already use, and compare the reports side by side.
- Put
--max-costin CI before you put the scan on every pull request.
Try it
pip install secfoo
secfoo run --skill sast --agent claude
secfoo serve
The repo is here: https://github.com/secfoo-com/secfoo. If it is useful to you, a star helps a small open-source project a lot, and issues telling us where the output is wrong help even more.
We are also part of the team organising a secfoo hackathon in Dublin in mid-November. If you would like to take part, or to help as a moderator, leave a comment and I will get back to you.
Top comments (2)
The 5x token jump between the direct API call and the Codex CLI on a 30-line file is the real takeaway here. Tool-using agent loops explore aggressively, so on a monorepo that multiplier easily blows past 20x as the agent traverses unrelated directories. Setting --max-cost in CI works as a decent circuit breaker, but until scans are scoped to changed git diffs, running full agent passes per PR gets expensive fast.
Thanks Reid, that matches what I saw. I have only measured small targets so far, so I can't put a number on a monorepo yet, but the direction makes sense: the agent CLI explores with tools, while the api agent gets one prepared prompt.
Two things help today: --exclude keeps unrelated paths out of the scan and --max-cost is the circuit breaker you describe.
Scanning only what changed since the last run is planned but not shipped. I agree that is what makes per-PR scans affordable. A larger repo is the next thing I want to measure.