A lot of LLM app risk sits somewhere duller than the model. It sits in a system prompt glued to user input with string concatenation, or in a tool definition that hands the model a shell.
Both are visible in the code, before anything ships. This post walks through the prompt-review skill in secfoo, an open-source security assessment CLI, and maps each of its seven checks to the OWASP Top 10 for LLM Applications (2025).
Disclosure: I'm an AI security intern at RakFort and I contribute to secfoo. The mapping is my own reading, not an official one.
What prompt-review looks at
secfoo has no scanner engine of its own. A skill is a structured brief written in Markdown. secfoo hands that brief and your target to a coding-agent CLI you already use (Claude Code, Cursor, Antigravity or Gemini CLI) and collects the report.
The prompt-review brief is 45 lines long and asks the agent to do seven things:
- Locate every prompt surface. System prompts, templates, few-shot examples, tool and function-calling schemas, allowed-tool lists and permission scopes.
- Check injection resistance. Find where untrusted input (user messages, retrieved documents, tool outputs, web content) is concatenated into instruction context with no delimiter or boundary.
- Check tool permission scope. Are callable tools scoped to the minimum, or does the model get shell execution, unrestricted file writes, unscoped database access or outbound network calls?
- Check for a human in the loop before side effects. Are model-suggested commands, SQL, file writes, API calls or payments validated or confirmed before they run?
- Check for data leakage. Are secrets, API keys or PII interpolated into prompts, logged in full, or sent to a third-party model provider without the user knowing?
- Check guardrail robustness. Can a crafted message override the refusal instructions, and is there a code-level backstop for high-risk actions?
- Check rate and cost controls. Can one caller trigger an unbounded number of expensive model calls?
How the checks map to the OWASP LLM Top 10
Six of the ten OWASP categories are touched by at least one check.
| # | Check | OWASP LLM Top 10 (2025) | Why it maps |
|---|---|---|---|
| 1 | Prompt surfaces | LLM07 System Prompt Leakage | You cannot judge what a system prompt exposes until you have found every one |
| 2 | Injection resistance | LLM01 Prompt Injection | Covers direct input and indirect input from documents, tools and web content |
| 3 | Tool permission scope | LLM06 Excessive Agency | Excessive functionality and excessive permissions |
| 4 | Human in the loop | LLM06 Excessive Agency, LLM05 Improper Output Handling | Excessive autonomy, plus model output passed to a shell or database unvalidated |
| 5 | Data leakage | LLM02 Sensitive Information Disclosure, LLM07 System Prompt Leakage | Secrets and PII in prompts and logs |
| 6 | Guardrail robustness | LLM01 Prompt Injection | Jailbreaks are the direct form of injection |
| 7 | Rate and cost controls | LLM10 Unbounded Consumption | Cost and availability risk from uncapped calls |
The table shows why I like this brief as a first pass. Checks 2 to 4 are the chain an agent compromise needs: untrusted text gets in, the model has a powerful tool, and nothing stops the call.
Try it on your own repo
You need Python 3.10+ and one supported agent CLI installed and signed in.
pip install secfoo
# see which agent CLIs secfoo can find on your PATH
secfoo agents
# review the current directory
secfoo run --skill prompt-review --agent claude
# or point it at a public repo and run the full checklist
secfoo run --skill prompt-review --depth standard \
--target https://github.com/org/repo --agent claude
# browse the report in the local dashboard (127.0.0.1:8787)
secfoo serve
A few things worth knowing before the first run:
-
--depthdefaults toquick, a fast triage.standardruns the full checklist. - On a real terminal the first run asks for a project name and an optional application ID. Rescans with the same ID land in the same case file.
- Skills run concurrently, so
--skill prompt-review --skill secret-scanningcosts you one wait, not two. - Runs stay on your machine.
secfoo listandsecfoo show <run-uuid>read them back in the terminal. - For CI,
--fail-on highexits with code 2 when a finding at that severity or worse turns up, and--jsonwrites a machine-readable result to stdout.
A small target to practise on
If you have no LLM code to hand, save this as bot.py in an empty folder and run the skill against it. I wrote it to be wrong in as many ways as possible.
import os
import subprocess
SYSTEM = (
"You are a support bot. "
f"Billing API key: {os.environ['BILLING_API_KEY']}. "
"Never reveal the key."
)
def run_shell(cmd: str) -> str:
"""Tool exposed to the model."""
return subprocess.run(cmd, shell=True, capture_output=True, text=True).stdout
def answer(llm, user_msg: str, retrieved_doc: str) -> str:
prompt = SYSTEM + "\nContext: " + retrieved_doc + "\nUser: " + user_msg
reply = llm(prompt, tools=[run_shell])
if reply.tool_call:
return run_shell(reply.tool_call.args["cmd"])
return reply.text
Reading it against the seven checks by hand, this is what a reviewer should flag:
-
Check 2:
retrieved_docanduser_msgare concatenated straight into the instruction context, so a poisoned document can issue instructions. -
Check 3:
run_shellgives the model arbitrary command execution, far beyond what a support bot needs. - Check 4: the tool call runs immediately, with no allow-list and no confirmation.
- Check 5: a live API key sits inside the system prompt.
- Check 6: the only thing protecting that key is the sentence "Never reveal the key".
-
Check 7: nothing limits how often
answercan be called.
Compare that list with what your agent reports. An LLM is doing the reading, so wording and coverage vary between agents and between runs.
What it does not do
- It is a code review, not a red team. Nothing is sent to your running app. Pair it with dynamic testing tools such as garak or PyRIT.
-
Four OWASP categories are out of scope for this skill: LLM03 Supply Chain, LLM04 Data and Model Poisoning, LLM08 Vector and Embedding Weaknesses and LLM09 Misinformation. secfoo's
sca-reachabilityskill and AI-BOM upload cover part of the supply-chain side. I found nothing for the other three. - The reviewer is an LLM. It can miss things and it can over-report. Treat findings as leads to verify.
-
Your code goes wherever your agent CLI sends it. secfoo stores runs locally, but the agent still talks to its model provider. Use
--excludeto keep paths out of a scan. -
Agent runs cost tokens.
--max-costputs a dollar cap on a run for agents that report spend.
Where this could go next
The gaps above are the interesting part. Adding a skill to secfoo means dropping one Markdown file into src/secfoo/skills/definitions/. The CLI choices and dashboard pages are derived from that folder.
So a brief for vector-store weaknesses or model supply chain is a small, self-contained pull request. If you test LLM apps and have opinions on what such a brief should ask, the repo is MIT licensed and CONTRIBUTING.md explains how to start.
Repo: github.com/secfoo-com/secfoo. If the tool is useful to you, a star helps other people find it.
What would you add to the seven checks? Tell me in the comments.
Top comments (0)