DEV Community

Prema Jyothi
Prema Jyothi

Posted on

I Planted 3 Sneaky Security Flaws in a Web App. Could AI Catch Them?

Every developer on my team has been talking about using AI for code reviews.

As an intern at RakFort in Dublin, I've been spending time exploring and testing security tooling. That got me wondering:

Can an AI security reviewer actually understand why code is dangerous, or does it just look for suspicious keywords?

So I decided to test it myself.

I built a small Python application and deliberately planted three security vulnerabilities inside it.

One hardcoded credential.
One SQL injection.
One command injection.

I already knew where all three were.

The real question was:

Could the AI find them too?

That led me to SecFoo, an open-source MIT-licensed CLI tool designed to run structured security reviews using AI coding agents.

Instead of giving it a clean project, I created my own little security challenge.


The Experiment

The goal was simple:

Give an AI security reviewer a deliberately vulnerable application and see whether it can identify the problems, explain the risks, and provide useful remediation guidance.

I created a small Python web application and intentionally introduced three vulnerable patterns.

I wasn't trying to build a production application. The purpose was to create a controlled experiment where I already knew the expected findings.

Trap #1 — Hardcoded Credential

First, I placed a fake credential directly in the source code:

AWS_SECRET_KEY = "FAKE_SECRET_FOR_TESTING"
Enter fullscreen mode Exit fullscreen mode

In a real application, credentials should not be committed directly into source code.

For this experiment, I used a fake value intentionally so that no real secret was exposed.

The question was whether the security review would recognize that the credential was embedded in application code and explain why that was risky.


Trap #2 — SQL Injection

Next, I created a database query using user-controlled input:

query = f"SELECT * FROM users WHERE id = '{user_id}'"
Enter fullscreen mode Exit fullscreen mode

The problem here isn't simply that the code contains an SQL query.

The important part is how the input reaches that query.

If user_id comes from an untrusted user, constructing the query this way can allow the input to alter the SQL statement.

A useful security review should therefore understand the relationship between the input and the database operation rather than simply searching for the word SELECT.


Trap #3 — Command Injection

Finally, I introduced a command execution pattern:

os.system(f"ping -c 1 {host}")
Enter fullscreen mode Exit fullscreen mode

Again, the interesting part is the data flow.

If host can be controlled by an untrusted user, passing it directly into a system command can create a command-injection risk.

This gave me three different types of problems to test:

  • Sensitive information exposed in source code
  • Untrusted input reaching a database query
  • Untrusted input reaching a system command

Running SecFoo

Once the intentionally vulnerable application was ready, I ran SecFoo's SAST skill against the project.

I configured the API key through an environment variable and then ran:

secfoo run --skill sast --agent api --target . --project-name vulnerability-challenge --app-id ""
Enter fullscreen mode Exit fullscreen mode

After the review completed, I opened the SecFoo dashboard:

secfoo serve
Enter fullscreen mode Exit fullscreen mode

The dashboard allowed me to inspect the findings from the security review.

So... Did It Catch Them?

This was the part I was actually interested in.

I wasn't looking for a fancy report.

I already knew the vulnerabilities existed.

I wanted to know whether the AI-assisted review could identify them and provide enough context to understand why they were vulnerabilities.

Results

Vulnerability Detected? What I was looking for
Hardcoded credential ✅ Recognition that sensitive credentials should not be embedded in source code
SQL injection ✅ Understanding that user-controlled input reaches the SQL query
Command injection ✅ Understanding that user-controlled input reaches command execution

In my test, all three intentionally planted vulnerabilities were identified.

That was more interesting to me than simply getting three green checkmarks.

The useful part was the context around the findings: the review could point toward the risky code and explain the security concern rather than simply saying that a particular function or keyword was "bad."

What I Found Interesting

One thing this experiment reminded me of is that security problems aren't always about individual lines of code.

Consider this:

query = f"SELECT * FROM users WHERE id = '{user_id}'"
Enter fullscreen mode Exit fullscreen mode

A pattern-based check can flag this code as suspicious.

But the more interesting security question is:

Where did user_id come from?

If it came directly from an HTTP request, the risk is very different from a value that was safely generated internally.

The same idea applies to command execution.

The security issue isn't simply:

os.system(...)
Enter fullscreen mode Exit fullscreen mode

It's the combination of:

untrusted input → application logic → dangerous operation

That's where AI-assisted analysis becomes interesting.


Here are the findings from the run:

Hardcoded credential finding

SQL injection and Command injection finding

Is This Replacing Traditional SAST?

I don't think so.

Traditional SAST tools are still extremely useful. They are fast, predictable, and excellent at detecting many known vulnerability patterns.

But some security findings require understanding how different pieces of an application interact.

That's where I think AI-assisted security reviews have an interesting role.

Instead of asking only:

"Does this line match a known vulnerability pattern?"

we can also ask:

"What is this code doing, where does this data come from, and what could happen if that data is malicious?"

That's a different way of looking at the problem.

And importantly, it doesn't mean an AI reviewer will always be correct.


What I Learned

1. Context matters

The interesting part wasn't simply detecting suspicious code.

The review needed to understand how input was being used and why that usage created a security risk.

2. Controlled experiments are useful

I already knew exactly where the vulnerabilities were.

That made it much easier to evaluate the result.

Instead of saying:

"The tool found some security issues."

I could ask a much more specific question:

"I planted three known vulnerabilities. How many did it find?"

That's a much better way to test a security tool.

3. AI security reviews still need human validation

Finding a vulnerability is only the beginning.

A real security review still needs a developer or security engineer to verify the finding, understand the application context, and decide whether the suggested remediation is appropriate.

AI can help with the investigation, but I wouldn't treat its output as an automatic security approval.


What I Want to Test Next

This experiment was intentionally simple.

The vulnerabilities were placed in obvious locations so that I could measure the result.

But real applications aren't always like that.

What happens when:

  • the input flows through several functions?
  • the vulnerable code is spread across multiple files?
  • the dangerous behavior is hidden behind an abstraction?
  • the vulnerability depends on application architecture?
  • two individually harmless components become dangerous when combined?

That's the experiment I'd like to try next.

Three deliberately vulnerable patterns are one thing.

Finding a vulnerability that requires understanding an entire application is a much harder challenge.


Wanna Try It?

If you're interested in experimenting with AI-assisted security reviews, you can check out SecFoo and run your own tests against a project.

https://github.com/secfoo-com/secfoo

You can start with:

pip install secfoo
Enter fullscreen mode Exit fullscreen mode

Then try creating your own small security challenge.

And here's my question for developers:

If you had to plant one vulnerability that an AI security reviewer would struggle to find, what would you choose?

Drop your idea in the comments.

I'd genuinely like to see what other developers come up with.

Top comments (1)

Collapse
 
indiainfranotes profile image
IndiaInfraNotes •

Planted flaws are a fair test, but the reviewer that only flags a single pattern will miss the one that needs the whole app. A detector that says "probably AI" is not the same as a signed record of who changed what. If the signature is missing, would you treat the build as untrusted or just as incomplete?
iin1006h10