DEV Community

Cover image for Your Pentest Report Says "SQL Injection on /api/search." It Won't Say Which Line.
Amartya Jha
Amartya Jha

Posted on Originally published at codeant.ai

Your Pentest Report Says "SQL Injection on /api/search." It Won't Say Which Line.

TL;DR: External penetration testing still matters, but the classic once-a-year model has six structural gaps that hurt teams shipping weekly. Here is what external testing actually covers, where it breaks down, and how a code-aware approach closes the gap between a perimeter finding and the pull request that fixes it.


Your security team just handed you a 73-page external penetration testing report with 50+ findings. "SQL injection on /api/search." But which query? Which file? What's the fix for your ORM? The report doesn't say.

By the time you ship a patch and wait six months for a re-test, your team has merged 200 more pull requests. Any one of them could have introduced the next vulnerability, and nobody is looking.

External penetration testing is still essential. It validates your perimeter, it satisfies auditors, and it tells you what an internet-based attacker sees before they attack. But six structural gaps stop the traditional model from keeping pace with modern release velocity. And the pressure is not theoretical: exploited vulnerabilities jumped to 20% of breaches as an initial access vector, up 34% year over year, in the Verizon 2025 Data Breach Investigations Report.

The Three Depths, In Plain Terms

Before the gaps, fix the three testing depths, because everything else turns on them.

Black box is how an attacker sees you from the outside with zero inside knowledge. Public URLs only, enumerate from scratch. If your network is not penetrable, no attacker can see your code. Great for validating the perimeter, useless for tracing authenticated flows or privilege boundaries.

White box is the opposite. A tester reads your entire codebase, every API call, every database call, and builds a threat model from your logic, dependencies, secrets, and infrastructure. Is there an unauthenticated API call hitting the database directly with no rate limiting? From outside you would never know. From the code you see it on the first read.

Gray box clubs the two together. You know from the code that a public API quietly triggers a downstream internal API that makes a database call. So you step outside the network, act like an adversary, and use that inside knowledge to reach the thing that was never meant to be reachable.

That last one is the interesting bit, and we go deeper on all three in Black Box vs White Box vs Gray Box Penetration Testing.

The Five-Phase Attack Lifecycle

A rigorous external engagement runs through five phases.

  1. Reconnaissance. Certificate Transparency log mining, DNS enumeration, cloud asset discovery, tech fingerprinting. This surfaces the forgotten admin.staging.yourapp.com and the S3 bucket nobody remembers.
  2. Service discovery. nmap, httpx fingerprinting, API discovery through JavaScript bundle analysis, auth boundary probing.
  3. Manual validation. Where professional testing separates from scanning. Custom payloads, chaining low-severity findings into critical ones, business logic a scanner cannot understand. Chaining alone can eat seven to eight hours, and then a human researcher re-validates every finding and pushes it further.
  4. Exploitation. Prove it with a working PoC:
# Authentication bypass PoC
curl -X POST https://api.yourapp.com/admin/users \
  -H "X-Forwarded-For: 127.0.0.1" \
  -d '{"role":"admin","email":"attacker@evil.com"}'
Enter fullscreen mode Exit fullscreen mode
  1. Evidence collection. Audit-grade output: curl PoCs, CVSS scores, control mapping, remediation guidance.

The finding classes that recur are the usual suspects: misconfigured cloud storage, missing auth on API endpoints, OWASP Top 10 issues like SQL injection, and IDOR/BOLA patterns. In every case the report tells you the endpoint, not the line. That gap is the whole point of this post.

The Six Gaps Traditional External Testing Leaves Open

These are not six unrelated problems. They are symptoms of one thing: external testing operates outside your codebase, while your vulnerabilities live inside it.

Gap 1: No code-level context. Testers can confirm /api/users returns user data. Without code access they cannot trace whether user input reaches a dangerous sink or whether an authorization check is bypassed internally.

Gap 2: Point-in-time testing. You pass in January. Over the next 51 weeks you ship 26 releases with new endpoints and auth flows that were never tested. Your clean January report still satisfies the auditor in December, on a fundamentally different attack surface.

Gap 3: The remediation disconnect. "Auth bypass on /api/users. CVSS 8.1. Implement proper authorization checks." Which middleware? Which controller? What data flow? The tester proved the exploit from outside and never saw your code, so engineering guesses, and the guessing loop runs 16 to 20 weeks.

Gap 4: Missing attack chains. Three findings, none critical: GraphQL introspection (Low), BOLA (Medium), weak passwords (Medium). Three months later an attacker chains all three into a full breach. Scanners match signatures one at a time. A two-week engagement does not have time to explore every chain.

Gap 5: Compliance evidence vs real security. An annual report cannot answer "how do you validate security between tests?" or "prove this fix never regressed?" Frameworks increasingly want a continuous evidence trail, not a yearly snapshot.

Gap 6: External-only misses insider knowledge. Real attackers scrape GitHub for leaked creds, read JS bundles for internal endpoints, and operate with partial inside knowledge. A pure black-box test never has that context.

The common thread: a perimeter finding is only half the story. The other half is the line of code that caused it.

Code-Aware Offensive Testing, The Short Version

Code-aware testing operates on both sides of the perimeter at once. The same engine that reviews your pull requests, understands your auth middleware, and traces your data flows also runs reconnaissance against your external surface. When it finds an auth bypass on /api/users, it already knows which middleware handles that route, because it has been reviewing that code for months.

The practical difference shows up in the finding itself. Instead of a CVSS number and a shrug, you get:

finding_id: AUTH-001
file: src/controllers/UserController.java:47
issue: "Missing resource ownership check"
source: req.params.userId
sink: db.query() without validation
remediation: |
  if (req.user.id !== parseInt(userId)) {
    return res.status(403).json({ error: "Forbidden" });
  }
Enter fullscreen mode Exit fullscreen mode

File, line, source, sink, fix. That is what turns a report into a merged pull request.

From Finding to PR Fix

The workflow that actually closes findings is not complicated, it just has to be code-precise.

Triage on exploitability, not raw CVSS. A working PoC makes it P0. Pre-auth jumps the queue. PHI, PII, or payment data in the blast radius means fix now.

Patch at the layer that survives a re-test. Enforce the invariant at the query, not in a decorator:

# Survives re-test: ownership enforced at the query level
def get_document(doc_id, current_user):
    return db.query(Document).filter(
        Document.id == doc_id,
        Document.owner_id == current_user.id  # explicit ownership
    ).first_or_404()
Enter fullscreen mode Exit fullscreen mode

Parameterize input rather than escaping it:

# Fails re-test: escaping is fragile
query = f"SELECT * FROM users WHERE email = '{sanitize(email)}'"

# Survives re-test: parameterization enforced by the driver
db.execute("SELECT * FROM users WHERE email = ?", [email])
Enter fullscreen mode Exit fullscreen mode

Lock the fix with a negative test, then re-scan:

def test_document_access_requires_ownership():
    user_a = create_user()
    user_b = create_user()
    doc = create_document(owner=user_a)
    response = client.get(f"/api/documents/{doc.id}", auth=user_b)
    assert response.status_code == 403  # negative test
Enter fullscreen mode Exit fullscreen mode

When Each Approach Fits

Not every team needs the same thing.

  • Traditional firms fit stable surfaces, quarterly-or-less releases, and bespoke red-team work. Trade-off: 2 to 4 week turnaround, separate re-test fees, findings as reports not tickets.
  • Continuous monitoring fits SaaS that needs 24/7 surface detection. Trade-off: flags potential issues but lacks exploitation depth and code-level remediation.
  • Code-aware offensive testing fits weekly deploys, API-driven architectures, and continuous-evidence compliance. Findings map to code, chains get built, re-scans are fast.

If you ship weekly, handle regulated data, and cannot easily translate reports into developer tickets, the code-aware model is probably the right fit.

The Takeaway

External penetration testing earns its place for compliance and perimeter validation. But a traditional external test alone cannot keep pace with weekly releases. The fix is not to replace it, it is to give it code memory, so a perimeter finding arrives with the line that caused it and the diff that closes it.

If you want to see the gap on your own app, the fastest way is to point a code-aware scan at your ten most sensitive internet-facing endpoints and compare what comes back to what your last report said.

This post was originally published on the CodeAnt AI blog, where the full version includes the complete six-gap breakdown, the compliance evidence workflow, and an FAQ. If you found this useful, a follow or a comment on what your team runs today is always appreciated.

Top comments (0)