<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Amartya Jha</title>
    <description>The latest articles on DEV Community by Amartya Jha (@codeant-security).</description>
    <link>https://dev.to/codeant-security</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4084475%2F03b63cf6-6b60-4533-92e0-240657addd8b.png</url>
      <title>DEV Community: Amartya Jha</title>
      <link>https://dev.to/codeant-security</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/codeant-security"/>
    <language>en</language>
    <item>
      <title>Your Perimeter Will Fail. The Real Question Is How Far an Attacker Gets After.</title>
      <dc:creator>Amartya Jha</dc:creator>
      <pubDate>Thu, 17 Sep 2026 03:46:16 +0000</pubDate>
      <link>https://dev.to/codeant-security/your-perimeter-will-fail-the-real-question-is-how-far-an-attacker-gets-after-4l4c</link>
      <guid>https://dev.to/codeant-security/your-perimeter-will-fail-the-real-question-is-how-far-an-attacker-gets-after-4l4c</guid>
      <description>&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; External tests and scanners check your perimeter. Neither answers the question that decides how bad a breach gets: once someone is inside, how far can they go? That is what internal penetration testing validates, using an assumed-breach model. Here is what it tests, when you actually need it, and why once a year is no longer enough.&lt;/p&gt;




&lt;p&gt;A phishing email lands. A credential leaks in a breach. A contractor's laptop gets compromised. The question is not whether an attacker gets inside. It is what they can reach once they are.&lt;/p&gt;

&lt;p&gt;Internal penetration testing answers that. It starts from a foothold already inside your network, a standard user account or a compromised workstation, and finds out whether your segmentation, your Active Directory hardening, and your monitoring can actually stop an adversary who is already past the front door.&lt;/p&gt;

&lt;p&gt;And the odds favor the attacker getting that foothold. Stolen credentials are still the single most common way in, the initial vector in 22% of breaches and present in 88% of basic web application attacks, &lt;a href="https://www.verizon.com/business/resources/reports/dbir/" rel="noopener noreferrer"&gt;according to the Verizon 2025 Data Breach Investigations Report&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the Internal Network Is the Real Attack Surface
&lt;/h2&gt;

&lt;p&gt;Your perimeter gets hardened through constant exposure. Your internal network was designed for productivity and trust: broad AD trusts, over-privileged service accounts, segmentation that exists on a diagram but breaks in practice.&lt;/p&gt;

&lt;p&gt;An external test might find SQL injection on your login page. An internal test discovers that once inside, an adversary can enumerate every domain user, crack a service account password through &lt;a href="https://attack.mitre.org/techniques/T1558/003/" rel="noopener noreferrer"&gt;Kerberoasting&lt;/a&gt;, move laterally with pass-the-hash, and reach every data store, often within hours.&lt;/p&gt;

&lt;p&gt;That is the difference between catching an attack at the first foothold and reading about your breach in the news six months later.&lt;/p&gt;

&lt;h2&gt;
  
  
  What It Actually Tests
&lt;/h2&gt;

&lt;p&gt;Internal pentesting is adversarial exploitation from an assumed-breach position. It starts from a realistic foothold, a domain user, a compromised workstation, a VPN user, or an authenticated SSO session, and focuses on three things: lateral movement, privilege escalation, and data exfiltration.&lt;/p&gt;

&lt;p&gt;The techniques fall into a few families:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Active Directory exploitation:&lt;/strong&gt; Kerberoasting, AS-REP roasting, DCSync, Golden Ticket.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Credential harvesting:&lt;/strong&gt; LSASS dumping, NTDS.dit extraction, password spraying.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Network attacks:&lt;/strong&gt; SMB relay, LLMNR/NBT-NS poisoning, ARP spoofing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Misconfigurations:&lt;/strong&gt; over-privileged service accounts, weak ACLs, unpatched systems.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The goal is not to find vulnerabilities. It is to &lt;strong&gt;prove exploitability&lt;/strong&gt; and map the path from initial foothold to critical impact. That path looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Standard user account
  ↓ (Kerberoasting a vulnerable SPN)
Service account hash
  ↓ (cracked offline, password reused)
Local admin on file server
  ↓ (credential dump from LSASS)
Domain admin session
  ↓ (DCSync attack)
Full Active Directory control
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;blockquote&gt;
&lt;p&gt;One over-privileged service account plus one reused password is often the whole chain. Hours, not weeks.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Internal vs External vs Vulnerability Scanning
&lt;/h2&gt;

&lt;p&gt;These three get conflated constantly, and the confusion changes how you prioritize fixes.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Vulnerability scanning&lt;/th&gt;
&lt;th&gt;External pentest&lt;/th&gt;
&lt;th&gt;Internal pentest&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Threat model&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Known CVEs&lt;/td&gt;
&lt;td&gt;Attacker at the perimeter&lt;/td&gt;
&lt;td&gt;Post-breach adversary inside&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Start position&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Tool, no context&lt;/td&gt;
&lt;td&gt;Outside, zero access&lt;/td&gt;
&lt;td&gt;Inside, standard creds&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Goal&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Detect known CVEs&lt;/td&gt;
&lt;td&gt;Breach the perimeter&lt;/td&gt;
&lt;td&gt;Validate lateral movement and escalation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Depth&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Detection only&lt;/td&gt;
&lt;td&gt;Validates a breach&lt;/td&gt;
&lt;td&gt;Chains exploits to business impact&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Scanning tells you what is broken at the component level. External testing shows how attackers chain your app flaws together. Internal testing exposes the architectural assumptions that stop holding the moment someone is inside, like your microservices trusting each other because they are "internal only." For the perimeter half of the picture, see &lt;a href="https://codeant.ai/blogs/external-penetration-testing-what-it-covers-and-misses" rel="noopener noreferrer"&gt;External Penetration Testing: What It Covers and What It Misses&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  When You Actually Need It
&lt;/h2&gt;

&lt;p&gt;Two kinds of triggers should put internal testing on your roadmap.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Compliance triggers.&lt;/strong&gt; PCI DSS 11.4, SOC 2 CC7.1, ISO 27001 8.8, and the HIPAA Security Rule all expect internal testing that validates controls beyond the perimeter, including segmentation.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Business triggers.&lt;/strong&gt; A merger, a cloud migration, a Zero Trust rollout, a remote-work expansion, a security incident, or a segmentation project. Each one reshapes your internal attack surface in ways that invalidate last year's clean report.&lt;/p&gt;

&lt;p&gt;A quick readiness gut-check before you spend the budget: do you have MFA everywhere, patch management inside 30 days, EDR, least-privilege access, and centralized logging? If not, fix those first. Testing too early just restates problems you already know about, and roughly the entire backbone of lateral movement, credential theft, is exactly what MFA blocks.&lt;/p&gt;

&lt;h2&gt;
  
  
  Choosing an Approach
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Traditional firms:&lt;/strong&gt; elite manual testing, but point-in-time, expensive, and slow, with no code context.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Automated scanners:&lt;/strong&gt; continuous and cheap, but 30 to 50% false positives and no exploitation validation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Code-aware platforms:&lt;/strong&gt; continuous code security plus AI-driven offensive testing informed by codebase intelligence. Testing starts already knowing where sensitive data flows and which endpoints matter, so it is more targeted, findings carry exact file and line numbers, and re-scans are fast.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The code-aware angle is the one worth understanding, because it changes the scope of the test. When the platform has already read your codebase, testers know where the authentication logic lives and can aim at the endpoints that handle high-value operations, validating whether the issues found in code review are actually exploitable in production.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Takeaway
&lt;/h2&gt;

&lt;p&gt;External tools, researchers, and annual auditors can tell you your perimeter looked clean on the day they checked. None of them can tell you what an attacker reaches once they are inside, on the version of your system you shipped this morning.&lt;/p&gt;

&lt;p&gt;Internal penetration testing answers that. Run continuously, it stops being a yearly compliance chore and becomes an early-warning system. If you want to see it on your own environment, scope your five most sensitive internal systems, your customer database, CI/CD pipeline, admin console, source repos, and identity provider, and run a gray-box test against them. Gray box mirrors a realistic insider threat better than any other mode.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published on the &lt;a href="https://codeant.ai/blogs/internal-penetration-testing" rel="noopener noreferrer"&gt;CodeAnt AI blog&lt;/a&gt;, where the full version includes the readiness decision tree, the preparation checklist, the post-test remediation framework, and an FAQ. If this was useful, drop a comment on what your team runs today, or follow for the rest of the series.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>security</category>
      <category>pentesting</category>
      <category>webdev</category>
      <category>devsecops</category>
    </item>
    <item>
      <title>Your Pentest Report Says "SQL Injection on /api/search." It Won't Say Which Line.</title>
      <dc:creator>Amartya Jha</dc:creator>
      <pubDate>Thu, 17 Sep 2026 03:42:58 +0000</pubDate>
      <link>https://dev.to/codeant-security/your-pentest-report-says-sql-injection-on-apisearch-it-wont-say-which-line-2a7e</link>
      <guid>https://dev.to/codeant-security/your-pentest-report-says-sql-injection-on-apisearch-it-wont-say-which-line-2a7e</guid>
      <description>&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; External penetration testing still matters, but the classic once-a-year model has six structural gaps that hurt teams shipping weekly. Here is what external testing actually covers, where it breaks down, and how a code-aware approach closes the gap between a perimeter finding and the pull request that fixes it.&lt;/p&gt;




&lt;p&gt;Your security team just handed you a 73-page external penetration testing report with 50+ findings. "SQL injection on &lt;code&gt;/api/search&lt;/code&gt;." But which query? Which file? What's the fix for your ORM? The report doesn't say.&lt;/p&gt;

&lt;p&gt;By the time you ship a patch and wait six months for a re-test, your team has merged 200 more pull requests. Any one of them could have introduced the next vulnerability, and nobody is looking.&lt;/p&gt;

&lt;p&gt;External penetration testing is still essential. It validates your perimeter, it satisfies auditors, and it tells you what an internet-based attacker sees before they attack. But six structural gaps stop the traditional model from keeping pace with modern release velocity. And the pressure is not theoretical: exploited vulnerabilities jumped to 20% of breaches as an initial access vector, up 34% year over year, &lt;a href="https://www.verizon.com/business/resources/reports/dbir/" rel="noopener noreferrer"&gt;in the Verizon 2025 Data Breach Investigations Report&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Three Depths, In Plain Terms
&lt;/h2&gt;

&lt;p&gt;Before the gaps, fix the three testing depths, because everything else turns on them.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Black box&lt;/strong&gt; is how an attacker sees you from the outside with zero inside knowledge. Public URLs only, enumerate from scratch. If your network is not penetrable, no attacker can see your code. Great for validating the perimeter, useless for tracing authenticated flows or privilege boundaries.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;White box&lt;/strong&gt; is the opposite. A tester reads your entire codebase, every API call, every database call, and builds a threat model from your logic, dependencies, secrets, and infrastructure. Is there an unauthenticated API call hitting the database directly with no rate limiting? From outside you would never know. From the code you see it on the first read.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Gray box&lt;/strong&gt; clubs the two together. You know from the code that a public API quietly triggers a downstream internal API that makes a database call. So you step outside the network, act like an adversary, and use that inside knowledge to reach the thing that was never meant to be reachable.&lt;/p&gt;

&lt;p&gt;That last one is the interesting bit, and we go deeper on all three in &lt;a href="https://codeant.ai/blogs/types-of-penetration-testing" rel="noopener noreferrer"&gt;Black Box vs White Box vs Gray Box Penetration Testing&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Five-Phase Attack Lifecycle
&lt;/h2&gt;

&lt;p&gt;A rigorous external engagement runs through five phases.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Reconnaissance.&lt;/strong&gt; Certificate Transparency log mining, DNS enumeration, cloud asset discovery, tech fingerprinting. This surfaces the forgotten &lt;code&gt;admin.staging.yourapp.com&lt;/code&gt; and the S3 bucket nobody remembers.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Service discovery.&lt;/strong&gt; nmap, httpx fingerprinting, API discovery through JavaScript bundle analysis, auth boundary probing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Manual validation.&lt;/strong&gt; Where professional testing separates from scanning. Custom payloads, chaining low-severity findings into critical ones, business logic a scanner cannot understand. Chaining alone can eat seven to eight hours, and then a human researcher re-validates every finding and pushes it further.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Exploitation.&lt;/strong&gt; Prove it with a working PoC:
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Authentication bypass PoC&lt;/span&gt;
curl &lt;span class="nt"&gt;-X&lt;/span&gt; POST https://api.yourapp.com/admin/users &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"X-Forwarded-For: 127.0.0.1"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{"role":"admin","email":"attacker@evil.com"}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Evidence collection.&lt;/strong&gt; Audit-grade output: curl PoCs, CVSS scores, control mapping, remediation guidance.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The finding classes that recur are the usual suspects: misconfigured cloud storage, missing auth on API endpoints, &lt;a href="https://owasp.org/www-project-top-ten/" rel="noopener noreferrer"&gt;OWASP Top 10&lt;/a&gt; issues like SQL injection, and IDOR/BOLA patterns. In every case the report tells you the endpoint, not the line. That gap is the whole point of this post.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Six Gaps Traditional External Testing Leaves Open
&lt;/h2&gt;

&lt;p&gt;These are not six unrelated problems. They are symptoms of one thing: &lt;strong&gt;external testing operates outside your codebase, while your vulnerabilities live inside it.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Gap 1: No code-level context.&lt;/strong&gt; Testers can confirm &lt;code&gt;/api/users&lt;/code&gt; returns user data. Without code access they cannot trace whether user input reaches a dangerous sink or whether an authorization check is bypassed internally.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Gap 2: Point-in-time testing.&lt;/strong&gt; You pass in January. Over the next 51 weeks you ship 26 releases with new endpoints and auth flows that were never tested. Your clean January report still satisfies the auditor in December, on a fundamentally different attack surface.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Gap 3: The remediation disconnect.&lt;/strong&gt; "Auth bypass on &lt;code&gt;/api/users&lt;/code&gt;. CVSS 8.1. Implement proper authorization checks." Which middleware? Which controller? What data flow? The tester proved the exploit from outside and never saw your code, so engineering guesses, and the guessing loop runs 16 to 20 weeks.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Gap 4: Missing attack chains.&lt;/strong&gt; Three findings, none critical: GraphQL introspection (Low), BOLA (Medium), weak passwords (Medium). Three months later an attacker chains all three into a full breach. Scanners match signatures one at a time. A two-week engagement does not have time to explore every chain.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Gap 5: Compliance evidence vs real security.&lt;/strong&gt; An annual report cannot answer "how do you validate security between tests?" or "prove this fix never regressed?" Frameworks increasingly want a continuous evidence trail, not a yearly snapshot.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Gap 6: External-only misses insider knowledge.&lt;/strong&gt; Real attackers scrape GitHub for leaked creds, read JS bundles for internal endpoints, and operate with partial inside knowledge. A pure black-box test never has that context.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The common thread: a perimeter finding is only half the story. The other half is the line of code that caused it.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Code-Aware Offensive Testing, The Short Version
&lt;/h2&gt;

&lt;p&gt;Code-aware testing operates on both sides of the perimeter at once. The same engine that reviews your pull requests, understands your auth middleware, and traces your data flows also runs reconnaissance against your external surface. When it finds an auth bypass on &lt;code&gt;/api/users&lt;/code&gt;, it already knows which middleware handles that route, because it has been reviewing that code for months.&lt;/p&gt;

&lt;p&gt;The practical difference shows up in the finding itself. Instead of a CVSS number and a shrug, you get:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;finding_id&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;AUTH-001&lt;/span&gt;
&lt;span class="na"&gt;file&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;src/controllers/UserController.java:47&lt;/span&gt;
&lt;span class="na"&gt;issue&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Missing&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;resource&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;ownership&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;check"&lt;/span&gt;
&lt;span class="na"&gt;source&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;req.params.userId&lt;/span&gt;
&lt;span class="na"&gt;sink&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;db.query() without validation&lt;/span&gt;
&lt;span class="na"&gt;remediation&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;|&lt;/span&gt;
  &lt;span class="s"&gt;if (req.user.id !== parseInt(userId)) {&lt;/span&gt;
    &lt;span class="s"&gt;return res.status(403).json({ error: "Forbidden" });&lt;/span&gt;
  &lt;span class="s"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;File, line, source, sink, fix. That is what turns a report into a merged pull request.&lt;/p&gt;

&lt;h2&gt;
  
  
  From Finding to PR Fix
&lt;/h2&gt;

&lt;p&gt;The workflow that actually closes findings is not complicated, it just has to be code-precise.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Triage on exploitability, not raw CVSS.&lt;/strong&gt; A working PoC makes it P0. Pre-auth jumps the queue. PHI, PII, or payment data in the blast radius means fix now.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Patch at the layer that survives a re-test.&lt;/strong&gt; Enforce the invariant at the query, not in a decorator:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Survives re-test: ownership enforced at the query level
&lt;/span&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;get_document&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;doc_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;current_user&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;query&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;Document&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;filter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;Document&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nb"&gt;id&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;doc_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;Document&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;owner_id&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;current_user&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nb"&gt;id&lt;/span&gt;  &lt;span class="c1"&gt;# explicit ownership
&lt;/span&gt;    &lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;first_or_404&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Parameterize input rather than escaping it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Fails re-test: escaping is fragile
&lt;/span&gt;&lt;span class="n"&gt;query&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;SELECT * FROM users WHERE email = &lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nf"&gt;sanitize&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;email&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;'"&lt;/span&gt;

&lt;span class="c1"&gt;# Survives re-test: parameterization enforced by the driver
&lt;/span&gt;&lt;span class="n"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;execute&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;SELECT * FROM users WHERE email = ?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;email&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Lock the fix with a negative test, then re-scan:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;test_document_access_requires_ownership&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="n"&gt;user_a&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;create_user&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="n"&gt;user_b&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;create_user&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="n"&gt;doc&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;create_document&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;owner&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;user_a&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/api/documents/&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;doc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nb"&gt;id&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;auth&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;user_b&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;assert&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;status_code&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="mi"&gt;403&lt;/span&gt;  &lt;span class="c1"&gt;# negative test
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  When Each Approach Fits
&lt;/h2&gt;

&lt;p&gt;Not every team needs the same thing.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Traditional firms&lt;/strong&gt; fit stable surfaces, quarterly-or-less releases, and bespoke red-team work. Trade-off: 2 to 4 week turnaround, separate re-test fees, findings as reports not tickets.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Continuous monitoring&lt;/strong&gt; fits SaaS that needs 24/7 surface detection. Trade-off: flags potential issues but lacks exploitation depth and code-level remediation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Code-aware offensive testing&lt;/strong&gt; fits weekly deploys, API-driven architectures, and continuous-evidence compliance. Findings map to code, chains get built, re-scans are fast.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you ship weekly, handle regulated data, and cannot easily translate reports into developer tickets, the code-aware model is probably the right fit.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Takeaway
&lt;/h2&gt;

&lt;p&gt;External penetration testing earns its place for compliance and perimeter validation. But a traditional external test alone cannot keep pace with weekly releases. The fix is not to replace it, it is to give it code memory, so a perimeter finding arrives with the line that caused it and the diff that closes it.&lt;/p&gt;

&lt;p&gt;If you want to see the gap on your own app, the fastest way is to point a code-aware scan at your ten most sensitive internet-facing endpoints and compare what comes back to what your last report said.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;This post was originally published on the &lt;a href="https://codeant.ai/blogs/external-penetration-testing-what-it-covers-and-misses" rel="noopener noreferrer"&gt;CodeAnt AI blog&lt;/a&gt;, where the full version includes the complete six-gap breakdown, the compliance evidence workflow, and an FAQ. If you found this useful, a follow or a comment on what your team runs today is always appreciated.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>security</category>
      <category>pentesting</category>
      <category>webdev</category>
      <category>devsecops</category>
    </item>
    <item>
      <title>GitLab vs GitHub Is Not a Code Review Decision</title>
      <dc:creator>Amartya Jha</dc:creator>
      <pubDate>Wed, 16 Sep 2026 06:12:13 +0000</pubDate>
      <link>https://dev.to/codeant-security/gitlab-vs-github-is-not-a-code-review-decision-5a25</link>
      <guid>https://dev.to/codeant-security/gitlab-vs-github-is-not-a-code-review-decision-5a25</guid>
      <description>&lt;p&gt;Every GitLab vs GitHub comparison is really a pricing comparison wearing a feature table. They are useful if you are picking a platform from nothing, which is almost nobody reading them.&lt;/p&gt;

&lt;p&gt;The question a team already on one of these platforms actually has is narrower. How does review differ, and is the difference worth doing anything about?&lt;/p&gt;

&lt;p&gt;Having looked at both properly, the answer is that they are far closer than the comparisons suggest, and where they differ they differ in opposite directions.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start With What Makes This Comparison Different
&lt;/h2&gt;

&lt;p&gt;On Bitbucket and Azure DevOps, adding a review tool fills a hole the platform left open. Neither ships meaningful native security scanning.&lt;/p&gt;

&lt;p&gt;GitLab and GitHub are the two platforms that do. GitLab has SAST, secret detection, dependency scanning, container scanning and DAST built in as pipeline templates. GitHub has Advanced Security covering secret scanning, dependency review and CodeQL.&lt;/p&gt;

&lt;p&gt;So this is not a comparison about who has scanning. It is about tier, depth and noise, and both platforms put their deeper capabilities behind their upper plans.&lt;/p&gt;

&lt;h2&gt;
  
  
  GitLab's Real Advantage Is Policy Expression
&lt;/h2&gt;

&lt;p&gt;GitHub's model is essentially one rule: require N approvals, and require code owners where CODEOWNERS matches.&lt;/p&gt;

&lt;p&gt;GitLab layers named approval rules on top of CODEOWNERS. You can define a Security rule requiring two approvals from one group and a Database rule requiring one from another, both applying independently to the same merge request.&lt;/p&gt;

&lt;p&gt;It also supports sections in the CODEOWNERS file itself:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight codeowners"&gt;&lt;code&gt;&lt;span class="nn"&gt;[Security]&lt;/span&gt;&lt;span class="m"&gt;[2]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="n"&gt;/src/auth/&lt;/span&gt;&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="nf"&gt;@security-team&lt;/span&gt;&lt;span class="w"&gt;

&lt;/span&gt;&lt;span class="nn"&gt;[Database]&lt;/span&gt;&lt;span class="m"&gt;[1]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="n"&gt;/db/migrations/&lt;/span&gt;&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="nf"&gt;@data-team&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;[2]&lt;/code&gt; requires two approvals from that section. There is no GitHub equivalent, and on GitHub the same intent takes two places: a file saying who, and a branch protection setting saying how many.&lt;/p&gt;

&lt;p&gt;For a small team that difference is academic. For an organisation whose review policy was negotiated with a compliance function, it is the difference between the policy being enforced and the policy being approximated.&lt;/p&gt;

&lt;h2&gt;
  
  
  GitLab Also Has Two Scanners GitHub Does Not
&lt;/h2&gt;

&lt;p&gt;Container scanning and DAST have no direct GitHub equivalent, and this almost never appears in these comparisons.&lt;/p&gt;

&lt;p&gt;There is a structural difference in how the scanning runs, too. GitLab makes security a pipeline job:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;include&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;template&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Jobs/SAST.gitlab-ci.yml&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;template&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Jobs/Secret-Detection.gitlab-ci.yml&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;template&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Jobs/Dependency-Scanning.gitlab-ci.yml&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is the whole basic setup, and the elegance is real: security config is versioned in the repo, reviewed in a merge request, and one hosted template can be shared across an entire estate and changed centrally.&lt;/p&gt;

&lt;p&gt;GitHub runs most of Advanced Security as platform settings rather than jobs. Which means the audit question changes from "show me the commit" to "show me who had admin".&lt;/p&gt;

&lt;h2&gt;
  
  
  GitHub's Advantage Is Everything Around It
&lt;/h2&gt;

&lt;p&gt;The Actions marketplace is substantially larger, and that matters more than it sounds like it should.&lt;/p&gt;

&lt;p&gt;A security vendor with a polished GitHub Action and a thin GitLab template will produce a materially worse experience on GitLab, regardless of how good the underlying analysis is. That is a vendor decision rather than a platform one, and it is worth checking per tool rather than assuming parity.&lt;/p&gt;

&lt;p&gt;Ecosystem familiarity is the other half. If your open-source contributors, your new hires and your tooling all assume GitHub, that is a real cost to being elsewhere.&lt;/p&gt;

&lt;h2&gt;
  
  
  Both Native AI Reviewers Have the Same Two Gaps
&lt;/h2&gt;

&lt;p&gt;GitLab Duo and GitHub Copilot code review are more alike than either vendor suggests.&lt;/p&gt;

&lt;p&gt;Both are general models reading a diff. Neither performs dedicated static security analysis. And on neither platform does an AI comment satisfy a required approval.&lt;/p&gt;

&lt;p&gt;That last point is the one worth internalising. An AI reviewer that comments improves the conversation on a change. It does not alter what can merge. On both platforms, the things that gate are a human approval and a pipeline or status check, and the AI reviewer is neither.&lt;/p&gt;

&lt;p&gt;The licensing differs. Duo is an add-on on top of a paid tier, so two per-seat line items. Copilot is a subscription. Neither replaces the platform's separate security product.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Thing Neither Platform Solves
&lt;/h2&gt;

&lt;p&gt;Both are good at producing findings. Neither tells you which findings an attacker could actually reach in your codebase.&lt;/p&gt;

&lt;p&gt;Severity measures impact if exploited. It says nothing about whether exploitation is possible on a path your code actually calls. The gap between those two is where every security backlog comes from, and it grows at whatever rate your scanners run.&lt;/p&gt;

&lt;p&gt;That is the problem worth solving, and it is not a platform choice. It is a tooling one. &lt;a href="https://codeant.ai/code-security" rel="noopener noreferrer"&gt;CodeAnt AI&lt;/a&gt; runs review and security in the same pass on either platform, checks findings for reachability before a developer sees them, and posts a status check the merge can gate on.&lt;/p&gt;

&lt;h2&gt;
  
  
  So Which One
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;GitLab&lt;/strong&gt; if you want one application covering source, CI and security rather than assembling them, if self-managed is a requirement rather than a fallback, or if your review policy is complex enough that approval rules express it better than a reviewer count.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;GitHub&lt;/strong&gt; if ecosystem reach matters, if the Actions marketplace saves you building integrations, or if open source and contributor familiarity are part of the calculation.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;If you already run one&lt;/strong&gt;, the review differences are not a migration case. Approval rules are nice. Container scanning is genuinely useful. Neither justifies moving a platform every workflow in your organisation is built around.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Actually Useful Exercise
&lt;/h2&gt;

&lt;p&gt;Forget the comparison. Open your most active repository and list what runs on a merge request today.&lt;/p&gt;

&lt;p&gt;AI review, yes or no. SAST, yes or no. Secret detection, yes or no. Dependency and container scanning, yes or no.&lt;/p&gt;

&lt;p&gt;Then check which of those can block the merge rather than only reporting.&lt;/p&gt;

&lt;p&gt;Most teams find the list shorter than they expected, and the gap is almost never the platform.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://codeant.ai/blogs/gitlab-vs-github-code-review" rel="noopener noreferrer"&gt;codeant.ai&lt;/a&gt;, with the full feature comparison, a side-by-side walkthrough of the same change on both, and a migration mapping if you are moving between them.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>gitlab</category>
      <category>github</category>
      <category>ai</category>
      <category>devops</category>
    </item>
    <item>
      <title>I Automated Copilot Code Review on Every Pull Request. Here's What It Still Can't Do</title>
      <dc:creator>Amartya Jha</dc:creator>
      <pubDate>Wed, 16 Sep 2026 06:04:23 +0000</pubDate>
      <link>https://dev.to/codeant-security/i-automated-copilot-code-review-on-every-pull-request-heres-what-it-still-cant-do-4fa1</link>
      <guid>https://dev.to/codeant-security/i-automated-copilot-code-review-on-every-pull-request-heres-what-it-still-cant-do-4fa1</guid>
      <description>&lt;p&gt;The first time I turned on Copilot code review, I watched it comment on a pull request within ninety seconds, felt genuinely impressed, and then spent ten minutes working out why the merge button was still greyed out.&lt;/p&gt;

&lt;p&gt;That gap between what it does and what it changes is the whole story. Copilot reviews. It does not approve, it does not gate, and it does not look for the class of problem that actually ends up in an incident report.&lt;/p&gt;

&lt;p&gt;None of that makes it a bad tool. It makes it a tool with a shape, and knowing the shape before you build a workflow around it saves the conversation I had with my own team three weeks later.&lt;/p&gt;

&lt;p&gt;Here is the setup, and then the honest part.&lt;/p&gt;

&lt;h2&gt;
  
  
  Getting It Running
&lt;/h2&gt;

&lt;p&gt;You need a paid Copilot plan, and GitHub CLI v2.88.0 or later if you want to do this from a terminal.&lt;/p&gt;

&lt;p&gt;Requesting a review is the same gesture as asking a colleague. Open the pull request, find Reviewers in the sidebar, pick Copilot. Or from the command line:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;gh &lt;span class="nb"&gt;pr &lt;/span&gt;create &lt;span class="nt"&gt;--fill&lt;/span&gt; &lt;span class="nt"&gt;--reviewer&lt;/span&gt; @copilot
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;@&lt;/code&gt; matters. Older CLI versions will not accept it at all.&lt;/p&gt;

&lt;p&gt;One thing nobody tells you: the review runs as a GitHub Actions workflow. When a review does not appear, the Actions tab has a run called &lt;em&gt;Copilot code review&lt;/em&gt; with real logs, and that is where the answer usually is. I spent an embarrassing amount of time looking everywhere else first.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Part That Actually Saves Time
&lt;/h2&gt;

&lt;p&gt;Copilot's comments often come with a suggestion block you can commit straight from the pull request. Useful, and applying them one at a time produces a commit per fix, which turns a tidy branch into fourteen commits called "Potential fix for pull request finding".&lt;/p&gt;

&lt;p&gt;Use the Files changed tab and &lt;strong&gt;Add suggestion to batch&lt;/strong&gt; instead. Queue everything, then commit once.&lt;/p&gt;

&lt;p&gt;If several comments have no ready-made suggestion, or a fix spans files, there is a third option on the summary comment: &lt;strong&gt;Fix batch with Copilot&lt;/strong&gt;. The coding agent works server-side and pushes to your branch a few minutes later. That handles what suggestion blocks cannot, because it writes changes rather than committing pre-written ones.&lt;/p&gt;

&lt;p&gt;You can also argue with it. Reply mentioning &lt;code&gt;@copilot&lt;/code&gt; in the thread and it responds, having read the whole conversation. Sometimes it concedes. Sometimes it explains something you had not considered. Occasionally it is just wrong, and you resolve the thread and move on.&lt;/p&gt;

&lt;h2&gt;
  
  
  Making It Automatic Without Making It Annoying
&lt;/h2&gt;

&lt;p&gt;Manual requesting works until someone forgets, which means review coverage depends on who opened the pull request. That is a strange property for a quality process.&lt;/p&gt;

&lt;p&gt;Repository Settings, then Rules, then Rulesets. Create a branch ruleset targeting your default branch and enable &lt;strong&gt;Automatically request Copilot code review&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The sub-setting underneath it is the one people miss: re-run on new commits. It is off by default, which is why teams end up with reviews that describe code from three pushes ago.&lt;/p&gt;

&lt;p&gt;Two other levels exist and interact. An organisation policy overrides everything, so if an admin has turned it off nobody can opt in. And there is a personal toggle in your own Copilot settings that applies to pull requests you open, which the org policy also overrides.&lt;/p&gt;

&lt;h2&gt;
  
  
  Tuning It So It Stops Repeating Itself
&lt;/h2&gt;

&lt;p&gt;Out of the box Copilot reviews against a generic idea of good code. A &lt;code&gt;.github/copilot-instructions.md&lt;/code&gt; file changes that.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gh"&gt;# Review instructions&lt;/span&gt;
&lt;span class="p"&gt;
-&lt;/span&gt; This is an ASP.NET Core API. Model binding validates inputs at the
  framework level, so do not flag interpolated strings in parameterised
  query builders.
&lt;span class="p"&gt;-&lt;/span&gt; Prefer flagging missing null checks on public method parameters over
  naming preferences.
&lt;span class="p"&gt;-&lt;/span&gt; Do not comment on formatting. Prettier handles it.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Keep it short. Three to five bullets beat a long document, because long instructions get squeezed out of the context budget and stop mattering.&lt;/p&gt;

&lt;p&gt;The highest-value thing you can write in that file is what your framework already guarantees. That is the biggest single source of comments you will otherwise resolve every week for the rest of the project.&lt;/p&gt;

&lt;p&gt;For larger codebases, per-path files under &lt;code&gt;.github/instructions/&lt;/code&gt; with an &lt;code&gt;applyTo&lt;/code&gt; glob in the front matter let you give frontend and backend different guidance. They stack on the project-wide file.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Billing Thing Nobody Mentions Until the Invoice
&lt;/h2&gt;

&lt;p&gt;Copilot moved to usage-based billing in June 2026. Premium request units were replaced by GitHub AI Credits, metered on tokens, and code review carries a published multiplier of 13 against the requests meter.&lt;/p&gt;

&lt;p&gt;Separately, because code review now runs on Actions, reviews on private repositories consume Actions minutes at the normal rate. Public repositories are exempt.&lt;/p&gt;

&lt;p&gt;Which means if you enable automatic reviews on a repository where Dependabot opens forty pull requests a week, you are now spending from the same Actions pool your CI uses. Worth checking your spending limits before rolling it out broadly rather than after.&lt;/p&gt;

&lt;h2&gt;
  
  
  Now the Honest Part
&lt;/h2&gt;

&lt;p&gt;Four things Copilot code review does not do, and each one is why teams end up adding something alongside it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It cannot approve.&lt;/strong&gt; Its output is comments, never an approval event. A branch protection rule requiring one approving review is not satisfied by it, and it cannot satisfy a &lt;code&gt;CODEOWNERS&lt;/code&gt; requirement either. A critical issue sitting in a Copilot comment merges exactly as easily as no comment at all, unless a person reads it and acts.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It is not a security scanner.&lt;/strong&gt; It reasons about the change. It does not perform taint analysis, secret detection, dependency scanning, or infrastructure-as-code checks. GitHub has those as Advanced Security, a separately licensed product Copilot neither replaces nor surfaces. If you adopted Copilot thinking AI review covered security, that gap is invisible until it is not.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It is scoped to the diff.&lt;/strong&gt; It reads the change and nearby context. It is not indexing your repository to trace how a changed function is called three services away, and no amount of instruction tuning fixes that.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It does not track anything.&lt;/strong&gt; No finding state, no ownership, no resolution timestamp. A comment is a comment. If you need to prove findings were resolved within a window, that is a different category of tool entirely.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Same Feature Is Different on Azure DevOps
&lt;/h2&gt;

&lt;p&gt;Worth knowing if your organisation runs both, which plenty do after an acquisition.&lt;/p&gt;

&lt;p&gt;Copilot code review on Azure DevOps is a separate implementation with the same name. It is in public preview with staged regional rollout, and it carries hard caps: 10 GB repository, 100 changed files, the pull request must be active with no merge conflicts. Enabling it takes three levels, organisation then project then repository, and the automatic-review setting lives under branch policies rather than anywhere you would look first.&lt;/p&gt;

&lt;p&gt;It also cannot satisfy a required reviewer or block a merge, same as on GitHub. So on the platform whose main strength is policy enforcement, the native AI reviewer sits entirely outside the policy layer.&lt;/p&gt;

&lt;p&gt;GitLab's Duo and Bitbucket's Rovo Dev have the same two properties. Different vendors, different licensing, and none of them does dedicated security analysis or changes what can merge.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Actually Do Now
&lt;/h2&gt;

&lt;p&gt;Copilot runs on every pull request through a ruleset, with a short instructions file telling it what our framework already handles. It gives every change a baseline read within a minute, which on a distributed team is genuinely better than waiting three hours for someone to open the tab.&lt;/p&gt;

&lt;p&gt;And separately, a tool that posts a status check rather than a comment handles the part a comment cannot: security scanning that runs in the same pass, findings checked for whether they are actually reachable, and a gate that holds the merge when something critical shows up. That is &lt;a href="https://codeant.ai/ai-code-review" rel="noopener noreferrer"&gt;what CodeAnt AI does&lt;/a&gt;, across GitHub, GitLab, Bitbucket and Azure DevOps, including self-hosted.&lt;/p&gt;

&lt;p&gt;The distinction I would offer anyone setting this up: a fast reviewer and a control are two different things, and Copilot is emphatically the first one. Work out which you were trying to buy.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://codeant.ai/blogs/github-copilot-code-review" rel="noopener noreferrer"&gt;codeant.ai&lt;/a&gt;, where there is a fuller version including troubleshooting and the &lt;a href="https://codeant.ai/blogs/azure-devops-vs-github-code-review" rel="noopener noreferrer"&gt;platform-by-platform comparison&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>githubcopilot</category>
      <category>github</category>
      <category>ai</category>
      <category>programming</category>
    </item>
    <item>
      <title>Why Your $6,500 Penetration Test Will Cost You More Than the $47,000 One</title>
      <dc:creator>Amartya Jha</dc:creator>
      <pubDate>Wed, 09 Sep 2026 11:09:58 +0000</pubDate>
      <link>https://dev.to/codeant/why-your-6500-penetration-test-will-cost-you-more-than-the-47000-one-2fhe</link>
      <guid>https://dev.to/codeant/why-your-6500-penetration-test-will-cost-you-more-than-the-47000-one-2fhe</guid>
      <description>&lt;p&gt;Two penetration testing quotes land in your inbox.&lt;/p&gt;

&lt;p&gt;One says &lt;strong&gt;$6,500&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The other says &lt;strong&gt;$22,000&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Both vendors say they use AI. Both promise comprehensive coverage. Both have credible-looking reports, experienced consultants, and a convincing sales call.&lt;/p&gt;

&lt;p&gt;So why is one more than 3× the price?&lt;/p&gt;

&lt;p&gt;Because &lt;strong&gt;“penetration test” isn't a standardized unit of work.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The difference usually isn't the PDF at the end. It's everything that happens before the PDF exists: how much of the attack surface gets mapped, how deeply authentication and authorization are tested, whether findings are manually validated, whether vulnerabilities are chained into attack paths, whether source code is reviewed, and what happens after the report lands.&lt;/p&gt;

&lt;p&gt;This matters even more when the pentest is being used for compliance, enterprise security reviews, or a customer audit.&lt;/p&gt;

&lt;p&gt;A cheap assessment can absolutely be useful. A $6,500 test isn't automatically bad, just as a $22,000 test isn't automatically good.&lt;/p&gt;

&lt;p&gt;The question is:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What are you actually buying at each price point?&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The Quick Numbers
&lt;/h2&gt;

&lt;p&gt;For planning purposes, penetration testing in 2026 can range from roughly &lt;strong&gt;$3,000 for a limited external assessment to $100,000+ for a large, integrated engagement&lt;/strong&gt; covering applications, APIs, infrastructure, source code, and multiple testing perspectives.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Penetration test type&lt;/th&gt;
&lt;th&gt;Typical 2026 planning range&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Limited black box / external assessment&lt;/td&gt;
&lt;td&gt;$3,000–$8,000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Full black box web application assessment&lt;/td&gt;
&lt;td&gt;$8,000–$25,000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gray box / authenticated assessment&lt;/td&gt;
&lt;td&gt;$10,000–$40,000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;White box / source code assessment&lt;/td&gt;
&lt;td&gt;$15,000–$60,000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Integrated black + gray + white box&lt;/td&gt;
&lt;td&gt;$25,000–$100,000+&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Continuous / recurring assessment&lt;/td&gt;
&lt;td&gt;$5,000–$20,000/month&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;These aren't standardized market prices. Actual pricing varies substantially by scope, application complexity, number of assets, number of roles, source-code size, infrastructure, methodology, tester involvement, and deliverables.&lt;/p&gt;

&lt;p&gt;And that's exactly why comparing two quotes by the final number alone doesn't work.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Scope Is Usually Where the Difference Starts
&lt;/h2&gt;

&lt;p&gt;“Web application penetration test” sounds specific.&lt;/p&gt;

&lt;p&gt;It isn't.&lt;/p&gt;

&lt;p&gt;One vendor might mean:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;One production application and its public login page.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Another might mean:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The web application, API, authenticated endpoints, staging environment, cloud assets, JavaScript bundles, authentication flows, and all discovered subdomains.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Those are completely different engagements.&lt;/p&gt;

&lt;p&gt;Before comparing prices, turn the scope into something measurable.&lt;/p&gt;

&lt;p&gt;Ask:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;How many applications are included?&lt;/li&gt;
&lt;li&gt;Which production and staging environments?&lt;/li&gt;
&lt;li&gt;Which domains and subdomains?&lt;/li&gt;
&lt;li&gt;How many APIs and endpoints?&lt;/li&gt;
&lt;li&gt;Are authenticated endpoints included?&lt;/li&gt;
&lt;li&gt;How many user roles?&lt;/li&gt;
&lt;li&gt;Is multi-tenant isolation tested?&lt;/li&gt;
&lt;li&gt;Is cloud infrastructure included?&lt;/li&gt;
&lt;li&gt;Are source repositories included?&lt;/li&gt;
&lt;li&gt;Are third-party integrations in scope?&lt;/li&gt;
&lt;li&gt;What is explicitly excluded?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A written scope document is more useful than a paragraph in a sales proposal.&lt;/p&gt;

&lt;h2&gt;
  
  
  Methodology Depth Changes the Price
&lt;/h2&gt;

&lt;p&gt;The second major variable is &lt;strong&gt;how the testing is performed&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Two vendors can test the same application and produce very different results.&lt;/p&gt;

&lt;p&gt;A shallow engagement might look like:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Run automated scanners.&lt;/li&gt;
&lt;li&gt;Identify vulnerable versions.&lt;/li&gt;
&lt;li&gt;Match responses against known signatures.&lt;/li&gt;
&lt;li&gt;Generate findings.&lt;/li&gt;
&lt;li&gt;Deliver a PDF.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;A deeper assessment might involve:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Passive reconnaissance.&lt;/li&gt;
&lt;li&gt;Attack-surface enumeration.&lt;/li&gt;
&lt;li&gt;Subdomain and cloud asset discovery.&lt;/li&gt;
&lt;li&gt;API endpoint mapping.&lt;/li&gt;
&lt;li&gt;Authentication testing.&lt;/li&gt;
&lt;li&gt;Authorization testing.&lt;/li&gt;
&lt;li&gt;Manual business-logic testing.&lt;/li&gt;
&lt;li&gt;Controlled exploitation.&lt;/li&gt;
&lt;li&gt;Privilege escalation.&lt;/li&gt;
&lt;li&gt;Cross-tenant access testing.&lt;/li&gt;
&lt;li&gt;Attack-path construction.&lt;/li&gt;
&lt;li&gt;Source-level investigation.&lt;/li&gt;
&lt;li&gt;Remediation validation.&lt;/li&gt;
&lt;li&gt;Retesting.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The target can be identical.&lt;/p&gt;

&lt;p&gt;The amount of work is not.&lt;/p&gt;

&lt;p&gt;That's why &lt;strong&gt;price alone is a poor proxy for pentest quality&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  $3,000–$8,000: Baseline External Testing
&lt;/h2&gt;

&lt;p&gt;At the lower end of the market, you're generally getting a limited external assessment.&lt;/p&gt;

&lt;p&gt;The exact methodology varies, but this tier may emphasize:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Automated vulnerability discovery&lt;/li&gt;
&lt;li&gt;Known-CVE detection&lt;/li&gt;
&lt;li&gt;Technology fingerprinting&lt;/li&gt;
&lt;li&gt;Basic port and service enumeration&lt;/li&gt;
&lt;li&gt;TLS configuration checks&lt;/li&gt;
&lt;li&gt;Common web vulnerabilities&lt;/li&gt;
&lt;li&gt;Structured reporting&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This can be valuable for establishing baseline visibility.&lt;/p&gt;

&lt;p&gt;But buyers should check whether the engagement includes &lt;strong&gt;manual exploitation and validation&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;A scanner saying:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“Potential SQL injection detected”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;is not equivalent to a tester demonstrating that a controllable parameter reaches a SQL query and validating the impact safely.&lt;/p&gt;

&lt;p&gt;The distinction is:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;potential vulnerability vs. confirmed vulnerability.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;If the quote is inexpensive because manual validation isn't included, that's important to know before comparing it with a much more hands-on assessment.&lt;/p&gt;

&lt;h2&gt;
  
  
  $8,000–$25,000: Full External / Black Box Testing
&lt;/h2&gt;

&lt;p&gt;This is where the engagement can become a genuine methodology-driven external pentest.&lt;/p&gt;

&lt;p&gt;Depending on scope, testing may include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;DNS enumeration&lt;/li&gt;
&lt;li&gt;Subdomain discovery&lt;/li&gt;
&lt;li&gt;Certificate transparency analysis&lt;/li&gt;
&lt;li&gt;Port and service enumeration&lt;/li&gt;
&lt;li&gt;Web application fingerprinting&lt;/li&gt;
&lt;li&gt;JavaScript analysis&lt;/li&gt;
&lt;li&gt;Endpoint discovery&lt;/li&gt;
&lt;li&gt;API enumeration&lt;/li&gt;
&lt;li&gt;Authentication testing&lt;/li&gt;
&lt;li&gt;Authorization testing&lt;/li&gt;
&lt;li&gt;Common injection classes&lt;/li&gt;
&lt;li&gt;SSRF&lt;/li&gt;
&lt;li&gt;File upload vulnerabilities&lt;/li&gt;
&lt;li&gt;Access-control weaknesses&lt;/li&gt;
&lt;li&gt;Business-logic testing&lt;/li&gt;
&lt;li&gt;Cloud exposure&lt;/li&gt;
&lt;li&gt;Manual exploitation&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The important question isn't whether a proposal contains the word &lt;strong&gt;“manual.”&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Ask &lt;strong&gt;what is actually manual&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Are authenticated API endpoints manually tested?&lt;/p&gt;

&lt;p&gt;Are authorization boundaries tested across roles?&lt;/p&gt;

&lt;p&gt;Are findings reproduced by a human tester?&lt;/p&gt;

&lt;p&gt;Are exploit chains investigated?&lt;/p&gt;

&lt;p&gt;Are business-logic workflows tested?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Those questions reveal far more than a generic “OWASP methodology” statement.&lt;/p&gt;

&lt;h2&gt;
  
  
  $10,000–$40,000: Gray Box / Authenticated Testing
&lt;/h2&gt;

&lt;p&gt;Gray box testing introduces something external testing can't reproduce:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;legitimate application context.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Instead of asking only:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“What can an anonymous attacker access?”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;the tester can ask:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“What can this authenticated user access that they shouldn't?”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This is where some of the most interesting authorization vulnerabilities appear.&lt;/p&gt;

&lt;p&gt;Consider a SaaS application with:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Admin&lt;/li&gt;
&lt;li&gt;Manager&lt;/li&gt;
&lt;li&gt;Employee&lt;/li&gt;
&lt;li&gt;Read-only user&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;and thousands of tenant-specific resources.&lt;/p&gt;

&lt;p&gt;Testing isn't simply:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;GET /api/invoices/1234
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The tester needs to establish:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;User A → Tenant A → Resource A
User A → Tenant B → Resource B
Manager → Admin functionality
Employee → Manager functionality
Read-only → Write functionality
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The goal is to determine whether the application's &lt;strong&gt;authorization model matches the intended trust boundaries&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;This is where vulnerabilities such as BOLA/IDOR, privilege escalation, broken function-level authorization, and cross-tenant access become particularly important.&lt;/p&gt;

&lt;p&gt;And the number of roles and workflows can dramatically increase the testing effort.&lt;/p&gt;

&lt;h2&gt;
  
  
  $15,000–$60,000: White Box / Source Code Testing
&lt;/h2&gt;

&lt;p&gt;Source access changes the problem entirely.&lt;/p&gt;

&lt;p&gt;A black box tester sees:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;HTTP request
      ↓
HTTP response
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A white box tester can investigate:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;HTTP request
      ↓
Controller
      ↓
Middleware
      ↓
Authorization check
      ↓
Service
      ↓
Database query
      ↓
Dangerous sink
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That makes vulnerabilities visible that may be difficult or impossible to identify externally.&lt;/p&gt;

&lt;p&gt;A serious source-code assessment can include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Authentication implementation&lt;/li&gt;
&lt;li&gt;Authorization middleware&lt;/li&gt;
&lt;li&gt;API controllers&lt;/li&gt;
&lt;li&gt;Input validation&lt;/li&gt;
&lt;li&gt;Dataflow analysis&lt;/li&gt;
&lt;li&gt;Sensitive sinks&lt;/li&gt;
&lt;li&gt;Cryptographic usage&lt;/li&gt;
&lt;li&gt;Secrets&lt;/li&gt;
&lt;li&gt;Dependency usage&lt;/li&gt;
&lt;li&gt;Git history&lt;/li&gt;
&lt;li&gt;Infrastructure-as-code&lt;/li&gt;
&lt;li&gt;CI/CD configuration&lt;/li&gt;
&lt;li&gt;Security-sensitive framework configuration&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The depth of this work matters enormously.&lt;/p&gt;

&lt;p&gt;“Source code included” doesn't necessarily mean:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“Every line of the repository will be manually reviewed.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Ask what is actually being analyzed and how findings are traced back to the implementation.&lt;/p&gt;

&lt;h2&gt;
  
  
  $25,000–$100,000+: Integrated Testing
&lt;/h2&gt;

&lt;p&gt;The most useful full assessments aren't simply:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;black box + gray box + white box.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;They are integrated.&lt;/p&gt;

&lt;p&gt;Information discovered during one phase informs another.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;External recon
      ↓
API discovered
      ↓
Authenticated access obtained
      ↓
Authorization weakness identified
      ↓
Source code traced
      ↓
Attack path constructed
      ↓
Fix implemented
      ↓
Production retest
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That produces a much stronger understanding of the actual risk than three disconnected reports.&lt;/p&gt;

&lt;p&gt;For a mid-sized SaaS application, an integrated engagement might fall somewhere around $25,000–$50,000.&lt;/p&gt;

&lt;p&gt;Larger applications with multiple services, environments, APIs, cloud infrastructure, and complex authorization models can move toward $50,000–$100,000+.&lt;/p&gt;

&lt;p&gt;Enterprise-scale testing can go beyond that.&lt;/p&gt;

&lt;h2&gt;
  
  
  Continuous Testing: $5,000–$20,000/Month
&lt;/h2&gt;

&lt;p&gt;Annual pentesting has an obvious limitation.&lt;/p&gt;

&lt;p&gt;It tells you what was exposed &lt;strong&gt;when the test happened&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;If your engineering team deploys every week, the application tested in January may not resemble the application running in December.&lt;/p&gt;

&lt;p&gt;A new API gets deployed.&lt;/p&gt;

&lt;p&gt;A new subdomain appears.&lt;/p&gt;

&lt;p&gt;A cloud bucket gets created.&lt;/p&gt;

&lt;p&gt;An authentication flow changes.&lt;/p&gt;

&lt;p&gt;A service gets exposed.&lt;/p&gt;

&lt;p&gt;A dependency changes.&lt;/p&gt;

&lt;p&gt;The attack surface moves continuously.&lt;/p&gt;

&lt;p&gt;Recurring testing addresses that problem by shortening the time between changes and security validation.&lt;/p&gt;

&lt;p&gt;It doesn't necessarily replace a comprehensive periodic assessment.&lt;/p&gt;

&lt;p&gt;Instead:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Periodic pentest = deep assessment&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Continuous testing = ongoing validation&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The two can complement each other.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cost vs. Coverage
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Engagement&lt;/th&gt;
&lt;th&gt;Typical range&lt;/th&gt;
&lt;th&gt;Strongest at finding&lt;/th&gt;
&lt;th&gt;Common gap&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Automated assessment&lt;/td&gt;
&lt;td&gt;$3K–$8K&lt;/td&gt;
&lt;td&gt;Known technical weaknesses&lt;/td&gt;
&lt;td&gt;Business logic and attack chains&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Black box&lt;/td&gt;
&lt;td&gt;$8K–$25K&lt;/td&gt;
&lt;td&gt;External attack surface&lt;/td&gt;
&lt;td&gt;Source-level vulnerabilities&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gray box&lt;/td&gt;
&lt;td&gt;$10K–$40K&lt;/td&gt;
&lt;td&gt;Authorization and authenticated workflows&lt;/td&gt;
&lt;td&gt;Internal implementation details&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;White box&lt;/td&gt;
&lt;td&gt;$15K–$60K&lt;/td&gt;
&lt;td&gt;Code-level vulnerabilities&lt;/td&gt;
&lt;td&gt;Runtime behavior without execution&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Integrated&lt;/td&gt;
&lt;td&gt;$25K–$100K+&lt;/td&gt;
&lt;td&gt;Cross-layer attack paths&lt;/td&gt;
&lt;td&gt;Primarily constrained by scope&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Continuous&lt;/td&gt;
&lt;td&gt;$5K–$20K/mo&lt;/td&gt;
&lt;td&gt;Changes introduced over time&lt;/td&gt;
&lt;td&gt;Depth depends on cadence and methodology&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The important word in this table is &lt;strong&gt;gap&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Every testing model sees some things better than others.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Finding Quality Test
&lt;/h2&gt;

&lt;p&gt;Here's one of the easiest ways to compare vendors.&lt;/p&gt;

&lt;p&gt;Ask for a &lt;strong&gt;redacted sample technical finding&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Look at what you're actually getting.&lt;/p&gt;

&lt;p&gt;A useful finding should tell an engineer:&lt;/p&gt;

&lt;h3&gt;
  
  
  What is vulnerable?
&lt;/h3&gt;

&lt;p&gt;Not:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“Application may be vulnerable to authorization issues.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;But:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;GET /api/v1/invoices/{id}&lt;/code&gt; does not verify that the authenticated user belongs to the invoice's tenant.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  How was it validated?
&lt;/h3&gt;

&lt;p&gt;Show the relevant request, response, prerequisites, and safe proof of impact.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why does it happen?
&lt;/h3&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Request
  ↓
Controller
  ↓
Invoice ID accepted
  ↓
Database lookup by ID
  ↓
Tenant ownership never checked
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  What is the impact?
&lt;/h3&gt;

&lt;p&gt;Can an attacker read another tenant's invoice?&lt;/p&gt;

&lt;p&gt;Modify it?&lt;/p&gt;

&lt;p&gt;Delete it?&lt;/p&gt;

&lt;p&gt;Access sensitive customer data?&lt;/p&gt;

&lt;h3&gt;
  
  
  How should it be fixed?
&lt;/h3&gt;

&lt;p&gt;Ideally, the remediation guidance points engineers toward the actual control that needs to change.&lt;/p&gt;

&lt;p&gt;This is where two reports with the same number of findings can have radically different value.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Attack-Path Question
&lt;/h2&gt;

&lt;p&gt;A pentest shouldn't always treat every vulnerability as an isolated row in a spreadsheet.&lt;/p&gt;

&lt;p&gt;Consider:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Low-privilege account
        ↓
Information disclosure
        ↓
Internal endpoint discovered
        ↓
Authorization bypass
        ↓
Sensitive API access
        ↓
Credential exposure
        ↓
Cloud privilege escalation
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each individual finding might have a moderate severity.&lt;/p&gt;

&lt;p&gt;Together, they can represent a serious attack path.&lt;/p&gt;

&lt;p&gt;Ask the provider:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Do you identify and validate chained attack paths, or do you report vulnerabilities independently?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That answer can be more revealing than the number of tools listed in the proposal.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Retest Question
&lt;/h2&gt;

&lt;p&gt;This is another major difference between quotes.&lt;/p&gt;

&lt;p&gt;Ask:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Is retesting included?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;And more importantly:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What does “retest” actually mean?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A useful retest should verify that:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The vulnerability is no longer exploitable.&lt;/li&gt;
&lt;li&gt;The underlying security control works as intended.&lt;/li&gt;
&lt;li&gt;The fix didn't introduce another bypass.&lt;/li&gt;
&lt;li&gt;The change is present in the environment that actually matters.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;A staging retest doesn't necessarily prove that production is fixed.&lt;/p&gt;

&lt;p&gt;Also clarify how many retest cycles are included and whether they are billed separately.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Compliance Deliverables Gap
&lt;/h2&gt;

&lt;p&gt;The pentest itself isn't always the only deliverable you need.&lt;/p&gt;

&lt;p&gt;Depending on the framework and auditor, you may need supporting evidence around:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Scope&lt;/li&gt;
&lt;li&gt;Testing dates&lt;/li&gt;
&lt;li&gt;Findings&lt;/li&gt;
&lt;li&gt;Severity&lt;/li&gt;
&lt;li&gt;Remediation&lt;/li&gt;
&lt;li&gt;Retesting&lt;/li&gt;
&lt;li&gt;Final status&lt;/li&gt;
&lt;li&gt;Risk acceptance&lt;/li&gt;
&lt;li&gt;Evidence handling&lt;/li&gt;
&lt;li&gt;Data retention&lt;/li&gt;
&lt;li&gt;Data deletion&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Don't assume these documents are automatically included because the proposal says &lt;strong&gt;“compliance-ready report.”&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Ask what you actually receive.&lt;/p&gt;

&lt;p&gt;And ask how sensitive information collected during testing is handled after the engagement.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Checklist to Put Beside Every Quote
&lt;/h2&gt;

&lt;p&gt;Before signing a pentest SOW, make every vendor answer the same questions.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Scope&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Which applications?&lt;/li&gt;
&lt;li&gt;Which environments?&lt;/li&gt;
&lt;li&gt;Which domains and subdomains?&lt;/li&gt;
&lt;li&gt;Which APIs?&lt;/li&gt;
&lt;li&gt;Which third-party integrations?&lt;/li&gt;
&lt;li&gt;What's explicitly excluded?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Authentication&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Are login flows tested?&lt;/li&gt;
&lt;li&gt;MFA?&lt;/li&gt;
&lt;li&gt;Password reset?&lt;/li&gt;
&lt;li&gt;Session management?&lt;/li&gt;
&lt;li&gt;Tokens?&lt;/li&gt;
&lt;li&gt;Account recovery?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Authorization&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Horizontal privilege escalation?&lt;/li&gt;
&lt;li&gt;Vertical privilege escalation?&lt;/li&gt;
&lt;li&gt;BOLA/IDOR?&lt;/li&gt;
&lt;li&gt;Cross-tenant access?&lt;/li&gt;
&lt;li&gt;Role boundaries?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Application logic&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Manual business-logic testing?&lt;/li&gt;
&lt;li&gt;Workflow manipulation?&lt;/li&gt;
&lt;li&gt;Race conditions?&lt;/li&gt;
&lt;li&gt;Abuse cases?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Infrastructure&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;AWS?&lt;/li&gt;
&lt;li&gt;Azure?&lt;/li&gt;
&lt;li&gt;GCP?&lt;/li&gt;
&lt;li&gt;Kubernetes?&lt;/li&gt;
&lt;li&gt;Containers?&lt;/li&gt;
&lt;li&gt;IAM?&lt;/li&gt;
&lt;li&gt;Cloud storage?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Source&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Is source code included?&lt;/li&gt;
&lt;li&gt;How much is reviewed?&lt;/li&gt;
&lt;li&gt;Is dataflow analysis performed?&lt;/li&gt;
&lt;li&gt;Is Git history analyzed?&lt;/li&gt;
&lt;li&gt;Is IaC reviewed?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Validation&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Are findings manually confirmed?&lt;/li&gt;
&lt;li&gt;Is controlled exploitation performed?&lt;/li&gt;
&lt;li&gt;Are attack chains investigated?&lt;/li&gt;
&lt;li&gt;Is exploitability demonstrated safely?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Retesting&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Is it included?&lt;/li&gt;
&lt;li&gt;How many cycles?&lt;/li&gt;
&lt;li&gt;Is production retested?&lt;/li&gt;
&lt;li&gt;Is a final retest report provided?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Evidence&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;What reports are included?&lt;/li&gt;
&lt;li&gt;What compliance documentation ships with the engagement?&lt;/li&gt;
&lt;li&gt;How is sensitive test data stored?&lt;/li&gt;
&lt;li&gt;When is it deleted?&lt;/li&gt;
&lt;li&gt;Is deletion documented?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This turns a vague pricing comparison into an actual &lt;strong&gt;coverage comparison&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  So, Is the $22,000 Pentest Worth It?
&lt;/h2&gt;

&lt;p&gt;Not necessarily.&lt;/p&gt;

&lt;p&gt;The $22,000 test isn't automatically better.&lt;/p&gt;

&lt;p&gt;And the $6,500 test isn't automatically a waste of money.&lt;/p&gt;

&lt;p&gt;The better question is:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What risk does each engagement leave untested?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;If the $6,500 assessment covers only unauthenticated external scanning while the $22,000 engagement includes authenticated APIs, authorization testing, manual business logic testing, cloud assets, exploit validation, attack-path analysis, and retesting, you're not comparing two prices for the same product.&lt;/p&gt;

&lt;p&gt;You're comparing two different security assessments.&lt;/p&gt;

&lt;p&gt;That's the part of the pricing conversation that often gets lost.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Don't compare the number at the top of the quote. Compare the attack surface, methodology, validation, evidence, and lifecycle coverage underneath it.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;h3&gt;
  
  
  How much does a penetration test cost in 2026?
&lt;/h3&gt;

&lt;p&gt;Planning ranges can start around $3,000 for limited external testing and exceed $100,000 for large integrated assessments. The actual price depends heavily on scope, application complexity, authenticated workflows, infrastructure, source-code access, methodology, and deliverables.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why can two pentest quotes differ by 3×?
&lt;/h3&gt;

&lt;p&gt;Because “penetration test” doesn't define a standardized amount of work. One quote may include automated scanning and limited validation, while another may include authenticated testing, source review, manual business logic testing, exploit validation, attack-path analysis, and retesting.&lt;/p&gt;

&lt;h3&gt;
  
  
  Is a cheap penetration test useless?
&lt;/h3&gt;

&lt;p&gt;No. A limited assessment can provide useful baseline visibility. The important thing is understanding what it is designed to detect and what it explicitly doesn't cover.&lt;/p&gt;

&lt;h3&gt;
  
  
  Should retesting be included?
&lt;/h3&gt;

&lt;p&gt;Ideally, yes, or at least the retesting process and cost should be explicit in the SOW. Otherwise, the initial quote may not represent the total cost of completing the testing lifecycle.&lt;/p&gt;

&lt;h3&gt;
  
  
  Is black box, gray box, or white box testing best?
&lt;/h3&gt;

&lt;p&gt;None is universally “best.” Black box provides an outside-in view, gray box is particularly useful for authenticated authorization and workflow testing, and white box exposes implementation-level weaknesses. For high-risk applications, combining perspectives can provide much stronger coverage.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;This piece is adapted from a longer breakdown on the CodeAnt AI blog, including a detailed quote-comparison framework and an analysis of alternative penetration testing pricing models.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>pentesting</category>
      <category>vulnerabilities</category>
      <category>cybersecurity</category>
      <category>penetrationtesting</category>
    </item>
    <item>
      <title>External Penetration Testing in 2026: A Technical Methodology, Tool Stack, and Attack Surface Guide</title>
      <dc:creator>Amartya Jha</dc:creator>
      <pubDate>Wed, 09 Sep 2026 10:46:46 +0000</pubDate>
      <link>https://dev.to/codeant/external-penetration-testing-in-2026-a-technical-methodology-tool-stack-and-attack-surface-guide-10g1</link>
      <guid>https://dev.to/codeant/external-penetration-testing-in-2026-a-technical-methodology-tool-stack-and-attack-surface-guide-10g1</guid>
      <description>&lt;p&gt;External penetration testing is often reduced to a familiar sequence:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;nmap → nuclei → Burp → report
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That workflow is useful, but it misses the hardest part of the engagement.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Finding the assets worth testing.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A production application is rarely the entire external attack surface.&lt;/p&gt;

&lt;p&gt;There may be a staging environment on another subdomain, an old API version on a different hostname, a forgotten VPN portal, a cloud storage bucket, a CI/CD interface, an exposed database, or a service that was intended to be internal but is reachable from the internet.&lt;/p&gt;

&lt;p&gt;If the tester doesn't discover those assets, none of the vulnerability scanners matter.&lt;/p&gt;

&lt;p&gt;This guide presents a technical external penetration testing methodology from the perspective of an attacker starting with only a target organization and its public footprint.&lt;/p&gt;

&lt;p&gt;The focus is on:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Passive reconnaissance&lt;/li&gt;
&lt;li&gt;DNS and certificate intelligence&lt;/li&gt;
&lt;li&gt;Attack-surface discovery&lt;/li&gt;
&lt;li&gt;Active host and port enumeration&lt;/li&gt;
&lt;li&gt;Web and API discovery&lt;/li&gt;
&lt;li&gt;Vulnerability identification&lt;/li&gt;
&lt;li&gt;Manual validation&lt;/li&gt;
&lt;li&gt;Safe exploitation&lt;/li&gt;
&lt;li&gt;Cloud exposure&lt;/li&gt;
&lt;li&gt;Commonly missed attack paths&lt;/li&gt;
&lt;li&gt;Tool selection&lt;/li&gt;
&lt;li&gt;Reporting evidence&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The commands below are intended for systems you are authorized to test. Production systems should be tested within the agreed scope, rate limits, and rules of engagement.&lt;/p&gt;




&lt;h1&gt;
  
  
  1. Define the Attack Surface Before Testing It
&lt;/h1&gt;

&lt;p&gt;The first mistake in external penetration testing is treating the client's domain as the attack surface.&lt;/p&gt;

&lt;p&gt;It isn't.&lt;/p&gt;

&lt;p&gt;The domain is the starting point.&lt;/p&gt;

&lt;p&gt;The actual attack surface can look more like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;company.com
├── www.company.com
├── app.company.com
├── api.company.com
├── api-v2.company.com
├── staging.company.com
├── dev.company.com
├── vpn.company.com
├── sso.company.com
├── git.company.com
├── ci.company.com
├── monitoring.company.com
└── legacy.company.com

Cloud
├── public S3 buckets
├── public load balancers
├── exposed databases
├── Kubernetes APIs
└── forgotten public IPs
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The first objective is therefore:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Build an inventory of externally reachable assets before attempting to exploit them.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This is where passive OSINT and attack-surface discovery become more important than simply running a vulnerability scanner.&lt;/p&gt;




&lt;h1&gt;
  
  
  2. Phase One: Passive Reconnaissance
&lt;/h1&gt;

&lt;p&gt;Passive reconnaissance gathers information without directly probing the target infrastructure.&lt;/p&gt;

&lt;p&gt;The goal is to answer:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;What domains belong to the organization?&lt;/li&gt;
&lt;li&gt;What IP ranges are associated with it?&lt;/li&gt;
&lt;li&gt;What subdomains exist?&lt;/li&gt;
&lt;li&gt;What technologies are publicly visible?&lt;/li&gt;
&lt;li&gt;What third-party services are being used?&lt;/li&gt;
&lt;li&gt;What infrastructure has historically existed?&lt;/li&gt;
&lt;li&gt;Have credentials or secrets appeared in public repositories?&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  2.1 DNS Enumeration
&lt;/h2&gt;

&lt;p&gt;Start with the basic DNS records.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;dig +short company.com A
dig +short company.com AAAA
dig +short company.com MX
dig +short company.com NS
dig +short company.com TXT
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each record can provide a different piece of infrastructure intelligence.&lt;/p&gt;

&lt;h3&gt;
  
  
  A / AAAA
&lt;/h3&gt;

&lt;p&gt;These identify IPv4 and IPv6 addresses associated with the domain.&lt;/p&gt;

&lt;h3&gt;
  
  
  MX
&lt;/h3&gt;

&lt;p&gt;Mail records can reveal:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Microsoft 365&lt;/li&gt;
&lt;li&gt;Google Workspace&lt;/li&gt;
&lt;li&gt;third-party email providers&lt;/li&gt;
&lt;li&gt;dedicated mail infrastructure&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  TXT
&lt;/h3&gt;

&lt;p&gt;TXT records commonly contain:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;SPF configuration&lt;/li&gt;
&lt;li&gt;domain verification records&lt;/li&gt;
&lt;li&gt;SaaS integrations&lt;/li&gt;
&lt;li&gt;cloud-provider verification&lt;/li&gt;
&lt;li&gt;email security configuration&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  NS
&lt;/h3&gt;

&lt;p&gt;Nameservers identify the DNS infrastructure responsible for the domain.&lt;/p&gt;




&lt;h1&gt;
  
  
  3. Test for DNS Zone Transfers
&lt;/h1&gt;

&lt;p&gt;A misconfigured authoritative DNS server may allow an AXFR request.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;dig axfr company.com @ns1.company.com
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A successful response can disclose the entire DNS zone.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;dev.company.com
staging.company.com
internal.company.com
db01.company.com
vpn.company.com
old-api.company.com
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A failed AXFR is expected.&lt;/p&gt;

&lt;p&gt;A successful transfer is materially different because it can expose internal naming conventions and infrastructure that may not appear through ordinary enumeration.&lt;/p&gt;




&lt;h1&gt;
  
  
  4. Certificate Transparency
&lt;/h1&gt;

&lt;p&gt;Certificate Transparency logs are one of the most useful passive sources for discovering forgotten hostnames.&lt;/p&gt;

&lt;p&gt;A simple &lt;code&gt;crt.sh&lt;/code&gt; query:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-s&lt;/span&gt; &lt;span class="s2"&gt;"https://crt.sh/?q=%25.company.com&amp;amp;output=json"&lt;/span&gt; |
    jq &lt;span class="nt"&gt;-r&lt;/span&gt; &lt;span class="s1"&gt;'.[].name_value'&lt;/span&gt; |
    &lt;span class="nb"&gt;sort&lt;/span&gt; &lt;span class="nt"&gt;-u&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Normalize the results before continuing.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-s&lt;/span&gt; &lt;span class="s2"&gt;"https://crt.sh/?q=%25.company.com&amp;amp;output=json"&lt;/span&gt; |
    jq &lt;span class="nt"&gt;-r&lt;/span&gt; &lt;span class="s1"&gt;'.[].name_value'&lt;/span&gt; |
    &lt;span class="nb"&gt;sed&lt;/span&gt; &lt;span class="s1"&gt;'s/\*\.//g'&lt;/span&gt; |
    &lt;span class="nb"&gt;sort&lt;/span&gt; &lt;span class="nt"&gt;-u&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Certificate data can reveal hosts that aren't linked from the organization's main website.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;app.company.com
api.company.com
staging.company.com
qa.company.com
legacy.company.com
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important point is that &lt;strong&gt;discovery is not vulnerability confirmation&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;staging.company.com&lt;/code&gt; isn't a finding simply because it exists.&lt;/p&gt;

&lt;p&gt;It is an asset that now needs testing.&lt;/p&gt;




&lt;h1&gt;
  
  
  5. ASN and IP Range Discovery
&lt;/h1&gt;

&lt;p&gt;If the engagement includes network infrastructure, determine which public ranges are associated with the organization.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;whois &lt;span class="nt"&gt;-h&lt;/span&gt; whois.radb.net &lt;span class="nt"&gt;--&lt;/span&gt; &lt;span class="s1"&gt;'-i origin AS12345'&lt;/span&gt; |
    &lt;span class="nb"&gt;grep &lt;/span&gt;route:
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;ASN information can reveal infrastructure that isn't directly associated with the primary domain.&lt;/p&gt;

&lt;p&gt;You can then use the resulting ranges as input to the active reconnaissance phase.&lt;/p&gt;

&lt;p&gt;This is particularly useful for organizations operating their own infrastructure rather than relying entirely on cloud providers.&lt;/p&gt;




&lt;h1&gt;
  
  
  6. Historical URL Discovery
&lt;/h1&gt;

&lt;p&gt;Current crawling only tells you what exists now.&lt;/p&gt;

&lt;p&gt;Historical sources can reveal endpoints that existed previously.&lt;/p&gt;

&lt;p&gt;Two common tools are:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;gau company.com
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;waybackurls company.com
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then filter for interesting paths:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;gau company.com |
    &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-Ei&lt;/span&gt; &lt;span class="s1"&gt;'\.(json|xml|config|bak|sql|zip|env)$'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Or search for administrative functionality:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;gau company.com |
    &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-Ei&lt;/span&gt; &lt;span class="s1"&gt;'(admin|debug|internal|api|graphql|swagger)'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Historical URLs are particularly useful because applications evolve.&lt;/p&gt;

&lt;p&gt;An endpoint may disappear from the current interface while remaining deployed.&lt;/p&gt;




&lt;h1&gt;
  
  
  7. Public Repository Reconnaissance
&lt;/h1&gt;

&lt;p&gt;Public source repositories can expose much more than source code.&lt;/p&gt;

&lt;p&gt;Search for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;API keys&lt;/li&gt;
&lt;li&gt;cloud credentials&lt;/li&gt;
&lt;li&gt;internal hostnames&lt;/li&gt;
&lt;li&gt;database URLs&lt;/li&gt;
&lt;li&gt;deployment scripts&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;.env&lt;/code&gt; files&lt;/li&gt;
&lt;li&gt;private package registries&lt;/li&gt;
&lt;li&gt;CI/CD configuration&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;"company.com" API_KEY
"company.com" AWS_SECRET
"company-internal"
"database.company.com"
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Secret scanning tools can help automate this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;trufflehog github &lt;span class="nt"&gt;--org&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;company
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A discovered credential still needs validation.&lt;/p&gt;

&lt;p&gt;Don't assume that every string matching &lt;code&gt;AWS_SECRET&lt;/code&gt; is an active credential.&lt;/p&gt;

&lt;p&gt;The tester should establish:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Is it syntactically valid?&lt;/li&gt;
&lt;li&gt;Is it still active?&lt;/li&gt;
&lt;li&gt;What identity does it belong to?&lt;/li&gt;
&lt;li&gt;What permissions does it have?&lt;/li&gt;
&lt;li&gt;Does it provide access to in-scope resources?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That turns a potential secret leak into a measurable security finding.&lt;/p&gt;




&lt;h1&gt;
  
  
  8. Build the Initial Asset Inventory
&lt;/h1&gt;

&lt;p&gt;At this stage, consolidate the passive results.&lt;/p&gt;

&lt;p&gt;A useful working inventory might look like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;hostname                  IP              source
----------------------------------------------------------
www.company.com           203.0.113.10    DNS
api.company.com           203.0.113.20    CT
staging.company.com       203.0.113.30    CT
vpn.company.com           203.0.113.40    DNS
legacy.company.com        203.0.113.50    Wayback
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Don't immediately start exploiting.&lt;/p&gt;

&lt;p&gt;First remove:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;duplicates&lt;/li&gt;
&lt;li&gt;third-party assets outside scope&lt;/li&gt;
&lt;li&gt;CDN infrastructure&lt;/li&gt;
&lt;li&gt;unrelated shared hosting&lt;/li&gt;
&lt;li&gt;dead DNS records&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Then resolve the remaining hosts.&lt;/p&gt;




&lt;h1&gt;
  
  
  9. Phase Two: Active Reconnaissance
&lt;/h1&gt;

&lt;p&gt;Now the tester starts interacting directly with the target.&lt;/p&gt;

&lt;p&gt;The objective is to determine:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Which hosts are alive?&lt;/li&gt;
&lt;li&gt;Which ports are open?&lt;/li&gt;
&lt;li&gt;What services are running?&lt;/li&gt;
&lt;li&gt;Which applications are exposed?&lt;/li&gt;
&lt;li&gt;Which technologies are being used?&lt;/li&gt;
&lt;li&gt;Which endpoints exist?&lt;/li&gt;
&lt;/ul&gt;




&lt;h1&gt;
  
  
  10. Subdomain Enumeration
&lt;/h1&gt;

&lt;p&gt;Use multiple discovery sources where appropriate.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;subfinder &lt;span class="nt"&gt;-d&lt;/span&gt; company.com &lt;span class="nt"&gt;-silent&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; subdomains.txt
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Amass can provide additional enumeration:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;amass enum &lt;span class="nt"&gt;-passive&lt;/span&gt; &lt;span class="nt"&gt;-d&lt;/span&gt; company.com &lt;span class="nt"&gt;-o&lt;/span&gt; amass.txt
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Merge the results:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;cat &lt;/span&gt;subdomains.txt amass.txt |
    &lt;span class="nb"&gt;sort&lt;/span&gt; &lt;span class="nt"&gt;-u&lt;/span&gt; |
    &lt;span class="nb"&gt;tee &lt;/span&gt;all_subdomains.txt
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then resolve and probe them:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;cat &lt;/span&gt;all_subdomains.txt |
    httpx &lt;span class="nt"&gt;-silent&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;-status-code&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;-title&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;-tech-detect&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;-follow-redirects&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;-o&lt;/span&gt; live_hosts.txt
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now the tester has something much more useful than a raw subdomain list:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;https://app.company.com       200   Customer Portal
https://api.company.com       200   API
https://staging.company.com   200   Staging
https://vpn.company.com       302   VPN Login
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h1&gt;
  
  
  11. Full Port Enumeration
&lt;/h1&gt;

&lt;p&gt;For identified IP addresses:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;nmap &lt;span class="nt"&gt;-p-&lt;/span&gt; &lt;span class="nt"&gt;--open&lt;/span&gt; &lt;span class="nt"&gt;-sV&lt;/span&gt; &lt;span class="nt"&gt;-sC&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;-oA&lt;/span&gt; nmap_full_scan &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="o"&gt;[&lt;/span&gt;TARGET_IP]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For large ranges, scanners such as RustScan can accelerate initial port discovery before handing results to Nmap for service detection.&lt;/p&gt;

&lt;p&gt;The important distinction is:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Port discovery tells you where something is listening.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Service enumeration tells you what it is.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;22/tcp    SSH
80/tcp    HTTP
443/tcp   HTTPS
3306/tcp  MySQL
6379/tcp  Redis
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each one becomes a separate testing path.&lt;/p&gt;




&lt;h1&gt;
  
  
  12. Prioritize the Interesting Ports
&lt;/h1&gt;

&lt;p&gt;Not every open port deserves equal attention.&lt;/p&gt;

&lt;p&gt;High-interest services include:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;21      FTP
22      SSH
23      Telnet
25      SMTP
53      DNS
80      HTTP
443     HTTPS
445     SMB
1433    MSSQL
3306    MySQL
3389    RDP
5432    PostgreSQL
5900    VNC
6379    Redis
9200    Elasticsearch
2375    Docker API
6443    Kubernetes API
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;But don't turn this list into a severity checklist.&lt;/p&gt;

&lt;p&gt;An exposed port is an &lt;strong&gt;observation&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The actual finding depends on what is accessible through it.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;5432/tcp open PostgreSQL
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;is not equivalent to:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;5432/tcp open PostgreSQL
Authentication disabled
Unauthenticated database access confirmed
Production data accessible
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The second is a security finding.&lt;/p&gt;




&lt;h1&gt;
  
  
  13. Web Technology Fingerprinting
&lt;/h1&gt;

&lt;p&gt;For HTTP services:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;httpx &lt;span class="nt"&gt;-u&lt;/span&gt; https://company.com &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;-status-code&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;-title&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;-tech-detect&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;-server&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;-content-length&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Look for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;framework fingerprints&lt;/li&gt;
&lt;li&gt;server versions&lt;/li&gt;
&lt;li&gt;exposed headers&lt;/li&gt;
&lt;li&gt;reverse proxies&lt;/li&gt;
&lt;li&gt;CDNs&lt;/li&gt;
&lt;li&gt;application frameworks&lt;/li&gt;
&lt;li&gt;CMS platforms&lt;/li&gt;
&lt;li&gt;API technologies&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Technology identification can influence the next stage of testing.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;nginx
PHP
Laravel
GraphQL
WordPress
Spring Boot
ASP.NET
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;each suggests a different set of likely endpoints and vulnerabilities.&lt;/p&gt;




&lt;h1&gt;
  
  
  14. Directory and Endpoint Discovery
&lt;/h1&gt;

&lt;p&gt;Once an application is identified, enumerate paths.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ffuf &lt;span class="nt"&gt;-u&lt;/span&gt; https://company.com/FUZZ &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;-w&lt;/span&gt; /usr/share/wordlists/dirb/common.txt &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;-mc&lt;/span&gt; 200,204,301,302,307,401,403
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For extensions:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ffuf &lt;span class="nt"&gt;-u&lt;/span&gt; https://company.com/FUZZ &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;-w&lt;/span&gt; wordlist.txt &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;-e&lt;/span&gt; .php,.json,.xml,.txt,.bak,.zip
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Interesting responses include:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;200 OK
401 Unauthorized
403 Forbidden
500 Internal Server Error
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A &lt;code&gt;403&lt;/code&gt; can be interesting because it proves the endpoint exists.&lt;/p&gt;

&lt;p&gt;A &lt;code&gt;500&lt;/code&gt; can also be useful because error handling sometimes exposes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;stack traces&lt;/li&gt;
&lt;li&gt;framework versions&lt;/li&gt;
&lt;li&gt;database errors&lt;/li&gt;
&lt;li&gt;internal paths&lt;/li&gt;
&lt;li&gt;debug information&lt;/li&gt;
&lt;/ul&gt;




&lt;h1&gt;
  
  
  15. API Enumeration
&lt;/h1&gt;

&lt;p&gt;Modern external attack surfaces are increasingly API-heavy.&lt;/p&gt;

&lt;p&gt;Start by identifying:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;/api
/api/v1
/api/v2
/graphql
/swagger
/openapi.json
/api-docs
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Search historical URLs:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;gau company.com |
    &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-Ei&lt;/span&gt; &lt;span class="s1"&gt;'/api/|graphql|swagger|openapi'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If an OpenAPI specification is publicly accessible:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-s&lt;/span&gt; https://api.company.com/openapi.json
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You may immediately obtain:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;endpoint names&lt;/li&gt;
&lt;li&gt;parameters&lt;/li&gt;
&lt;li&gt;object identifiers&lt;/li&gt;
&lt;li&gt;authentication schemes&lt;/li&gt;
&lt;li&gt;request methods&lt;/li&gt;
&lt;li&gt;API versions&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This can dramatically improve manual testing.&lt;/p&gt;




&lt;h1&gt;
  
  
  16. Authentication Testing
&lt;/h1&gt;

&lt;p&gt;External authentication testing should go beyond:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“Does the login page exist?”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Test the complete authentication boundary.&lt;/p&gt;

&lt;p&gt;Questions include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Is MFA enforced?&lt;/li&gt;
&lt;li&gt;Can MFA be bypassed?&lt;/li&gt;
&lt;li&gt;Are password-reset tokens predictable?&lt;/li&gt;
&lt;li&gt;Do reset links expire?&lt;/li&gt;
&lt;li&gt;Can sessions be reused?&lt;/li&gt;
&lt;li&gt;Are tokens invalidated after logout?&lt;/li&gt;
&lt;li&gt;Are API tokens scoped?&lt;/li&gt;
&lt;li&gt;Are legacy authentication endpoints still active?&lt;/li&gt;
&lt;li&gt;Do mobile and web authentication paths behave differently?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A common mistake is testing only the primary web login.&lt;/p&gt;

&lt;p&gt;The real authentication surface may include:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;/web/login
/api/login
/mobile/auth
/oauth/token
/sso/login
/password/reset
/admin/login
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h1&gt;
  
  
  17. Authorization Testing
&lt;/h1&gt;

&lt;p&gt;Authorization flaws are often more valuable than generic version-based findings because they depend on application logic.&lt;/p&gt;

&lt;p&gt;Consider an API:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight http"&gt;&lt;code&gt;&lt;span class="err"&gt;GET /api/users/1001
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Change the identifier:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight http"&gt;&lt;code&gt;&lt;span class="err"&gt;GET /api/users/1002
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If user 1001 can access user 1002's information, the issue is an Insecure Direct Object Reference or broader broken object-level authorization problem.&lt;/p&gt;

&lt;p&gt;The same principle applies to:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;/orders/1001
/invoices/1001
/projects/1001
/files/1001
/payments/1001
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The tester should ask:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Does the server verify that this authenticated principal is actually authorized to access this object?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Changing an ID and receiving a &lt;code&gt;200&lt;/code&gt; response isn't enough by itself.&lt;/p&gt;

&lt;p&gt;The response needs to demonstrate unauthorized access.&lt;/p&gt;




&lt;h1&gt;
  
  
  18. Test Alternate API Versions
&lt;/h1&gt;

&lt;p&gt;One of the easiest blind spots is assuming that:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;/api/v2
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;has the same security controls as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;/api/v1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Test both.&lt;/p&gt;

&lt;p&gt;Look for differences in:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;authentication&lt;/li&gt;
&lt;li&gt;authorization&lt;/li&gt;
&lt;li&gt;input validation&lt;/li&gt;
&lt;li&gt;rate limiting&lt;/li&gt;
&lt;li&gt;object-level access controls&lt;/li&gt;
&lt;li&gt;error handling&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;An old API doesn't need to be linked from the frontend to remain exploitable.&lt;/p&gt;

&lt;p&gt;If it is internet-accessible, it is part of the attack surface.&lt;/p&gt;




&lt;h1&gt;
  
  
  19. Vulnerability Scanning
&lt;/h1&gt;

&lt;p&gt;Once the attack surface is mapped, automated vulnerability detection becomes much more useful.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;nuclei &lt;span class="nt"&gt;-l&lt;/span&gt; live_hosts.txt &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;-severity&lt;/span&gt; critical,high,medium &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;-tags&lt;/span&gt; cve,exposure,misconfig &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;-o&lt;/span&gt; nuclei.txt
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Run scanners against known technology where possible rather than blindly scanning everything.&lt;/p&gt;

&lt;p&gt;The output should become a triage queue.&lt;/p&gt;

&lt;p&gt;Not the final report.&lt;/p&gt;




&lt;h1&gt;
  
  
  20. Scanner Finding vs Real Finding
&lt;/h1&gt;

&lt;p&gt;Consider:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Nuclei:
Apache CVE detected
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's a lead.&lt;/p&gt;

&lt;p&gt;The validation workflow should be:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Scanner match
     ↓
Identify exact version
     ↓
Confirm affected component
     ↓
Determine whether vulnerable feature is reachable
     ↓
Reproduce safely
     ↓
Establish impact
     ↓
Collect evidence
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Only then should it become a confirmed finding.&lt;/p&gt;

&lt;p&gt;This matters because scanners can produce false positives from:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;backported patches&lt;/li&gt;
&lt;li&gt;incorrect version detection&lt;/li&gt;
&lt;li&gt;reverse proxies&lt;/li&gt;
&lt;li&gt;custom builds&lt;/li&gt;
&lt;li&gt;disabled vulnerable modules&lt;/li&gt;
&lt;li&gt;unreachable vulnerable functionality&lt;/li&gt;
&lt;/ul&gt;




&lt;h1&gt;
  
  
  21. Validate Exposed Redis
&lt;/h1&gt;

&lt;p&gt;If Redis is exposed:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;redis-cli &lt;span class="nt"&gt;-h&lt;/span&gt; &lt;span class="o"&gt;[&lt;/span&gt;TARGET_IP] ping
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A response of:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;PONG
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;shows connectivity.&lt;/p&gt;

&lt;p&gt;It doesn't automatically prove unauthenticated administrative access.&lt;/p&gt;

&lt;p&gt;The next question is whether commands requiring authentication are permitted.&lt;/p&gt;

&lt;p&gt;For example, safely establish the access level permitted under the engagement rules.&lt;/p&gt;

&lt;p&gt;The finding should ultimately describe &lt;strong&gt;what access was demonstrated&lt;/strong&gt;, not merely that port 6379 was open.&lt;/p&gt;




&lt;h1&gt;
  
  
  22. Validate Database Exposure
&lt;/h1&gt;

&lt;p&gt;For a PostgreSQL service:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;psql &lt;span class="nt"&gt;-h&lt;/span&gt; &lt;span class="o"&gt;[&lt;/span&gt;TARGET_IP] &lt;span class="nt"&gt;-U&lt;/span&gt; &lt;span class="o"&gt;[&lt;/span&gt;TEST_USER] &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="o"&gt;[&lt;/span&gt;DATABASE]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For MySQL:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;mysql &lt;span class="nt"&gt;-h&lt;/span&gt; &lt;span class="o"&gt;[&lt;/span&gt;TARGET_IP] &lt;span class="nt"&gt;-u&lt;/span&gt; &lt;span class="o"&gt;[&lt;/span&gt;TEST_USER] &lt;span class="nt"&gt;-p&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The purpose is to establish whether:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;authentication is required&lt;/li&gt;
&lt;li&gt;credentials work&lt;/li&gt;
&lt;li&gt;access is restricted&lt;/li&gt;
&lt;li&gt;the exposed account has excessive privileges&lt;/li&gt;
&lt;li&gt;sensitive data is accessible&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Never dump unnecessary production data simply to prove access.&lt;/p&gt;

&lt;p&gt;A small, controlled proof is generally enough.&lt;/p&gt;




&lt;h1&gt;
  
  
  23. Exposed Management Interfaces
&lt;/h1&gt;

&lt;p&gt;Management systems deserve special attention because they often have significantly more privileges than ordinary applications.&lt;/p&gt;

&lt;p&gt;Common examples:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Jenkins
Grafana
Kibana
Argo CD
Rancher
GitLab
Prometheus
Docker
Kubernetes
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The tester should determine:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Is the interface publicly accessible?&lt;/li&gt;
&lt;li&gt;Is authentication required?&lt;/li&gt;
&lt;li&gt;What identity is created after authentication?&lt;/li&gt;
&lt;li&gt;What actions can that identity perform?&lt;/li&gt;
&lt;li&gt;Can the interface reach internal infrastructure?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;A publicly reachable Jenkins login page isn't automatically a vulnerability.&lt;/p&gt;

&lt;p&gt;An unauthenticated Jenkins instance allowing job execution is an entirely different problem.&lt;/p&gt;




&lt;h1&gt;
  
  
  24. Cloud Attack Surface
&lt;/h1&gt;

&lt;p&gt;External testing increasingly means testing cloud infrastructure rather than traditional perimeter devices.&lt;/p&gt;

&lt;p&gt;Common areas include:&lt;/p&gt;

&lt;h3&gt;
  
  
  Object storage
&lt;/h3&gt;

&lt;p&gt;Check for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;public listing&lt;/li&gt;
&lt;li&gt;public reads&lt;/li&gt;
&lt;li&gt;public writes&lt;/li&gt;
&lt;li&gt;unintended object exposure&lt;/li&gt;
&lt;li&gt;backup files&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Public load balancers
&lt;/h3&gt;

&lt;p&gt;Identify:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;backend services&lt;/li&gt;
&lt;li&gt;alternate listeners&lt;/li&gt;
&lt;li&gt;forgotten ports&lt;/li&gt;
&lt;li&gt;administrative endpoints&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Cloud-hosted databases
&lt;/h3&gt;

&lt;p&gt;Look for publicly reachable:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;RDS
Cloud SQL
MongoDB
Redis
Elasticsearch
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Kubernetes
&lt;/h3&gt;

&lt;p&gt;Look for exposed:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Kubernetes API
Dashboard
Ingress controllers
Metrics
Management interfaces
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Cloud security failures often come from resources that were created temporarily and never removed.&lt;/p&gt;




&lt;h1&gt;
  
  
  25. Public S3 Bucket Discovery
&lt;/h1&gt;

&lt;p&gt;Bucket discovery can begin with known naming patterns, but DNS and application source code can provide better candidates.&lt;/p&gt;

&lt;p&gt;For a bucket you are authorized to test:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;aws s3 &lt;span class="nb"&gt;ls &lt;/span&gt;s3://bucket-name &lt;span class="nt"&gt;--no-sign-request&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A successful listing demonstrates public listing access.&lt;/p&gt;

&lt;p&gt;Then determine whether objects are publicly readable.&lt;/p&gt;

&lt;p&gt;Do not assume that because a bucket exists, all of its contents are public.&lt;/p&gt;

&lt;p&gt;Test the actual permission boundary.&lt;/p&gt;

&lt;p&gt;The distinction between:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;bucket exists
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;anonymous principal can list and download objects
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;is the difference between asset discovery and a confirmed exposure.&lt;/p&gt;




&lt;h1&gt;
  
  
  26. Exposed &lt;code&gt;.git&lt;/code&gt; Directories
&lt;/h1&gt;

&lt;p&gt;A public &lt;code&gt;.git&lt;/code&gt; directory can potentially expose repository history.&lt;/p&gt;

&lt;p&gt;Check:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-I&lt;/span&gt; https://company.com/.git/HEAD
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If accessible:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl https://company.com/.git/HEAD
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Depending on what is exposed, an attacker may be able to reconstruct repository information.&lt;/p&gt;

&lt;p&gt;The security impact depends on what the repository contains.&lt;/p&gt;

&lt;p&gt;Potentially sensitive material includes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;source code&lt;/li&gt;
&lt;li&gt;credentials&lt;/li&gt;
&lt;li&gt;deployment configuration&lt;/li&gt;
&lt;li&gt;internal endpoints&lt;/li&gt;
&lt;li&gt;historical secrets&lt;/li&gt;
&lt;/ul&gt;




&lt;h1&gt;
  
  
  27. Exposed &lt;code&gt;.env&lt;/code&gt; Files
&lt;/h1&gt;

&lt;p&gt;A simple check:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-i&lt;/span&gt; https://company.com/.env
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A response containing environment configuration can be severe if it exposes active credentials.&lt;/p&gt;

&lt;p&gt;Potential entries include:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;DB_HOST=
DB_USERNAME=
DB_PASSWORD=
AWS_ACCESS_KEY_ID=
AWS_SECRET_ACCESS_KEY=
STRIPE_SECRET_KEY=
JWT_SECRET=
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Again, the key question is whether the values are actually active and what access they provide.&lt;/p&gt;

&lt;p&gt;A leaked secret should be validated carefully and reported with enough evidence to demonstrate impact without unnecessarily exposing the secret itself.&lt;/p&gt;




&lt;h1&gt;
  
  
  28. Subdomain Takeover Testing
&lt;/h1&gt;

&lt;p&gt;A typical workflow:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;subdomain discovered
        ↓
DNS record inspected
        ↓
CNAME points to third-party service
        ↓
resource no longer exists
        ↓
service confirms resource can potentially be claimed
        ↓
controlled validation
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A dangling CNAME alone is not sufficient evidence.&lt;/p&gt;

&lt;p&gt;The tester needs to establish whether the third-party resource is actually claimable.&lt;/p&gt;




&lt;h1&gt;
  
  
  29. The Findings Automation Commonly Misses
&lt;/h1&gt;

&lt;p&gt;The most interesting external pentest findings aren't always CVEs.&lt;/p&gt;

&lt;p&gt;They are often inconsistencies.&lt;/p&gt;

&lt;h3&gt;
  
  
  Different authorization between UI and API
&lt;/h3&gt;

&lt;p&gt;The UI checks permission A.&lt;/p&gt;

&lt;p&gt;The API checks permission B.&lt;/p&gt;

&lt;h3&gt;
  
  
  Production and staging have different security controls
&lt;/h3&gt;

&lt;p&gt;Production requires MFA.&lt;/p&gt;

&lt;p&gt;Staging doesn't.&lt;/p&gt;

&lt;h3&gt;
  
  
  API v1 and v2 implement authorization differently
&lt;/h3&gt;

&lt;p&gt;One version checks object ownership.&lt;/p&gt;

&lt;p&gt;The other trusts the object ID supplied by the client.&lt;/p&gt;

&lt;h3&gt;
  
  
  Legacy endpoints remain accessible
&lt;/h3&gt;

&lt;p&gt;The frontend stopped using &lt;code&gt;/api/v1&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The server never stopped serving it.&lt;/p&gt;

&lt;h3&gt;
  
  
  Forgotten infrastructure has stronger connectivity than expected
&lt;/h3&gt;

&lt;p&gt;A staging server is internet-facing but also has access to internal production services.&lt;/p&gt;

&lt;p&gt;These findings require understanding &lt;strong&gt;relationships&lt;/strong&gt;, not just signatures.&lt;/p&gt;

&lt;p&gt;That's why manual testing remains essential even when automated scanning is extensive.&lt;/p&gt;




&lt;h1&gt;
  
  
  30. Attack-Path Thinking
&lt;/h1&gt;

&lt;p&gt;The strongest pentests don't treat findings as isolated rows in a spreadsheet.&lt;/p&gt;

&lt;p&gt;They ask whether findings can be chained.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Forgotten subdomain
        ↓
Staging application
        ↓
Debug endpoint
        ↓
Cloud credential exposure
        ↓
Overprivileged IAM identity
        ↓
Production storage access
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;None of the individual observations necessarily represents the complete impact.&lt;/p&gt;

&lt;p&gt;The chain does.&lt;/p&gt;

&lt;p&gt;Another example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Exposed management interface
        ↓
Weak authentication
        ↓
Administrative access
        ↓
CI/CD job execution
        ↓
Cloud credentials
        ↓
Production environment
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is why external penetration testing is different from simply producing a vulnerability scan.&lt;/p&gt;

&lt;p&gt;The tester is trying to understand &lt;strong&gt;what an attacker can do with the access they obtain&lt;/strong&gt;.&lt;/p&gt;




&lt;h1&gt;
  
  
  31. External Pentesting Tool Stack
&lt;/h1&gt;

&lt;p&gt;A practical toolchain might look like this:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Phase&lt;/th&gt;
&lt;th&gt;Tools&lt;/th&gt;
&lt;th&gt;Purpose&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;DNS&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;dig&lt;/code&gt;, &lt;code&gt;dnsx&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;DNS enumeration&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Subdomains&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;subfinder&lt;/code&gt;, &lt;code&gt;Amass&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Host discovery&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Certificate data&lt;/td&gt;
&lt;td&gt;&lt;code&gt;crt.sh&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Passive hostname discovery&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;HTTP probing&lt;/td&gt;
&lt;td&gt;&lt;code&gt;httpx&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Live host identification&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Port scanning&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;Nmap&lt;/code&gt;, RustScan&lt;/td&gt;
&lt;td&gt;Port and service discovery&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Web discovery&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;ffuf&lt;/code&gt;, Gobuster&lt;/td&gt;
&lt;td&gt;Endpoint enumeration&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Historical discovery&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;gau&lt;/code&gt;, &lt;code&gt;waybackurls&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Old endpoints&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OSINT&lt;/td&gt;
&lt;td&gt;Shodan, Censys, theHarvester&lt;/td&gt;
&lt;td&gt;Public infrastructure&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Vulnerability scanning&lt;/td&gt;
&lt;td&gt;Nuclei&lt;/td&gt;
&lt;td&gt;CVEs and misconfigurations&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Secret detection&lt;/td&gt;
&lt;td&gt;TruffleHog&lt;/td&gt;
&lt;td&gt;Public credential discovery&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Web testing&lt;/td&gt;
&lt;td&gt;Burp Suite&lt;/td&gt;
&lt;td&gt;Manual HTTP testing&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Exploitation&lt;/td&gt;
&lt;td&gt;Metasploit&lt;/td&gt;
&lt;td&gt;Controlled exploitation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cloud&lt;/td&gt;
&lt;td&gt;AWS/Azure/GCP CLI tools&lt;/td&gt;
&lt;td&gt;Cloud validation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SMB/Windows&lt;/td&gt;
&lt;td&gt;enum4linux&lt;/td&gt;
&lt;td&gt;SMB enumeration&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The exact stack should change according to the target.&lt;/p&gt;

&lt;p&gt;A SaaS application doesn't need the same workflow as an enterprise network containing VPN concentrators, Windows infrastructure, and exposed management systems.&lt;/p&gt;




&lt;h1&gt;
  
  
  32. External vs Internal Penetration Testing
&lt;/h1&gt;

&lt;p&gt;These engagements begin from fundamentally different trust assumptions.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;External&lt;/th&gt;
&lt;th&gt;Internal&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Starting point&lt;/td&gt;
&lt;td&gt;Internet&lt;/td&gt;
&lt;td&gt;Internal network&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Credentials&lt;/td&gt;
&lt;td&gt;Usually none&lt;/td&gt;
&lt;td&gt;Usually provided or obtained&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Recon&lt;/td&gt;
&lt;td&gt;OSINT, DNS, public infrastructure&lt;/td&gt;
&lt;td&gt;Network and directory enumeration&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Primary target&lt;/td&gt;
&lt;td&gt;Perimeter&lt;/td&gt;
&lt;td&gt;Internal systems&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Typical findings&lt;/td&gt;
&lt;td&gt;Exposed services, web/API flaws&lt;/td&gt;
&lt;td&gt;AD, privilege escalation, lateral movement&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Main question&lt;/td&gt;
&lt;td&gt;Can an attacker get in?&lt;/td&gt;
&lt;td&gt;What can they do after getting in?&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;An external pentest shouldn't be treated as a cheaper version of an internal pentest.&lt;/p&gt;

&lt;p&gt;They model different attack positions.&lt;/p&gt;




&lt;h1&gt;
  
  
  33. Reporting a Finding Properly
&lt;/h1&gt;

&lt;p&gt;A useful finding should answer five questions:&lt;/p&gt;

&lt;h3&gt;
  
  
  What is vulnerable?
&lt;/h3&gt;

&lt;p&gt;Identify the exact asset and component.&lt;/p&gt;

&lt;h3&gt;
  
  
  Where is it?
&lt;/h3&gt;

&lt;p&gt;Provide the hostname, endpoint, port, or resource.&lt;/p&gt;

&lt;h3&gt;
  
  
  How was it validated?
&lt;/h3&gt;

&lt;p&gt;Give a reproducible proof.&lt;/p&gt;

&lt;h3&gt;
  
  
  What can an attacker do?
&lt;/h3&gt;

&lt;p&gt;Explain the actual impact.&lt;/p&gt;

&lt;h3&gt;
  
  
  How should it be fixed?
&lt;/h3&gt;

&lt;p&gt;Provide a practical remediation.&lt;/p&gt;

&lt;p&gt;For example, instead of:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Redis exposed to internet — Critical&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;write:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Unauthenticated Redis instance accessible from the public internet&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;code&gt;redis.company.com:6379&lt;/code&gt; accepts unauthenticated commands from an external network. During validation, the tester was able to enumerate the accessible Redis environment without credentials. Network exposure should be restricted to trusted application hosts and authentication should be enforced.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is substantially more useful.&lt;/p&gt;




&lt;h1&gt;
  
  
  34. Evidence Matters
&lt;/h1&gt;

&lt;p&gt;Good evidence should allow another engineer to reproduce the finding.&lt;/p&gt;

&lt;p&gt;Useful evidence includes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Request
Response
Hostname
Port
Timestamp
Authenticated identity
Relevant configuration
Minimal proof of impact
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Avoid collecting unnecessary sensitive data.&lt;/p&gt;

&lt;p&gt;For example, if a database contains 10 million customer records, you don't need to export all 10 million records to prove unauthorized database access.&lt;/p&gt;

&lt;p&gt;A single authorized test record or metadata response may be enough.&lt;/p&gt;




&lt;h1&gt;
  
  
  35. A Practical External Pentest Workflow
&lt;/h1&gt;

&lt;p&gt;Putting everything together:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                 TARGET
                    │
                    ▼
             Passive OSINT
                    │
          ┌─────────┼─────────┐
          ▼         ▼         ▼
         DNS       CT       ASN/IP
          │         │         │
          └─────────┼─────────┘
                    ▼
             Asset Inventory
                    │
                    ▼
          Active Reconnaissance
                    │
          ┌─────────┼─────────┐
          ▼         ▼         ▼
       HTTP       Ports      APIs
          │         │         │
          └─────────┼─────────┘
                    ▼
          Vulnerability Testing
                    │
                    ▼
             Manual Validation
                    │
                    ▼
          Controlled Exploitation
                    │
                    ▼
             Attack-Path Analysis
                    │
                    ▼
                Reporting
                    │
                    ▼
                Retesting
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important part is that each phase informs the next.&lt;/p&gt;

&lt;p&gt;Certificate Transparency discovers a hostname.&lt;/p&gt;

&lt;p&gt;The hostname resolves to an IP.&lt;/p&gt;

&lt;p&gt;The IP exposes a service.&lt;/p&gt;

&lt;p&gt;The service identifies an application.&lt;/p&gt;

&lt;p&gt;The application exposes an API.&lt;/p&gt;

&lt;p&gt;The API contains an authorization flaw.&lt;/p&gt;

&lt;p&gt;The authorization flaw provides access to another object.&lt;/p&gt;

&lt;p&gt;That is the attack path.&lt;/p&gt;




&lt;h1&gt;
  
  
  36. What a Good External Pentest Actually Proves
&lt;/h1&gt;

&lt;p&gt;At the end of the engagement, the important question isn't:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;How many vulnerabilities did we find?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It's:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What could an external attacker actually accomplish?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A good external penetration test should establish:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;What infrastructure is publicly reachable&lt;/li&gt;
&lt;li&gt;What applications are exposed&lt;/li&gt;
&lt;li&gt;Which services are vulnerable&lt;/li&gt;
&lt;li&gt;Which authentication controls can be bypassed&lt;/li&gt;
&lt;li&gt;Which authorization boundaries fail&lt;/li&gt;
&lt;li&gt;Which sensitive resources are accessible&lt;/li&gt;
&lt;li&gt;Which findings can be chained&lt;/li&gt;
&lt;li&gt;What level of access an attacker can obtain&lt;/li&gt;
&lt;li&gt;What remediation actually closes the attack path&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Ten validated findings can be more valuable than 500 scanner alerts.&lt;/p&gt;




&lt;h1&gt;
  
  
  37. External Penetration Testing in a Continuously Changing Environment
&lt;/h1&gt;

&lt;p&gt;The traditional model is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Pentest
   ↓
Report
   ↓
Fix
   ↓
Retest
   ↓
Wait
   ↓
Next annual pentest
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The problem is the environment changes during the waiting period.&lt;/p&gt;

&lt;p&gt;A new hostname can appear tomorrow.&lt;/p&gt;

&lt;p&gt;A new API can be deployed next week.&lt;/p&gt;

&lt;p&gt;A cloud resource can become public after a configuration change.&lt;/p&gt;

&lt;p&gt;A development environment can be exposed without ever entering the original pentest scope.&lt;/p&gt;

&lt;p&gt;This is why external attack-surface monitoring is increasingly complementary to periodic penetration testing.&lt;/p&gt;

&lt;p&gt;The methodology doesn't fundamentally change.&lt;/p&gt;

&lt;p&gt;The cadence does.&lt;/p&gt;

&lt;p&gt;Instead of:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;discover → test → report → wait
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;the model becomes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;discover → test → validate → remediate → retest → repeat
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;CodeAnt AI's analysis of &lt;a href="https://codeant.ai/blogs/continuous-vs-annual-pentesting" rel="noopener noreferrer"&gt;continuous vs. annual penetration testing&lt;/a&gt; goes deeper into this difference in testing cadence.&lt;/p&gt;




&lt;h1&gt;
  
  
  38. External Penetration Testing Checklist
&lt;/h1&gt;

&lt;h2&gt;
  
  
  Reconnaissance
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;[ ] DNS records&lt;/li&gt;
&lt;li&gt;[ ] Zone transfer&lt;/li&gt;
&lt;li&gt;[ ] Certificate Transparency&lt;/li&gt;
&lt;li&gt;[ ] ASN discovery&lt;/li&gt;
&lt;li&gt;[ ] IP ranges&lt;/li&gt;
&lt;li&gt;[ ] Public repositories&lt;/li&gt;
&lt;li&gt;[ ] Secret exposure&lt;/li&gt;
&lt;li&gt;[ ] Historical URLs&lt;/li&gt;
&lt;li&gt;[ ] Search-engine indexing&lt;/li&gt;
&lt;li&gt;[ ] Public cloud resources&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Attack Surface
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;[ ] Subdomains&lt;/li&gt;
&lt;li&gt;[ ] Live hosts&lt;/li&gt;
&lt;li&gt;[ ] IPv4&lt;/li&gt;
&lt;li&gt;[ ] IPv6&lt;/li&gt;
&lt;li&gt;[ ] Open ports&lt;/li&gt;
&lt;li&gt;[ ] Service versions&lt;/li&gt;
&lt;li&gt;[ ] Web technologies&lt;/li&gt;
&lt;li&gt;[ ] APIs&lt;/li&gt;
&lt;li&gt;[ ] Staging environments&lt;/li&gt;
&lt;li&gt;[ ] Legacy applications&lt;/li&gt;
&lt;li&gt;[ ] Management interfaces&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Web/API
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;[ ] Authentication&lt;/li&gt;
&lt;li&gt;[ ] MFA&lt;/li&gt;
&lt;li&gt;[ ] Password reset&lt;/li&gt;
&lt;li&gt;[ ] Session management&lt;/li&gt;
&lt;li&gt;[ ] Authorization&lt;/li&gt;
&lt;li&gt;[ ] IDOR/BOLA&lt;/li&gt;
&lt;li&gt;[ ] API versioning&lt;/li&gt;
&lt;li&gt;[ ] GraphQL&lt;/li&gt;
&lt;li&gt;[ ] File upload&lt;/li&gt;
&lt;li&gt;[ ] Debug endpoints&lt;/li&gt;
&lt;li&gt;[ ] Sensitive files&lt;/li&gt;
&lt;li&gt;[ ] Security headers&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Infrastructure
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;[ ] RDP&lt;/li&gt;
&lt;li&gt;[ ] VNC&lt;/li&gt;
&lt;li&gt;[ ] SSH&lt;/li&gt;
&lt;li&gt;[ ] SMB&lt;/li&gt;
&lt;li&gt;[ ] FTP&lt;/li&gt;
&lt;li&gt;[ ] Databases&lt;/li&gt;
&lt;li&gt;[ ] Redis&lt;/li&gt;
&lt;li&gt;[ ] Elasticsearch&lt;/li&gt;
&lt;li&gt;[ ] Docker API&lt;/li&gt;
&lt;li&gt;[ ] Kubernetes API&lt;/li&gt;
&lt;li&gt;[ ] VPN&lt;/li&gt;
&lt;li&gt;[ ] CI/CD&lt;/li&gt;
&lt;li&gt;[ ] Monitoring interfaces&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Cloud
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;[ ] Public buckets&lt;/li&gt;
&lt;li&gt;[ ] Public databases&lt;/li&gt;
&lt;li&gt;[ ] Public management interfaces&lt;/li&gt;
&lt;li&gt;[ ] Exposed credentials&lt;/li&gt;
&lt;li&gt;[ ] IAM permissions&lt;/li&gt;
&lt;li&gt;[ ] Forgotten resources&lt;/li&gt;
&lt;li&gt;[ ] Cloud metadata exposure&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Validation
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;[ ] Scanner findings manually validated&lt;/li&gt;
&lt;li&gt;[ ] False positives removed&lt;/li&gt;
&lt;li&gt;[ ] Exploitability demonstrated&lt;/li&gt;
&lt;li&gt;[ ] Impact established&lt;/li&gt;
&lt;li&gt;[ ] Attack paths analyzed&lt;/li&gt;
&lt;li&gt;[ ] Evidence collected&lt;/li&gt;
&lt;li&gt;[ ] Remediation documented&lt;/li&gt;
&lt;li&gt;[ ] Retesting performed&lt;/li&gt;
&lt;/ul&gt;




&lt;h1&gt;
  
  
  Frequently Asked Questions
&lt;/h1&gt;

&lt;h2&gt;
  
  
  What is external penetration testing?
&lt;/h2&gt;

&lt;p&gt;External penetration testing simulates an attacker operating from outside an organization's network. The tester begins with publicly available information and attempts to discover and exploit internet-facing applications, services, infrastructure, APIs, and cloud resources within the authorized scope.&lt;/p&gt;

&lt;h2&gt;
  
  
  What tools are used for external penetration testing?
&lt;/h2&gt;

&lt;p&gt;Common tools include Nmap for port and service discovery, subfinder and Amass for subdomain enumeration, httpx for HTTP probing, ffuf and Gobuster for endpoint discovery, Nuclei for automated vulnerability detection, Burp Suite for manual web testing, and tools such as TruffleHog for secret discovery.&lt;/p&gt;

&lt;p&gt;No individual tool provides complete coverage. Effective testing combines automated discovery with manual validation.&lt;/p&gt;

&lt;h2&gt;
  
  
  What should be tested first in an external pentest?
&lt;/h2&gt;

&lt;p&gt;Attack-surface discovery should come before deep vulnerability testing. Start with domains, DNS, certificates, subdomains, IP ranges, cloud resources, historical URLs, and public repositories. Then move into active host discovery, port scanning, application enumeration, and vulnerability testing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Is an open port automatically a vulnerability?
&lt;/h2&gt;

&lt;p&gt;No. An open port indicates that a service is reachable. Whether that represents a vulnerability depends on the service, authentication, configuration, network controls, patch level, and what an attacker can accomplish after connecting.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is the difference between vulnerability scanning and penetration testing?
&lt;/h2&gt;

&lt;p&gt;Vulnerability scanning primarily identifies potential weaknesses using automated detection techniques. Penetration testing goes further by manually validating vulnerabilities, testing application logic, demonstrating impact, and determining whether individual weaknesses can be chained into meaningful attack paths.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why is API testing important in external penetration testing?
&lt;/h2&gt;

&lt;p&gt;Modern applications often expose significant functionality through APIs. APIs may implement authentication and authorization differently from the web interface, and older API versions can remain accessible even after the frontend stops using them. Testing alternate API paths is therefore an important part of external attack-surface analysis.&lt;/p&gt;

&lt;h2&gt;
  
  
  How often should external penetration testing be performed?
&lt;/h2&gt;

&lt;p&gt;The required frequency depends on the organization's regulatory, contractual, and risk requirements. Periodic manual penetration testing can be supplemented by continuous external attack-surface monitoring to identify newly exposed assets and configuration drift between formal engagements.&lt;/p&gt;




&lt;h1&gt;
  
  
  Conclusion
&lt;/h1&gt;

&lt;p&gt;External penetration testing is not fundamentally a port-scanning exercise.&lt;/p&gt;

&lt;p&gt;It's an exercise in &lt;strong&gt;mapping trust boundaries from the outside&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The process starts with passive information:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;DNS
Certificates
ASN data
Repositories
Historical URLs
Cloud infrastructure
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That becomes an asset inventory.&lt;/p&gt;

&lt;p&gt;The inventory becomes active reconnaissance:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Subdomains
IPs
Ports
Services
Applications
APIs
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Those discoveries become testing targets.&lt;/p&gt;

&lt;p&gt;Then comes the part that requires the most judgment:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Is this actually vulnerable?
Can it be exploited?
What access does it provide?
Can it be chained with another weakness?
What can an attacker ultimately accomplish?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That distinction matters.&lt;/p&gt;

&lt;p&gt;A scanner can tell you that port 5432 is open.&lt;/p&gt;

&lt;p&gt;A penetration tester determines whether the database is authenticated, what account is accessible, what data is exposed, and whether that access creates a meaningful attack path.&lt;/p&gt;

&lt;p&gt;A scanner can discover &lt;code&gt;/api/v1&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;A tester determines whether the old API still works, whether authentication is enforced, whether authorization is correct, and whether it exposes functionality that &lt;code&gt;/api/v2&lt;/code&gt; properly protects.&lt;/p&gt;

&lt;p&gt;A scanner can identify a staging hostname.&lt;/p&gt;

&lt;p&gt;A tester asks why it exists, what it contains, what credentials it accepts, and what systems it can reach.&lt;/p&gt;

&lt;p&gt;That is the difference between &lt;strong&gt;finding vulnerabilities&lt;/strong&gt; and &lt;strong&gt;understanding an attack surface&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;And as cloud infrastructure, APIs, CI/CD systems, and ephemeral environments continue to change, the attack surface is no longer something that can be accurately measured once a year.&lt;/p&gt;

&lt;p&gt;The methodology remains:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Discover → Enumerate → Test → Validate → Exploit → Chain → Report → Retest.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The challenge is making sure you're still discovering the right things when the environment changes tomorrow.&lt;/p&gt;

</description>
      <category>security</category>
      <category>cybersecurity</category>
      <category>penetrationtesting</category>
      <category>vulnerabilities</category>
    </item>
    <item>
      <title>IDOR: The Vulnerability Class Your Scanners Will Never Find</title>
      <dc:creator>Amartya Jha</dc:creator>
      <pubDate>Tue, 08 Sep 2026 06:57:57 +0000</pubDate>
      <link>https://dev.to/codeant/idor-the-vulnerability-class-your-scanners-will-never-find-37li</link>
      <guid>https://dev.to/codeant/idor-the-vulnerability-class-your-scanners-will-never-find-37li</guid>
      <description>&lt;p&gt;In 2023, a researcher changed a single number in a URL, &lt;code&gt;/api/orders/10021&lt;/code&gt; to &lt;code&gt;/api/orders/10022&lt;/code&gt;, and got back a complete order record belonging to someone else. Name, address, items purchased, last four digits of a payment card. No exploit chain, no malware, no clever payload. Just a missing ownership check.&lt;/p&gt;

&lt;p&gt;That's IDOR: Insecure Direct Object Reference. In the OWASP API Security Top 10 it's called BOLA, Broken Object Level Authorization, and it's API1:2023. Different name, same failure: the application confirms you're logged in, but never confirms you're allowed to touch the specific object you asked for.&lt;/p&gt;

&lt;p&gt;It's arguably the most consistently present vulnerability class in modern web apps and APIs, and it causes more real SaaS data breaches than SQL injection, XSS, and RCE combined, not because it's sophisticated, but because it's invisible to most of the tooling teams rely on by default.&lt;/p&gt;

&lt;h2&gt;
  
  
  The precise definition
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fv8znha1441wz9q5g71t5.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fv8znha1441wz9q5g71t5.png" alt=" " width="799" height="410"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;An IDOR exists when three things are all true: the app accepts a user-supplied identifier for an internal object, uses that identifier to fetch the object, and never verifies the requesting user is actually authorized to access that specific object. The vulnerability isn't in the identifier existing or being guessable. It's in that third step, the missing gap between "an ID was supplied" and "this object got returned anyway."&lt;/p&gt;

&lt;p&gt;The right security question was never "can the user change the ID." It's "after the ID changes, does the server still check whether this specific authenticated user is allowed to touch this specific object." If the answer is no, you have an IDOR/BOLA.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it actually looks like, across six different shapes
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Sequential IDs.&lt;/strong&gt; &lt;code&gt;GET /api/orders/10021&lt;/code&gt; returns User A's order legitimately. Swap it to &lt;code&gt;/api/orders/10022&lt;/code&gt;, and if that returns User B's order without an ownership check, that's the classic pattern. The sequential IDs aren't the bug, the missing authorization check on top of them is.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;BOLA in an API.&lt;/strong&gt; A vehicle-control API: &lt;code&gt;POST /api/vehicles/ABC123/doors/unlock&lt;/code&gt; works fine for the owner. Swap the vehicle ID to one belonging to someone else, and if the server unlocks it anyway, that's API1:2023 by definition, OWASP uses almost this exact example.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;UUIDs don't fix it.&lt;/strong&gt; &lt;code&gt;GET /api/documents/f47ac10b-58cc-4372-a567-0e02b2c3d479&lt;/code&gt; looks unguessable, but User A can still get User B's UUID through a shared link, an API response, a notification, browser history, any legitimate flow. Predictability affects how easily an ID is &lt;em&gt;discovered&lt;/em&gt;. Authorization determines whether it can be &lt;em&gt;used&lt;/em&gt;. Those are separate properties, and switching ID format only fixes the first one.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;File references.&lt;/strong&gt; &lt;code&gt;GET /static/invoices/10021.pdf&lt;/code&gt; swapped to &lt;code&gt;10022.pdf&lt;/code&gt; is the same bug wearing a different hat, applies equally to profile images, exported reports, generated PDFs, chat transcripts.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;GraphQL.&lt;/strong&gt; A mutation like &lt;code&gt;deleteDocument(id: "doc-10022")&lt;/code&gt; can delete an object the caller doesn't own just as easily as a REST endpoint can leak one. The identifier's transport (URL path, query param, JSON body, GraphQL variable) is irrelevant. The server still has to authorize the object before acting on it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Multi-tenant.&lt;/strong&gt; &lt;code&gt;GET /api/organizations/acme/invoices&lt;/code&gt; swapped to &lt;code&gt;/api/organizations/other-company/invoices&lt;/code&gt; is the same failure at organizational scale, and it's the most damaging variant because a single missing check can expose an entire company's data instead of one user's.&lt;/p&gt;

&lt;h2&gt;
  
  
  The taxonomy that actually matters for triage
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Horizontal&lt;/strong&gt; — same privilege level, different user. The most common type. CVSS roughly 6.5 (read) to 8.8 (write/delete).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Vertical&lt;/strong&gt; — a standard user reaches admin or elevated-privilege data by supplying a privileged resource's ID. Roughly 7.5–9.1.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Blind&lt;/strong&gt; — the action succeeds but nothing comes back in the response. No visible data exposure, so it's chronically underestimated, but it causes real integrity failures: deleted content, modified state, disrupted workflows, with no trace unless you check the target resource afterward. Roughly 5.0–8.8.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Second-order&lt;/strong&gt; — an identifier gets accepted in one step of a workflow, stored, and reused later without re-validating ownership. No automated tool understands multi-step context this well; the endpoint looks completely correct in isolation. Often 7.5–9.5, because it tends to bypass the most sensitive step in a workflow.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Multi-tenant&lt;/strong&gt; — cross-organization access. A force multiplier: one horizontal IDOR leaks one user's data, one multi-tenant IDOR can leak an entire org's, potentially thousands of records in a single request. Almost always 8.8–9.5.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Mass assignment&lt;/strong&gt; — the request body includes fields like &lt;code&gt;user_id&lt;/code&gt;, &lt;code&gt;owner_id&lt;/code&gt;, or &lt;code&gt;role&lt;/code&gt; that the client shouldn't be able to set, and the server processes them unfiltered.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Scoring it correctly
&lt;/h2&gt;

&lt;p&gt;Most IDOR findings get under-scored because whoever's scoring focuses on the single HTTP request instead of what it enables at scale. The good news is that almost every parameter is fixed across web/API IDORs, so scoring mostly comes down to one variable: impact.&lt;/p&gt;

&lt;p&gt;Fixed CVSS 4.0 parameters for nearly every case: Attack Vector Network, Attack Complexity Low, Attack Requirements None, Privileges Required Low, User Interaction None. Two exceptions: second-order IDORs get Attack Requirements Present, since specific workflow state has to exist first, and unauthenticated IDORs (rare, but they happen) get Privileges Required None, which pushes the score straight to critical.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Scenario&lt;/th&gt;
&lt;th&gt;What changes&lt;/th&gt;
&lt;th&gt;CVSS 4.0&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Read, single low-sensitivity record&lt;/td&gt;
&lt;td&gt;VC: Low&lt;/td&gt;
&lt;td&gt;~5.3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Read, single high-sensitivity record (PII/financial)&lt;/td&gt;
&lt;td&gt;VC: High&lt;/td&gt;
&lt;td&gt;~6.9&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Read, all records for one user&lt;/td&gt;
&lt;td&gt;VC: High&lt;/td&gt;
&lt;td&gt;~7.1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Read, all records across all users&lt;/td&gt;
&lt;td&gt;VC: High + SC: High&lt;/td&gt;
&lt;td&gt;~8.8&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Read, multi-tenant, full org exposure&lt;/td&gt;
&lt;td&gt;VC: High + SC: High&lt;/td&gt;
&lt;td&gt;~9.1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Write, modify another user's resource&lt;/td&gt;
&lt;td&gt;VI: High&lt;/td&gt;
&lt;td&gt;~8.8&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Write, modify an admin-level resource&lt;/td&gt;
&lt;td&gt;VI: High + SC: High&lt;/td&gt;
&lt;td&gt;~9.1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Delete, another user's data (blind)&lt;/td&gt;
&lt;td&gt;VI: High&lt;/td&gt;
&lt;td&gt;~7.5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Delete, critical business records&lt;/td&gt;
&lt;td&gt;VI: High + SC: High&lt;/td&gt;
&lt;td&gt;~8.5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Multi-tenant, any of the above&lt;/td&gt;
&lt;td&gt;Always add SC: High&lt;/td&gt;
&lt;td&gt;8.8–9.5&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two questions decide almost the entire score: how many records are actually in scope, and can the attacker write or delete, or only read. Write and delete consistently score higher than read at equivalent scope, and multi-tenant always adds SC: High because impact extends beyond the attacker's own org.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why your existing tools structurally can't catch this
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsm3jmch92qurnqndkujr.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsm3jmch92qurnqndkujr.png" alt=" " width="800" height="437"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;DAST operates at the HTTP request level: send a request, look at the response, compare patterns. It has no concept of ownership. That's not a maturity gap in any particular tool, it's a structural limit of request-level testing without object-ownership context.&lt;/p&gt;

&lt;p&gt;Static analysis flags patterns like "identifier from user input used in a query without a visible authorization check nearby," and in practice this produces a false-positive rate that consistently exceeds 50%. Engineers burn time dismissing noise, and real IDORs buried in a complex call chain still get missed, because an authorization check exists &lt;em&gt;somewhere&lt;/em&gt; in the codebase, just not the right one at the right level for this specific resource.&lt;/p&gt;

&lt;p&gt;The actual test IDOR needs is a runtime, cross-identity question: is this authorization check actually enforced, for this resource type, from this identity, at this point in this specific workflow. That requires an authenticated identity, a resource it owns, a second identity that doesn't own it, and systematic testing of every endpoint with both, comparing results. That's a fundamentally different exercise than anything a single-request scanner does.&lt;/p&gt;

&lt;h2&gt;
  
  
  Testing methodology, phase by phase
&lt;/h2&gt;

&lt;p&gt;IDOR testing needs at least two accounts with known, distinct resource ownership, and ideally two separate tenants if the app is multi-tenant. From there:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Endpoint discovery.&lt;/strong&gt; Pull every endpoint from the OpenAPI spec and from JS bundle analysis, since bundles often expose internal paths the spec omits. Flag every endpoint that accepts an object identifier, note whether it's in a path, query param, body, or header, and classify by method, since write/delete IDORs generally outscore read-only ones.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Horizontal testing.&lt;/strong&gt; Replay Account A's authenticated request using Account B's resource IDs, across every method, not just GET. Test batch endpoints specifically by mixing owned and unowned IDs in one request, batch endpoints frequently check authorization once for the whole collection instead of per item. Test export/download endpoints, teams treat these as secondary and skip authorization review on them constantly. Test search/filter parameters using another user's ID as the filter value.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Vertical testing.&lt;/strong&gt; Look for admin resource IDs leaking through error messages or metadata, then try accessing them with a standard-user token.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Multi-tenant testing.&lt;/strong&gt; Find every place &lt;code&gt;org_id&lt;/code&gt;/&lt;code&gt;tenant_id&lt;/code&gt; shows up (path, query param, body) and try substituting another org's ID in each location independently. Path-based org scoping (&lt;code&gt;/api/orgs/{org_id}/resources&lt;/code&gt;) is the most common pattern and the one most often missing validation.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Second-order testing.&lt;/strong&gt; Map multi-step workflows, create a resource as Account A, then try to resume or complete the workflow as Account B. The common failure: ownership gets checked at creation but never re-checked at resumption.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Blind testing.&lt;/strong&gt; Hit every DELETE and action-POST endpoint with another user's resource ID. A 204 No Content is not a passing result by itself, the only way to confirm impact is to check the resource afterward as its actual owner and see whether it changed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Fixing it
&lt;/h2&gt;

&lt;p&gt;The fix is always server-side, object-level authorization, never identifier obscurity.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# VULNERABLE
&lt;/span&gt;&lt;span class="n"&gt;invoice&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;Invoice&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get_or_404&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;invoice_id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;jsonify&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;invoice&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;to_dict&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The identifier might be valid, but nothing here confirms the authenticated user owns this invoice.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# SECURE
&lt;/span&gt;&lt;span class="n"&gt;invoice&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;Invoice&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;filter_by&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="nb"&gt;id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;invoice_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;user_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;current_user&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nb"&gt;id&lt;/span&gt;
&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;first_or_404&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;jsonify&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;invoice&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;to_dict&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The change that matters isn't the identifier format, it's the ownership constraint applied server-side. A few patterns that generalize: derive ownership exclusively from the authenticated session, never from anything the client claims; make an ownership filter a structural requirement on every query touching user-owned data, not an optional extra; for multi-tenant apps, enforce row-level tenant isolation at the ORM or database layer so it can't be skipped by a careless handler; explicitly filter mass-assignment fields like &lt;code&gt;user_id&lt;/code&gt;, &lt;code&gt;owner_id&lt;/code&gt;, and &lt;code&gt;role&lt;/code&gt; out of client-writable input; and for multi-step workflows, re-validate ownership at every single step, never assume a check from step one still holds by step three.&lt;/p&gt;

&lt;h2&gt;
  
  
  Testing it manually with Burp
&lt;/h2&gt;

&lt;p&gt;Browse the app authenticated, capture requests carrying anything that looks like an object reference, not just numeric IDs: UUIDs, usernames, slugs, filenames, tenant IDs, GraphQL variables. Send a request for your own object to Repeater and confirm the baseline response. Then change only the object identifier, keep the same auth token, and see what comes back.&lt;/p&gt;

&lt;p&gt;Don't trust status code alone. A vulnerable endpoint can return &lt;code&gt;200 OK&lt;/code&gt; with someone else's data, and a properly protected one might also return &lt;code&gt;200 OK&lt;/code&gt; with an error message in the body instead of a 403. Compare status, response body content, whether the object identifier in the response actually matches what was requested, response length against a known-denied baseline, and any side effects (did it modify or delete something).&lt;/p&gt;

&lt;p&gt;Confirmation requires all of: User A is authenticated, User A can access Object A, Object B belongs to User B, User A swaps the reference to B, the server returns or modifies Object B, and no additional authorization check intervened. That mismatch between "who's authenticated" and "what they were allowed to touch" is the entire vulnerability.&lt;/p&gt;

&lt;h2&gt;
  
  
  The UUID misconception, one more time
&lt;/h2&gt;

&lt;p&gt;"We switched to UUIDs so IDOR isn't a problem anymore" is one of the most persistent wrong beliefs in web security. UUIDs prevent enumeration. They do nothing to prevent unauthorized access once an identifier is obtained through some legitimate path, a shared link, a leaked response, a notification email. The fix was never about hiding the identifier. It's ownership verification at the server, full stop, and it's identical work whether the ID is &lt;code&gt;10022&lt;/code&gt; or a UUID.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the blast radius here is different
&lt;/h2&gt;

&lt;p&gt;Most vulnerability classes have a bounded blast radius. XSS affects visitors to one page. A given SQL injection exposes one query's results. IDOR's blast radius scales directly with your user count and data sensitivity, because it's one missing check away from every record that check was supposed to protect. A SaaS app with 10,000 customer records has 10,000 records riding on a single ownership filter. A multi-tenant platform with 500 customer orgs has 500 organizations' worth of data one parameter swap away from any authenticated user who thinks to try it.&lt;/p&gt;

&lt;p&gt;It doesn't take exploit code or specialized tooling. It takes patience and arithmetic, and it lives entirely in application logic, which is exactly the layer most default scanning tooling doesn't reach. The only detection method that actually works is authentication-aware, ownership-tracking, multi-identity testing that asks "what if this belonged to someone else" at every endpoint, every method, every workflow step, systematically rather than opportunistically.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.codeant.ai/" rel="noopener noreferrer"&gt;CodeAnt AI's penetration testing platform&lt;/a&gt; is built around exactly that cross-identity testing model, authenticating as multiple real identities simultaneously and tracking ownership across the full API surface rather than scanning requests in isolation. For the multi-tenant variant specifically, which tends to be the highest-severity flavor of this bug, see &lt;a href="https://www.codeant.ai/blog/multi-tenant-saas-penetration-testing" rel="noopener noreferrer"&gt;CodeAnt's multi-tenant SaaS penetration testing guide&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>security</category>
      <category>webdev</category>
      <category>appsec</category>
      <category>api</category>
    </item>
    <item>
      <title>Black Box vs White Box vs Gray Box Pentesting: Picking the Right One (Most Teams Don't)</title>
      <dc:creator>Amartya Jha</dc:creator>
      <pubDate>Mon, 07 Sep 2026 11:16:50 +0000</pubDate>
      <link>https://dev.to/codeant/black-box-vs-white-box-vs-gray-box-pentesting-picking-the-right-one-most-teams-dont-3316</link>
      <guid>https://dev.to/codeant/black-box-vs-white-box-vs-gray-box-pentesting-picking-the-right-one-most-teams-dont-3316</guid>
      <description>&lt;p&gt;"What type of penetration test do we need" is usually the second question a team asks, right after "do we need a pentest." It deserves more thought than it gets, because the three methodologies, black box, white box, and gray box, don't just differ in thoroughness. They differ in what they're structurally capable of seeing at all. A clean report from one says nothing about the other two.&lt;/p&gt;

&lt;p&gt;The difference comes down to one variable: what the tester knows and can access before the engagement starts. That starting point determines which vulnerability classes are reachable and which are, by construction, invisible to that particular test.&lt;/p&gt;

&lt;h2&gt;
  
  
  Black box: attacker starting from nothing
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4mm83mf1ai4mmy2ct1mh.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4mm83mf1ai4mmy2ct1mh.png" alt=" " width="800" height="381"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A black box tester gets a domain. No credentials, no source, no architecture docs. The question this answers is precise: what can someone on the internet, with zero inside knowledge, actually do to your data?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Reconnaissance is the foundation, and it's more thorough than most teams expect.&lt;/strong&gt; Subdomain enumeration brute-forces DNS across 150+ prefix patterns, not just &lt;code&gt;www&lt;/code&gt; and &lt;code&gt;api&lt;/code&gt;, but &lt;code&gt;dev&lt;/code&gt;, &lt;code&gt;staging&lt;/code&gt;, &lt;code&gt;uat&lt;/code&gt;, &lt;code&gt;internal&lt;/code&gt;, &lt;code&gt;jenkins&lt;/code&gt;, &lt;code&gt;grafana&lt;/code&gt;, &lt;code&gt;admin&lt;/code&gt;. Certificate Transparency logs get queried too, since every TLS cert ever issued for the domain is publicly logged there, which surfaces historical subdomains DNS brute-forcing misses entirely, including ones nobody remembers are still running a server. Port scanning covers all TCP ports, not just 80/443, which is how an exposed Redis instance or an unauthenticated Elasticsearch cluster gets found. It happens more than it should.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cloud asset discovery extends the same logic to infrastructure that isn't the app itself:&lt;/strong&gt; S3 buckets checked for public read/write, Azure Blob containers for anonymous access, GCP buckets for &lt;code&gt;allUsers&lt;/code&gt; permissions, CI/CD dashboards (Jenkins, CircleCI, GitHub Actions) checked for missing auth, monitoring endpoints (Grafana, Kibana, Datadog) checked the same way.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;JS bundle analysis is the step most traditional black box engagements skip, and it's one of the highest-value ones.&lt;/strong&gt; A modern SPA ships 5–15MB of minified JavaScript to every visitor's browser, and that bundle gets statically analyzed for hardcoded secrets across 30+ pattern types, AWS keys, Stripe live keys, GitHub tokens, JWT secrets, Twilio and SendGrid credentials, each one verified for validity before it's reported. Comparing staging and production bundles also surfaces endpoints that were pulled from prod but are still live on a non-production URL with weaker controls, which is a surprisingly common way to find a forgotten API.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Every discovered endpoint gets hit unauthenticated first&lt;/strong&gt;, and the response gets classified: &lt;code&gt;200 OK&lt;/code&gt; with real data means no auth enforced, full stop. &lt;code&gt;500&lt;/code&gt; can mean the request got processed before the auth check ever ran. &lt;code&gt;403&lt;/code&gt; gets checked for a bypass rather than taken at face value. CORS gets tested against 7+ attacker-controlled origins, since misconfigured CORS shows up in production constantly.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Findings get chained, not reported in isolation.&lt;/strong&gt; A tenant ID leaking from a profile endpoint plus an IDOR on the records endpoint equals full cross-tenant access. A hardcoded internal hostname in the JS bundle plus an unauthenticated endpoint on that internal API equals unauthenticated access to something that was never meant to be reachable at all. The combination is consistently more dangerous than either finding alone.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What black box structurally cannot see:&lt;/strong&gt; auth bypasses buried in middleware config that still return normal-looking responses, business logic flaws behind a login wall, secrets sitting in Git history, anything on an internal service never exposed to the internet, dependency vulnerabilities that need code access to assess reachability. This isn't a weakness in the methodology, it's the direct consequence of the threat model. A black box test simulates someone with nothing. It can only ever see what nothing gets you.&lt;/p&gt;

&lt;h2&gt;
  
  
  White box: reading the implementation directly
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fblmj6o53a78faybzzkgd.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fblmj6o53a78faybzzkgd.png" alt=" " width="800" height="577"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;White box gives the tester the code. Configs, infrastructure definitions, dependency manifests, architecture docs, version history. The question shifts from "what can an attacker discover" to "what's actually wrong in the implementation, whether or not it's visible from outside."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Security configuration gets read directly, which catches things scanning never will.&lt;/strong&gt; A Spring Security filter chain excluded for an entire &lt;code&gt;/api/v2/&lt;/code&gt; namespace looks completely normal from outside, the endpoint just responds. An external scanner has no way to know the auth layer got skipped entirely for that namespace. Same story with Express middleware ordering: an admin endpoint can return &lt;code&gt;200 OK&lt;/code&gt; with real data to an unauthenticated request because the auth middleware got registered after the route handler, and nothing about the HTTP response gives that away.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Secrets scanning covers current HEAD and history separately.&lt;/strong&gt; A credential committed and later deleted from the working tree is still sitting in version control, recoverable by anyone with clone access, and that's a distinct check from scanning what's currently checked in.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Dataflow tracing is where white box earns its reputation for precision.&lt;/strong&gt; Instead of "SQL injection detected," a proper trace produces something like: &lt;code&gt;app/views/products.py&lt;/code&gt;, line 14, &lt;code&gt;search_products()&lt;/code&gt;, the &lt;code&gt;category&lt;/code&gt; parameter from &lt;code&gt;request.GET&lt;/code&gt; reaches a raw query via string formatting, payload &lt;code&gt;' OR '1'='1' --&lt;/code&gt;, effect: returns all products regardless of category or featured status, root cause: &lt;code&gt;Product.objects.raw()&lt;/code&gt; with f-string interpolation instead of a parameterized query. That specificity is the difference between an engineer fixing the actual line and an engineer guessing at what "SQL injection" means for their codebase.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Dependency analysis goes past CVE matching into reachability.&lt;/strong&gt; A vulnerable package that's imported but never actually called from a code path the app executes is a very different risk than the same CVE sitting in a function that runs on every file upload. Reachability analysis is what separates the two, and it's the single biggest lever for cutting dependency-scan false positives.&lt;/p&gt;

&lt;h2&gt;
  
  
  Gray box: what a legitimate account can get away with
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Feeau5mzv4thk3huwbvof.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Feeau5mzv4thk3huwbvof.png" alt=" " width="800" height="436"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Gray box starts with authenticated access, one or more real test accounts, sometimes limited docs or API specs on top. The question: what can someone do once they already have valid credentials?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Access control and privilege escalation testing hits every admin-level endpoint with non-admin credentials&lt;/strong&gt;, and JWT claims get manipulated directly to check whether role or tenant claims are actually re-validated server-side or just trusted from the token.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;IDOR testing here is systematic, not exploratory.&lt;/strong&gt; Every endpoint accepting an object identifier, sequential integer, UUID, slug, username, filename, tenant ID, or an ID buried in a JSON or GraphQL body, gets evaluated for object-level authorization. The core test: can a user authorized for Object A simply substitute Object B's identifier and read, modify, or delete it. In multi-tenant apps this same check has to run at the org boundary, not just the user boundary, since a missing tenant check leaks an entire company's data instead of one account's. (&lt;a href="https://www.codeant.ai/blogs/idor-vulnerabilities" rel="noopener noreferrer"&gt;We've written the technical breakdown of this vulnerability class in full&lt;/a&gt;, including every variant and how to test for it, if you want to go deeper than this section.)&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Business logic testing is where gray box finds things nothing else reaches&lt;/strong&gt;, because none of these produce an anomalous HTTP response or match a known CVE signature. They require understanding what the app is supposed to enforce, then checking whether it actually does, at every entry point:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Can the order total be modified in the request before payment confirms?&lt;/li&gt;
&lt;li&gt;Can a single-use discount code be replayed by resending the validation call?&lt;/li&gt;
&lt;li&gt;Can checkout step 5 be hit directly without completing steps 1 through 4?&lt;/li&gt;
&lt;li&gt;Can a free-tier account call a premium endpoint directly via the API?&lt;/li&gt;
&lt;li&gt;Can rate limits be evaded by rotating user IDs or spoofed IP headers?&lt;/li&gt;
&lt;li&gt;Can negative quantities reduce a total in an e-commerce flow?&lt;/li&gt;
&lt;li&gt;Does a race condition in inventory or balance checks survive two simultaneous requests?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of this shows up in a scanner's output. All of it ships to production regularly.&lt;/p&gt;

&lt;h2&gt;
  
  
  The comparison that actually matters
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Black box&lt;/th&gt;
&lt;th&gt;White box&lt;/th&gt;
&lt;th&gt;Gray box&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Simulates&lt;/td&gt;
&lt;td&gt;External attacker, no prior knowledge&lt;/td&gt;
&lt;td&gt;Insider/attacker with implementation access&lt;/td&gt;
&lt;td&gt;Legitimate user or compromised account&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Starting point&lt;/td&gt;
&lt;td&gt;Domain + public info&lt;/td&gt;
&lt;td&gt;Source, config, architecture, dependencies&lt;/td&gt;
&lt;td&gt;Test credentials + selected context&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Strong at&lt;/td&gt;
&lt;td&gt;Exposed services, unauth endpoints, cloud misconfig&lt;/td&gt;
&lt;td&gt;Code-level auth flaws, data flow issues, secrets, dependency risk&lt;/td&gt;
&lt;td&gt;IDOR, privilege escalation, broken access control, business logic&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Structurally misses&lt;/td&gt;
&lt;td&gt;Anything needing auth or internal visibility&lt;/td&gt;
&lt;td&gt;Runtime behavior specific to production conditions&lt;/td&gt;
&lt;td&gt;Unauthenticated external exposure&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Answers&lt;/td&gt;
&lt;td&gt;"What can an external attacker reach?"&lt;/td&gt;
&lt;td&gt;"What's actually wrong in the implementation?"&lt;/td&gt;
&lt;td&gt;"What can an authenticated user get away with?"&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;None of these are difficulty tiers stacked on top of each other. They're three different threat models with three different structural blind spots, and the important asymmetry is this: a clean result from any one of them tells you nothing about the other two. A spotless black box report doesn't mean the code is clean. A spotless white box audit doesn't mean nothing's exposed externally. A spotless gray box assessment says nothing about whether the unauthenticated surface holds up.&lt;/p&gt;

&lt;h2&gt;
  
  
  Picking one (or more)
&lt;/h2&gt;

&lt;p&gt;If you've never had a real assessment, run all three together. Partial coverage produces the worst outcome in security, which is false confidence from a report that never looked where the actual risk was sitting.&lt;/p&gt;

&lt;p&gt;If you're pre-launch with a product handling customer data, prioritize gray box plus white box. Business logic flaws and code-level auth bugs are exactly what ships in a first release, and the external surface can be addressed continuously once the app is actually live and has real exposure to test against.&lt;/p&gt;

&lt;p&gt;If you've already run black box tests before (most first pentests are black box by default) and never gone further, white box is very likely your highest-value next investment. Most prior engagements never touched the code, and that's where the deepest, highest-severity findings tend to live.&lt;/p&gt;

&lt;p&gt;If your specific concern is a compromised account crossing into another customer's data, gray box is the direct answer, since it's the only methodology built to test tenant and object-level boundaries under real authenticated conditions.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why chaining across all three beats running them separately
&lt;/h2&gt;

&lt;p&gt;The real value isn't additive coverage, it's that findings from one perspective give the other two something to test against. Black box finds an exposed API. White box traces that same endpoint into its authorization logic and shows exactly how (or whether) access gets enforced. Gray box then hits that endpoint with real credentials to check whether an authenticated user can swap an object ID and reach someone else's data.&lt;/p&gt;

&lt;p&gt;A vulnerability that spans layers like this is easy to miss with any single methodology and hard to miss once the three are cross-referenced against each other. An exposed endpoint only becomes dangerous paired with a missing authorization check behind it. A code-level weakness only becomes exploitable once you know the corresponding endpoint is actually reachable from outside. Treat black box, white box, and gray box as three independent security programs and you'll keep missing exactly this kind of finding, the one that only exists at the seam between them.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.codeant.ai/" rel="noopener noreferrer"&gt;CodeAnt AI's penetration testing platform&lt;/a&gt; runs black box, white box, and gray box as a single engagement with attack-chain validation across all three, rather than three disconnected reports that never talk to each other.&lt;/p&gt;

</description>
      <category>security</category>
      <category>devops</category>
      <category>pentesting</category>
      <category>appsec</category>
    </item>
    <item>
      <title>Red Team Authorization: Solving the Paradox of Testing People Who Can't Know They're Being Tested</title>
      <dc:creator>Amartya Jha</dc:creator>
      <pubDate>Mon, 07 Sep 2026 07:48:24 +0000</pubDate>
      <link>https://dev.to/codeant/red-team-authorization-solving-the-paradox-of-testing-people-who-cant-know-theyre-being-tested-2omb</link>
      <guid>https://dev.to/codeant/red-team-authorization-solving-the-paradox-of-testing-people-who-cant-know-theyre-being-tested-2omb</guid>
      <description>&lt;p&gt;A red team engagement exists to answer one question: can your security team detect and respond to a real attack. For that answer to mean anything, the security team being tested, the blue team, cannot know the test is happening.&lt;/p&gt;

&lt;p&gt;That single constraint breaks the normal authorization model. Standard authorization flows through the people who own the systems being tested, and for a pentest, that includes security and IT. For a red team, those are exactly the people who have to stay in the dark. Get this wrong in one direction and the blue team finds out, which ruins the test. Get it wrong in the other direction and nobody with real authority actually approved the engagement, which puts the testers in genuine legal exposure if something goes sideways mid-operation. Both failure modes defeat the entire point of running a red team in the first place.&lt;/p&gt;

&lt;p&gt;The fix is compartmentalization: authorization flows from a small group of executives who are explicitly, provably not part of the blue team being evaluated.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this isn't just a stricter pentest authorization
&lt;/h2&gt;

&lt;p&gt;A red team authorization letter is a different document with different legal and operational requirements, not a pentest letter with a tighter NDA bolted on. A red team may deliberately interact with people, facilities, and defensive systems that have no idea a test is running.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Standard pentest&lt;/th&gt;
&lt;th&gt;Red team engagement&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Objective&lt;/td&gt;
&lt;td&gt;Identify and validate vulnerabilities&lt;/td&gt;
&lt;td&gt;Test detection, response, and containment&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Who knows&lt;/td&gt;
&lt;td&gt;Security, IT, system owners&lt;/td&gt;
&lt;td&gt;Restricted executive/legal group; blue team often excluded&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Signing authority&lt;/td&gt;
&lt;td&gt;CISO, CTO, or delegated security exec&lt;/td&gt;
&lt;td&gt;Exec sponsor with authority over the target, independent of the blue team&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Scope&lt;/td&gt;
&lt;td&gt;Apps, APIs, infra, cloud assets&lt;/td&gt;
&lt;td&gt;Technical systems plus facilities, personnel, identities, defensive controls&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Social engineering&lt;/td&gt;
&lt;td&gt;Usually excluded or separately authorized&lt;/td&gt;
&lt;td&gt;May be explicitly authorized&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Physical access&lt;/td&gt;
&lt;td&gt;Usually excluded or separately authorized&lt;/td&gt;
&lt;td&gt;May include offices, badge systems, tailgating&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Blue-team awareness&lt;/td&gt;
&lt;td&gt;Usually expected&lt;/td&gt;
&lt;td&gt;Often intentionally restricted&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Emergency contact&lt;/td&gt;
&lt;td&gt;Testing firm + client security contact&lt;/td&gt;
&lt;td&gt;24/7 executive sponsor who can confirm the op immediately&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Law enforcement&lt;/td&gt;
&lt;td&gt;Addressed in emergency procedures&lt;/td&gt;
&lt;td&gt;Must explicitly define how authorization gets verified on the spot&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Who's actually allowed to sign
&lt;/h2&gt;

&lt;p&gt;The signing rule is simple to state and easy to get wrong in practice: the signer cannot be part of the group whose detection capability is being evaluated.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;If the CISO manages the SOC being tested, the CISO cannot sign.&lt;/li&gt;
&lt;li&gt;If the entire security function is in scope, it goes to CEO or board level.&lt;/li&gt;
&lt;li&gt;If security manages but doesn't operationally run the specific test target, the CISO may sign, provided they're genuinely separate from the blue team in question.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This isn't "the CEO always signs." It's that the signer needs actual authority over the assets in scope and needs to sit outside the operational group being evaluated. Most red team authorizations end up at CEO or board level for a simple reason: the engagement is testing the security function itself, and the person who runs that function can't objectively authorize their own team's evaluation.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cloud providers have their own rules, and they change
&lt;/h2&gt;

&lt;p&gt;Red team authorization from your organization doesn't extend to permission from your cloud provider. AWS, Azure, and GCP each publish their own acceptable-use and security-testing policies, and they get updated. Verify the current version immediately before the engagement rather than trusting whatever was true last year.&lt;/p&gt;

&lt;p&gt;For every cloud provider in scope, document: the provider and account/subscription/project identifier, specific regions and resources, the approved testing window, approved and explicitly prohibited techniques, any notification or approval requirements the provider imposes, emergency contacts, executive authorization, and stop conditions. Treat this as its own verification step, separate from the signed authorization letter, not something you assume is automatically covered by it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The sealed envelope model
&lt;/h2&gt;

&lt;p&gt;When getting an executive signature at engagement start is genuinely difficult for legal, timing, or organizational reasons, a sealed envelope approach works: the fully signed authorization exists, but it's only opened if a legal issue comes up, law enforcement gets involved, or the organization faces an inquiry that requires documented proof. This lets the engagement run under strict compartmentalization while real, verifiable authorization still exists if it's ever needed.&lt;/p&gt;

&lt;h2&gt;
  
  
  The get-out-of-jail card
&lt;/h2&gt;

&lt;p&gt;Anyone doing physical access or social engineering work in the field should be carrying a physical card, not a digital one, that lets them verify the engagement on the spot if security or police stop them.&lt;/p&gt;

&lt;p&gt;Requirements that actually matter here: a 24/7-reachable number, not an office line that goes to voicemail at 11pm; an authentication code the contact can verify against; the named contact is the executive sponsor, never someone on the security team; every field operator carries one; and it's a printed copy, because a phone can be confiscated, dead, or locked during exactly the moment you need it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the rules of engagement have to nail down
&lt;/h2&gt;

&lt;p&gt;The authorization letter grants permission. The rules of engagement define the edges of that permission, and for a red team both documents matter equally: objectives (detection, response, containment, physical security, identity, or full lifecycle), in-scope and out-of-scope targets down to specific systems and facilities, permitted and prohibited techniques, the exact testing window, how discovered sensitive data gets handled, whether persistence is allowed and when it must be removed, the detection protocol once the blue team notices, stop conditions, emergency contacts, the law-enforcement verification procedure, third-party boundaries, evidence preservation, and cleanup requirements at close.&lt;/p&gt;

&lt;h2&gt;
  
  
  What happens when the blue team actually catches you
&lt;/h2&gt;

&lt;p&gt;This is the part standard pentest authorization letters never have to address, and it's where most of the real risk lives.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Blue team detects but doesn't escalate to law enforcement.&lt;/strong&gt; Define in advance whether the red team keeps operating as if undetected, or stops. Most engagements specify continuing unless the executive sponsor says otherwise. This should never be improvised live.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Blue team escalates to the CISO.&lt;/strong&gt; If the CISO is in the "doesn't know" group, they might treat this as a real incident and call the police. The authorization needs a defined procedure for the executive sponsor to step in before that happens.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Blue team calls law enforcement.&lt;/strong&gt; This is the scenario that produced the well-known Coalfire arrests. Officers show up, testers get detained. The get-out-of-jail card is the first line of resolution. The 24/7 executive sponsor is the second, and the letter needs to state plainly that they're reachable to confirm the engagement directly to law enforcement, not to a lawyer, not by email, directly.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Blue team successfully contains the attack.&lt;/strong&gt; Decide ahead of time whether the engagement resets, ends, or triggers an immediate debrief, so nobody's making that call under pressure in the moment.&lt;/p&gt;

&lt;p&gt;At minimum, pre-define stop conditions for: production availability materially affected, customer data placed at risk, an unintended third party affected, law enforcement involvement, the blue team activating emergency response, the red team hitting a prohibited system, or the executive sponsor simply ordering a stop.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a red team letter needs that a standard one doesn't
&lt;/h2&gt;

&lt;p&gt;A compartmentalization section naming exactly who knows the engagement is happening and who must not be told, by name, not by role. A written confirmation from the executive sponsor that they're reachable for the full engagement window, treated as a primary operational requirement rather than a footnote. A defined blue-team response protocol so nobody improvises under pressure. A documented decision on whether local law enforcement was pre-notified, which is genuinely the single most effective way to avoid an arrest scenario. Precise physical scope, "the company's offices" is not scope, list addresses, floors, access types, and whether tailgating or badge cloning are in bounds. Verified cloud-provider authorization per account. An explicit reference to the signed rules of engagement so there's no ambiguity about what's actually authorized. And a named person with clear stop/terminate authority, plus the conditions under which the team must halt without waiting for sign-off.&lt;/p&gt;

&lt;h2&gt;
  
  
  Permission under secrecy is a different problem than permission
&lt;/h2&gt;

&lt;p&gt;A standard pentest is about getting permission. A red team is about getting permission while deliberately keeping most of the organization unaware it exists, and that constraint changes the shape of every document involved. Authorization can't be broad, can't be assumed, and can't reuse the pentest template with a confidentiality clause added. The moment the wrong people find out, the test is worthless. The moment the right people haven't formally signed off, the engagement is a real legal liability for everyone in the field.&lt;/p&gt;

&lt;p&gt;Limited awareness, explicit authority, and a clear escalation path, defined before testing starts, is the whole model. Get it right and the engagement runs exactly as designed: realistic, controlled, defensible. Get it wrong and you either lose the test or put your own team at risk. There's no middle version of this.&lt;/p&gt;

&lt;p&gt;For the baseline structure this builds on, see &lt;a href="https://www.codeant.ai/blog/penetration-test-authorization-letter" rel="noopener noreferrer"&gt;CodeAnt AI's guide to penetration test authorization letters&lt;/a&gt;. For how these engagements get scoped and reported end to end, see &lt;a href="https://www.codeant.ai/" rel="noopener noreferrer"&gt;CodeAnt AI's penetration testing platform&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>security</category>
      <category>informationsecurity</category>
      <category>redteam</category>
      <category>pentesting</category>
    </item>
    <item>
      <title>SOC 2 Penetration Testing: What Auditors Actually Check For (And Where Most Programs Fail)</title>
      <dc:creator>Amartya Jha</dc:creator>
      <pubDate>Mon, 07 Sep 2026 07:21:44 +0000</pubDate>
      <link>https://dev.to/codeant/soc-2-penetration-testing-what-auditors-actually-check-for-and-where-most-programs-fail-3neh</link>
      <guid>https://dev.to/codeant/soc-2-penetration-testing-what-auditors-actually-check-for-and-where-most-programs-fail-3neh</guid>
      <description>&lt;p&gt;Day three of a SOC 2 Type II audit. The auditor pulls up the penetration testing section of the audit program and asks: "Can you show me evidence that exploitable vulnerabilities identified during your penetration test were corrected, and that the corrections were tested?"&lt;/p&gt;

&lt;p&gt;Not "did you do a pentest." Not "can I see the report." The question is specifically about correction and verification.&lt;/p&gt;

&lt;p&gt;The engineering lead can produce the pentest report without much trouble. A retest report that confirms each finding was actually patched and verified in production is harder to find. A documented timeline from finding to fix to verification, harder still. Evidence that the retest ran against the same environment, same methodology, as the original test, almost nobody has that on hand.&lt;/p&gt;

&lt;p&gt;The control gets marked as an exception. The report ships with a qualified opinion on CC7.1. The enterprise sales cycle gets a little longer.&lt;/p&gt;

&lt;p&gt;This plays out constantly, not because companies skip security testing, most don't, but because they run testing without understanding what SOC 2 specifically needs that testing to produce as evidence.&lt;/p&gt;

&lt;h2&gt;
  
  
  SOC 2 is an opinion, not a certificate
&lt;/h2&gt;

&lt;p&gt;SOC 2 (Service Organization Control 2) is an AICPA auditing standard. It evaluates whether your controls meet the Trust Services Criteria relevant to your service. It is not a certification body handing out badges, it's an attestation: an independent CPA firm's formal opinion that your controls meet the criteria. That opinion is the report you hand to customers.&lt;/p&gt;

&lt;p&gt;Five TSC categories exist. Security (CC) is mandatory. The other four are optional add-ons depending on what you handle:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Category&lt;/th&gt;
&lt;th&gt;Required?&lt;/th&gt;
&lt;th&gt;Pentest relevance&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Security (CC)&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Direct — CC6, CC7, CC8, CC9&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Availability (A)&lt;/td&gt;
&lt;td&gt;Optional&lt;/td&gt;
&lt;td&gt;Infra resilience, DDoS&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Processing Integrity (PI)&lt;/td&gt;
&lt;td&gt;Optional&lt;/td&gt;
&lt;td&gt;Business logic, data validation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Confidentiality (C)&lt;/td&gt;
&lt;td&gt;Optional&lt;/td&gt;
&lt;td&gt;Data access, encryption&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Privacy (P)&lt;/td&gt;
&lt;td&gt;Optional&lt;/td&gt;
&lt;td&gt;PII handling, access verification&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Here's the part that trips people up: SOC 2 barely uses the phrase "penetration testing" anywhere in its text. Companies search for an explicit line item, don't find one, and conclude testing is optional. It isn't. The requirement is implicit in CC6.6 (protection against external threats) and CC7.1 (detection of security events), and AICPA's supplemental guidance explicitly names penetration testing as an appropriate way to satisfy them. Every Big Four firm and most regional CPA firms with a SOC 2 practice treat annual pentesting as table stakes for the Security category.&lt;/p&gt;

&lt;p&gt;The real question was never whether you need a pentest. It's whether the one you ran generates the specific evidence an auditor is trained to look for.&lt;/p&gt;

&lt;h2&gt;
  
  
  Type I vs Type II: different questions entirely
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Type I&lt;/th&gt;
&lt;th&gt;Type II&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Core question&lt;/td&gt;
&lt;td&gt;Are controls suitably designed?&lt;/td&gt;
&lt;td&gt;Did controls operate effectively over time?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;For pentest&lt;/td&gt;
&lt;td&gt;Do you have a program, and is it reasonably designed?&lt;/td&gt;
&lt;td&gt;Did the program actually run and produce results people acted on?&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Type I&lt;/strong&gt; checks that a program exists and looks reasonable. You need: a written, dated, approved pentest policy; one completed test report inside the observation window; evidence the tester was qualified and independent; a finding list with severities; a remediation plan (doesn't need to be finished, just documented). Type I auditors will not ask for a retest report, proof that remediation worked, or evidence testing covered everything critical. They're checking existence and design, not execution.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Type II&lt;/strong&gt; checks execution across the whole period, and the evidence bar jumps considerably:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Pentest report(s) covering the observation period&lt;/li&gt;
&lt;li&gt;Evidence critical/high findings were actually remediated&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A retest report confirming remediation worked&lt;/strong&gt; — the single most commonly missing item&lt;/li&gt;
&lt;li&gt;Timeline evidence: find date, fix date, retest date&lt;/li&gt;
&lt;li&gt;Remediation SLA documentation and proof you hit it&lt;/li&gt;
&lt;li&gt;Evidence the test covered systems that touch customer data&lt;/li&gt;
&lt;li&gt;Risk acceptance documentation for anything left unremediated&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Common failure patterns at Type II, roughly in order of frequency: a pentest report exists but no retest report does; the retest didn't cover every original finding; remediation happened in staging, not production (production is what's audited, staging evidence doesn't count); remediation took longer than the stated SLA, which is worse than having no SLA; risk acceptances existed only verbally; the test ran against an older version than what's in production; the gap between the last test and the audit exceeds twelve months.&lt;/p&gt;

&lt;h2&gt;
  
  
  Timing rules nobody reads until it's too late
&lt;/h2&gt;

&lt;p&gt;SOC 2 Type II has a minimum six-month observation period per AICPA guidance, though in practice enterprise buyers expect twelve. Your testing needs to produce evidence across that whole window, not one point inside it.&lt;/p&gt;

&lt;p&gt;Three structural options, in order of evidence strength:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Cadence&lt;/th&gt;
&lt;th&gt;Coverage&lt;/th&gt;
&lt;th&gt;Best for&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;A&lt;/td&gt;
&lt;td&gt;Annual (1 test + 1 retest)&lt;/td&gt;
&lt;td&gt;System as of test date&lt;/td&gt;
&lt;td&gt;Stable, low-change systems&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;B&lt;/td&gt;
&lt;td&gt;Semi-annual (2 tests + 2 retests)&lt;/td&gt;
&lt;td&gt;Two snapshots across the period&lt;/td&gt;
&lt;td&gt;Most SaaS teams&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;C&lt;/td&gt;
&lt;td&gt;Continuous (monthly)&lt;/td&gt;
&lt;td&gt;Full period&lt;/td&gt;
&lt;td&gt;High-velocity development&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Rules that catch teams off guard:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The test must occur &lt;em&gt;inside&lt;/em&gt; the observation period. A test run two months before your period started evidences the prior period, not this one.&lt;/li&gt;
&lt;li&gt;For a twelve-month period, at least one test needs to land inside those twelve months.&lt;/li&gt;
&lt;li&gt;The retest must also happen inside the period, after remediation.&lt;/li&gt;
&lt;li&gt;The gap between your last test and period-end shouldn't exceed roughly six months. Beyond that, expect the auditor to ask about it.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The single most common mistake: running the pentest in Q1 for a Q1–Q4 audit period and never scheduling a follow-up. The auditor eventually asks what covered Q3 and Q4. The honest answer, nothing, doesn't cause an automatic qualified opinion by itself, but it's a gap you'll have to explain, and it's rarely a satisfying explanation. Fix: schedule a second test, or at minimum a targeted retest, around the six-month mark.&lt;/p&gt;

&lt;h2&gt;
  
  
  Scope: follow the data, not the architecture diagram
&lt;/h2&gt;

&lt;p&gt;The scoping principle that matters more than any other: test everything that touches, stores, processes, or transmits the data your SOC 2 covers. Testing systems that don't touch customer data while skipping the ones that do is a scoping failure auditors catch quickly.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Component&lt;/th&gt;
&lt;th&gt;In scope?&lt;/th&gt;
&lt;th&gt;Controls&lt;/th&gt;
&lt;th&gt;Notes&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Public web app&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;CC6.6&lt;/td&gt;
&lt;td&gt;Primary requirement&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Authenticated API&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;CC6.1, CC6.3&lt;/td&gt;
&lt;td&gt;Document endpoint coverage&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Admin panel&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;CC6.1, CC6.3&lt;/td&gt;
&lt;td&gt;High-value target&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Customer data store&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;CC6.1, C1.1&lt;/td&gt;
&lt;td&gt;Often tested via the app layer&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cloud IAM&lt;/td&gt;
&lt;td&gt;Yes (if cloud-hosted)&lt;/td&gt;
&lt;td&gt;CC6.6, CC8.1&lt;/td&gt;
&lt;td&gt;IAM privilege audit&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CI/CD pipeline&lt;/td&gt;
&lt;td&gt;Recommended&lt;/td&gt;
&lt;td&gt;CC8.1&lt;/td&gt;
&lt;td&gt;High supply-chain risk&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Mobile apps&lt;/td&gt;
&lt;td&gt;Recommended&lt;/td&gt;
&lt;td&gt;CC6.6&lt;/td&gt;
&lt;td&gt;If customer-facing&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Staging/QA&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;Test production&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Third-party SaaS tools&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;CC9.2&lt;/td&gt;
&lt;td&gt;Review their SOC 2 instead&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two scope gaps auditors flag constantly: cloud infrastructure ("AWS manages that" is not a scope rationale, IAM misconfigurations and exposed buckets are CC6.1/CC6.6 findings and belong explicitly in scope) and APIs (most breaches happen at the API layer, but plenty of testing programs only cover what renders in a browser).&lt;/p&gt;

&lt;h2&gt;
  
  
  The retest report is the thing nobody has
&lt;/h2&gt;

&lt;p&gt;If there's one document to obsess over, it's this one. Fixing a bug isn't the same as evidencing the fix worked. The retest needs to come from the same external firm that ran the original test, re-executing the original test cases against the patched, production system, and producing a second report confirming which findings actually closed. Internal QA can supplement this, it can't replace it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What auditors actually ask, verbatim-adjacent
&lt;/h2&gt;

&lt;p&gt;These aren't hypothetical. They're the standard interview pattern:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;"Do you have a penetration testing policy?" — wants a written, approved, dated document, not "yes, we do pentests"&lt;/li&gt;
&lt;li&gt;"When was your last test?" — wants a date inside the observation period&lt;/li&gt;
&lt;li&gt;"Who conducted it?" — wants a named external firm, not "our developer ran some scans"&lt;/li&gt;
&lt;li&gt;"How do you know remediations worked?" — wants a retest report, not self-attestation&lt;/li&gt;
&lt;li&gt;"Show me a critical finding remediated within your SLA" — wants discovery, fix, and retest dates all inside the policy timeline&lt;/li&gt;
&lt;li&gt;"What about findings you didn't fix?" — wants a signed risk acceptance with a named owner and compensating controls&lt;/li&gt;
&lt;li&gt;"Was the test conducted against production?" — many teams test staging to avoid disruption; auditors know this trick and ask directly&lt;/li&gt;
&lt;li&gt;"Did the test cover cloud infrastructure?" — app testing and cloud config review are not the same exercise&lt;/li&gt;
&lt;li&gt;"What version was running when the test occurred?" — a test of v1.2 doesn't evidence v2.8; put the version number in the authorization letter&lt;/li&gt;
&lt;li&gt;"Are you getting better year over year?" — wants CVSS severity trend data across periods; a flat or worsening trend is a flag in renewal audits&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The pattern behind every good answer is the same: name the document, cite the date, name the firm, show the timeline. If you can't do that in one sentence, that's the gap.&lt;/p&gt;

&lt;p&gt;Nearly every exception traces back to one of three root causes: no retest report, timing (test ran outside the observation window or was never followed up), or scope (staging instead of production, no cloud layer, browser-only coverage while the API sat untested).&lt;/p&gt;

&lt;h2&gt;
  
  
  Building the program by stage
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;First Type I audit.&lt;/strong&gt; Start four months out, not four weeks. Month one: write and approve the policy, pick a qualified external vendor, scope every system touching customer data, run the test. Month two: triage findings with owners, remediate all criticals before the audit (highs strongly recommended), document risk acceptances for anything left, book the retest. Month three: get the retest report, assemble the evidence package, get management sign-off on findings (minuted), hand it to the auditor. Budget for a sub-50-endpoint SaaS: roughly $8K–$20K for the test, $2K–$6K for the retest, $1K–$3K for policy support if starting cold. The most common failure at this stage is simply starting too late; four weeks isn't enough runway for test, remediation, and retest.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Growth stage, Type II with active shipping.&lt;/strong&gt; A single annual test doesn't cover code shipped in month nine of a twelve-month period. Reasonable structure: one comprehensive annual test as the baseline, quarterly targeted testing of high-change areas, a security review on any auth/authz change before release, and continuous DAST/SAST/SCA scanning filling the gaps. The most common gap here is months seven through twelve with zero testing evidence, plus new features (a new payment flow, new auth method) shipped after the annual test with no review at all, which is a direct CC8.1 problem.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Enterprise, Type II with MSA obligations.&lt;/strong&gt; Baseline SOC 2 requirements plus whatever your customer contracts separately demand, and these are two different obligation sets, don't conflate them. Enhanced testing typically means external perimeter, internal network, API, and cloud reviewed together rather than separately, security testing embedded in the SDLC instead of bolted on annually, and a red team engagement every two to three years. Some enterprise customers contractually require quarterly testing, full report access instead of an executive summary, specific tester certifications (CREST, for instance), or a 24–48 hour critical-finding notification window. None of that is SOC 2, it's contract, and it needs its own evidence trail separate from the standard package.&lt;/p&gt;

&lt;h2&gt;
  
  
  Choosing a provider without guessing
&lt;/h2&gt;

&lt;p&gt;SOC 2 doesn't name specific certifications. It requires a "qualified" party, and your auditor applies judgment to decide whether that bar was met, so you need a defensible case for your choice.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Non-negotiable floor:&lt;/strong&gt; external to your org (internal or affiliated testers don't satisfy independence, full stop), a formal written report (a spreadsheet of findings is not audit evidence), and a formal retest included in scope, not "we'll review your remediation notes."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Signals worth paying for:&lt;/strong&gt; CREST organizational certification, OSCP or CISSP on the individual testers, prior SOC 2-specific engagement experience (they should already know what "per-finding verification" means to an auditor, you shouldn't be the engagement where they learn it).&lt;/p&gt;

&lt;p&gt;Questions worth asking before signing: can you show a sanitized sample report with unique finding IDs, CVSS scores, and proof-of-concept evidence (not scanner output with a paragraph slapped on)? Is retest included and does it produce its own separate report? Will you confirm in writing you're testing production? What's your process when you find a critical mid-engagement (there should be an out-of-band escalation, not "wait for the final report")? Will you provide a CVSS delta against last year's findings for renewal purposes?&lt;/p&gt;

&lt;p&gt;Red flags that end the conversation: automated scanner output passed off as a formal report, no retest service, no unique finding IDs, an inability to map findings to specific TSC control numbers, a default to staging, no professional liability insurance.&lt;/p&gt;

&lt;h2&gt;
  
  
  The twelve mistakes that actually cause exceptions
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Process failures, before the test runs:&lt;/strong&gt; testing staging instead of production; skipping the retest or letting an internal team do it; leaving critical findings open at audit time (start 90 days out, not 30); running the test before the observation period even opens; switching testing firms every year, which kills your ability to show a CVSS trend.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Documentation failures, after the test:&lt;/strong&gt; finding triage that happened but was never logged in a ticketing system; management review of the security posture that happened in someone's head instead of in minutes (CC5.3 needs the paper trail); risk acceptances that were a verbal "yeah, we'll live with that" instead of a signed document with a named owner; no sign-off that the scope actually covered customer-data systems.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Policy failures, the ones that look like organizational dysfunction:&lt;/strong&gt; an SLA policy that says 48 hours for criticals next to a remediation log showing 30 days, which is worse evidence than having no SLA at all; a scope document that excludes cloud infrastructure with "AWS handles that" as the justification.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a well-run engagement actually produces
&lt;/h2&gt;

&lt;p&gt;A pentest built for SOC 2 evidence, not just technical findings, structures itself around the control mapping from the start: reconnaissance covering every external asset for CC6.6; source-level tracing of auth and authz flows for CC6.1/CC6.3; JS bundle analysis for leaked secrets and exposed endpoints under CC6.7; active exploitation with every finding confirmed working before it's reported, chained findings reported as chains (three mediums that combine into one critical shouldn't ship as three tickets your team deprioritizes); a report where every finding carries proof-of-exploit, root cause down to file and line, CVSS 4.0 with metric justification, and a mapping to the specific control ID, not a general "this relates to SOC 2." Then unlimited retesting at no extra cost, with the retest confirming remediation in production specifically, and a named researcher sign-off, which is what satisfies the "qualified party" language in your own policy.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.codeant.ai/" rel="noopener noreferrer"&gt;CodeAnt AI's penetration testing platform&lt;/a&gt; is built around exactly this evidence structure, which is also what makes running the semi-annual or continuous testing cadence operationally realistic instead of a once-a-year scramble.&lt;/p&gt;

&lt;h2&gt;
  
  
  The actual takeaway
&lt;/h2&gt;

&lt;p&gt;Nobody fails a SOC 2 pentest control because their engineers can't find vulnerabilities. They fail it because the retest report doesn't exist, the test ran outside the observation window, or the scope quietly excluded the systems that actually touch customer data. All three are process problems, not technical ones, and all three are fixable months before an auditor ever sits down across the table.&lt;/p&gt;

&lt;p&gt;For a deeper breakdown of testing cadence tradeoffs referenced above, see &lt;a href="https://www.codeant.ai/blog/continuous-vs-annual-penetration-testing" rel="noopener noreferrer"&gt;continuous vs. annual penetration testing&lt;/a&gt;. For what a retest engagement should actually contain, see &lt;a href="https://www.codeant.ai/blog/pentest-retest" rel="noopener noreferrer"&gt;what happens during a pentest retest&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>security</category>
      <category>compliance</category>
      <category>cybersecurity</category>
      <category>soc2</category>
    </item>
    <item>
      <title>The Deployment Velocity Gap: Why Annual Pentesting Can't Keep Up With Modern SaaS</title>
      <dc:creator>Amartya Jha</dc:creator>
      <pubDate>Mon, 07 Sep 2026 06:03:28 +0000</pubDate>
      <link>https://dev.to/codeant/the-deployment-velocity-gap-why-annual-pentesting-cant-keep-up-with-modern-saas-d51</link>
      <guid>https://dev.to/codeant/the-deployment-velocity-gap-why-annual-pentesting-cant-keep-up-with-modern-saas-d51</guid>
      <description>&lt;p&gt;Most SaaS teams still pentest once a year. Almost none of them ship code once a year.&lt;/p&gt;

&lt;p&gt;That mismatch is the actual security problem worth talking about, more than any single vulnerability class. Security gets validated at one cadence. Risk gets introduced at another. The distance between those two cadences has a name: the deployment velocity gap.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the gap actually is
&lt;/h2&gt;

&lt;p&gt;The deployment velocity gap is the time between a security-relevant change hitting production and that change actually being evaluated by a penetration test.&lt;/p&gt;

&lt;p&gt;In an annual model, that gap can run for months. A system might be genuinely secure the day the assessment wraps, but every deploy after that day creates new, untested surface area: new endpoints, an updated auth flow, a new third-party integration, a reconfigured piece of infrastructure. None of it was in scope for the test that already happened. All of it is in scope for the breach that hasn't happened yet.&lt;/p&gt;

&lt;p&gt;This isn't a knock on pentesting as a discipline. It's a scoping problem. A penetration test evaluates a moment in time, and it evaluates a defined scope, both of which shrink in relative value as deployment frequency rises.&lt;/p&gt;

&lt;p&gt;You can put a number on your own exposure with something this simple:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;deployment_velocity_gap&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;test_interval_days&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;weekly_deployments&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;
    test_interval_days: 365 = annual, 90 = quarterly, 30 = monthly
    weekly_deployments: how many times code ships per week
    &lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="n"&gt;average_gap_days&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;test_interval_days&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;
    &lt;span class="n"&gt;total_deployments&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;test_interval_days&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;7&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;weekly_deployments&lt;/span&gt;

    &lt;span class="n"&gt;risk_level&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;CRITICAL&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;average_gap_days&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;90&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;HIGH&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;     &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;average_gap_days&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;45&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;MEDIUM&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;   &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;average_gap_days&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;14&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;LOW&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;average_exposure_window_days&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;average_gap_days&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;max_exposure_window_days&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;test_interval_days&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;untested_deployments_per_cycle&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;total_deployments&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;risk_level&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;risk_level&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Run it against a team shipping three times a week on an annual test cadence and you get a 182.5-day average exposure window and 156 untested deployments per cycle: CRITICAL. Drop the interval to monthly and the same team lands at a 15-day average window and roughly 13 untested deployments: MEDIUM. The application didn't get safer. The window just got shorter.&lt;/p&gt;

&lt;h2&gt;
  
  
  Continuous pentesting isn't "scan on every commit"
&lt;/h2&gt;

&lt;p&gt;Worth being precise here, because the term gets diluted fast: continuous penetration testing does not mean pointing an automated scanner at the app on every push and calling it a pentest. Scanners find known patterns at volume. A pentest investigates how a weakness can actually be chained and exploited in the context of a real, running application, with defined scope, authorization, controlled exploitation, and human oversight over the findings.&lt;/p&gt;

&lt;p&gt;What continuous testing does mean in practice:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Monthly testing&lt;/strong&gt; — a recurring assessment at a materially shorter interval than annual.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sprint-cadence testing&lt;/strong&gt; — assessments aligned to the development cycle itself.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Targeted testing&lt;/strong&gt; — auth changes, new APIs, payment flows, and infra changes trigger a focused assessment on top of the recurring one.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Integrated coverage&lt;/strong&gt; — pentesting sits alongside IDE, PR, and CI/CD controls rather than replacing them.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;There's no single "correct" frequency. The right cadence is a function of deployment velocity, data sensitivity, regulatory obligation, and how fast the team can actually act on a finding once it exists.&lt;/p&gt;

&lt;h2&gt;
  
  
  Scoping a sprint by risk, not by calendar
&lt;/h2&gt;

&lt;p&gt;The more useful version of "test every sprint" is test every sprint &lt;em&gt;proportionally to what changed&lt;/em&gt;. A simple risk-scoring pass over merged PRs gets you most of the way there:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;SprintSecurityTestingProgram&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;classify_test_depth&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;risk_score&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;risk_score&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;100&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;FULL_DEPTH — auth changes require complete auth chain review&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="k"&gt;elif&lt;/span&gt; &lt;span class="n"&gt;risk_score&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;50&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;TARGETED_DEEP — multiple security-relevant changes&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="k"&gt;elif&lt;/span&gt; &lt;span class="n"&gt;risk_score&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;20&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;TARGETED_STANDARD — specific components need focused testing&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="k"&gt;else&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;LIGHTWEIGHT — automated testing sufficient&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;identify_priority_areas&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;changed_components&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;priorities&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;changed_components&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;authentication&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
            &lt;span class="n"&gt;priorities&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;area&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Authentication&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;priority&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;test_focus&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;JWT validation, session management, MFA bypass, brute force&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="p"&gt;})&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;changed_components&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;authorization&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
            &lt;span class="n"&gt;priorities&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;area&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Authorization&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;priority&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;test_focus&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;RBAC, IDOR, cross-tenant access, role bypass&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="p"&gt;})&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;changed_components&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;data_access&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
            &lt;span class="n"&gt;priorities&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;area&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Data Access Layer&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;priority&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;test_focus&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;SQL/NoSQL injection, ownership filter presence&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="p"&gt;})&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;sorted&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;priorities&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="k"&gt;lambda&lt;/span&gt; &lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;priority&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Auth and authz changes get full-depth review every time. Everything else gets scoped by what it actually touches. That's the difference between a program that scales and one that turns into indiscriminate noise.&lt;/p&gt;

&lt;h2&gt;
  
  
  The economics people usually get wrong
&lt;/h2&gt;

&lt;p&gt;The lazy comparison is "one annual invoice" versus "a subscription." That comparison undercounts the annual model and overcounts the continuous one, because it ignores retesting, remediation effort, false-positive waste, emergency response, and the actual expected cost of a breach sitting inside a long exposure window.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Cost category&lt;/th&gt;
&lt;th&gt;Annual model&lt;/th&gt;
&lt;th&gt;Continuous model&lt;/th&gt;
&lt;th&gt;Delta&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Direct testing cost&lt;/td&gt;
&lt;td&gt;$25K–$50K&lt;/td&gt;
&lt;td&gt;$48K–$72K/yr (sub)&lt;/td&gt;
&lt;td&gt;+$10K–$25K&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Retest cost&lt;/td&gt;
&lt;td&gt;$8K–$15K&lt;/td&gt;
&lt;td&gt;Included&lt;/td&gt;
&lt;td&gt;-$12K&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Engineering remediation&lt;/td&gt;
&lt;td&gt;$19.2K (8 findings × 3 days)&lt;/td&gt;
&lt;td&gt;$14.4K (12 findings × 1.5 days)&lt;/td&gt;
&lt;td&gt;-$4.8K&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;False positive waste&lt;/td&gt;
&lt;td&gt;$9.6K (30% FP rate)&lt;/td&gt;
&lt;td&gt;$1.6K (5% FP rate)&lt;/td&gt;
&lt;td&gt;-$8K&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Emergency response&lt;/td&gt;
&lt;td&gt;$17.5K (35% probability)&lt;/td&gt;
&lt;td&gt;$4K (8% probability)&lt;/td&gt;
&lt;td&gt;-$13.5K&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Expected breach cost&lt;/td&gt;
&lt;td&gt;$24.7K (180-day window)&lt;/td&gt;
&lt;td&gt;$1.6K (14-day window)&lt;/td&gt;
&lt;td&gt;-$23K&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Total TCO&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;~$104K&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;~$80K&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;-$24K&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Illustrative numbers for a $10M ARR company shipping weekly, 15 engineers at $100/hr, 12% annual breach probability, $500K average breach cost, but the shape of the comparison holds more broadly: the direct testing line goes up, almost everything downstream of it goes down, and the breach-cost line is usually the biggest single swing because it scales directly with how long the exposure window is.&lt;/p&gt;

&lt;p&gt;The expected-breach-cost math is worth writing out, because it's the part most TCO comparisons skip:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;exposure_window_days_continuous&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;14&lt;/span&gt;  &lt;span class="c1"&gt;# sprint cadence
&lt;/span&gt;&lt;span class="n"&gt;adjusted_breach_probability&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;breach_probability_annual&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;exposure_window_days_continuous&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;365&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;expected_breach_cost&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;adjusted_breach_probability&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;avg_breach_cost&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Shrink the exposure window, shrink the probability term, shrink the expected cost. It's linear and it's easy to underweight if you're only looking at the invoice.&lt;/p&gt;

&lt;h2&gt;
  
  
  Picking a cadence: a decision framework, not a rule
&lt;/h2&gt;

&lt;p&gt;Four questions, roughly in order of how much weight they should carry:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How often do you deploy?&lt;/strong&gt; Weekly or more and annual testing is covering under 10% of your deployments by construction. Monthly or less and annual/semi-annual is probably fine.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What data do you handle?&lt;/strong&gt; PII at scale, payment data, or health data pushes you toward quarterly-minimum or continuous regardless of deploy cadence, because breach impact and regulatory exposure are both high.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What's the regulatory environment?&lt;/strong&gt; PCI DSS, SOC 2, HIPAA, and ISO 27001 all have their own expectations for testing evidence and cadence. Continuous testing can supplement that evidence; it doesn't automatically substitute for a compliance-mandated assessment, and that distinction matters when an auditor is asking.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can you actually act on findings?&lt;/strong&gt; A one-person security team that can triage in real time can run continuous. A team with no dedicated security function and no plan for continuous will just accumulate an ignored backlog, which is worse than not having found the issues at all.&lt;/p&gt;

&lt;p&gt;Rough mapping:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Deployment frequency&lt;/th&gt;
&lt;th&gt;Recommended model&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Less than monthly&lt;/td&gt;
&lt;td&gt;Annual / semi-annual&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Monthly to bi-weekly&lt;/td&gt;
&lt;td&gt;Quarterly minimum&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Weekly or more&lt;/td&gt;
&lt;td&gt;Sprint-based or continuous&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  SLAs are what keep a continuous program from collapsing
&lt;/h2&gt;

&lt;p&gt;Continuous testing produces continuous findings, and without a severity-tiered SLA, that turns into noise the engineering team eventually starts ignoring.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Severity&lt;/th&gt;
&lt;th&gt;CVSS&lt;/th&gt;
&lt;th&gt;Acknowledge&lt;/th&gt;
&lt;th&gt;Remediate&lt;/th&gt;
&lt;th&gt;Retest&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Critical&lt;/td&gt;
&lt;td&gt;9.0–10.0&lt;/td&gt;
&lt;td&gt;4h&lt;/td&gt;
&lt;td&gt;48h&lt;/td&gt;
&lt;td&gt;within 24h of fix&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;td&gt;7.0–8.9&lt;/td&gt;
&lt;td&gt;24h&lt;/td&gt;
&lt;td&gt;7d&lt;/td&gt;
&lt;td&gt;within 48h of fix&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Medium&lt;/td&gt;
&lt;td&gt;4.0–6.9&lt;/td&gt;
&lt;td&gt;72h&lt;/td&gt;
&lt;td&gt;30d&lt;/td&gt;
&lt;td&gt;within sprint&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Low&lt;/td&gt;
&lt;td&gt;0.1–3.9&lt;/td&gt;
&lt;td&gt;1 week&lt;/td&gt;
&lt;td&gt;90d&lt;/td&gt;
&lt;td&gt;next quarterly&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Why continuous programs actually die
&lt;/h2&gt;

&lt;p&gt;Most don't fail because the testing was bad. They fail for one of four operational reasons:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Finding fatigue without triage.&lt;/strong&gt; Everything gets flagged with the same urgency, engineers tune it out. Fix: CVSS-based SLAs, a security champion per team doing first-pass triage, a weekly standup instead of ad-hoc tickets.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Testing is time-based instead of change-based.&lt;/strong&gt; A calendar cadence misses whatever shipped between checkpoints. Fix: scope each cycle off actual PR/change data, not the calendar (this is what the sprint risk-scoring class above is for).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Surface monitoring without ownership.&lt;/strong&gt; New subdomains and endpoints get flagged and nobody owns the follow-up. Fix: a rotation that owns the alert queue, with an explicit SLA (72h is reasonable) for investigating new surface.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Compliance-minimum thinking.&lt;/strong&gt; The program quietly reverts to "just enough for the auditor." Fix: report on breach-probability and incident-cost terms at the board level, not just pass/fail testing status, so the program's ROI stays visible.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Metrics that actually tell you if it's working
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Annual baseline&lt;/th&gt;
&lt;th&gt;Continuous target&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Mean time to detection&lt;/td&gt;
&lt;td&gt;~180 days&lt;/td&gt;
&lt;td&gt;&amp;lt;14 days&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Mean time to remediation&lt;/td&gt;
&lt;td&gt;~45 days (batch quarterly)&lt;/td&gt;
&lt;td&gt;&amp;lt;7 days&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Vulnerability escape rate&lt;/td&gt;
&lt;td&gt;~25% (found by someone else first)&lt;/td&gt;
&lt;td&gt;&amp;lt;5%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SLA compliance rate&lt;/td&gt;
&lt;td&gt;~60%&lt;/td&gt;
&lt;td&gt;&amp;gt;95%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;False positive rate&lt;/td&gt;
&lt;td&gt;~40%&lt;/td&gt;
&lt;td&gt;&amp;lt;10%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Attack surface coverage&lt;/td&gt;
&lt;td&gt;~70%&lt;/td&gt;
&lt;td&gt;&amp;gt;95%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;None of these should be optimized in isolation. The point of tracking them together is to see whether the gap between "change ships" and "change gets validated" is actually shrinking over time, not just whether a test happened.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where AI actually changes the equation
&lt;/h2&gt;

&lt;p&gt;The reason traditional pentest firms can't support a monthly cadence isn't unwillingness, it's structural: a human consultant works a target sequentially over one to two weeks, and that timeline doesn't compress just because you want it to.&lt;/p&gt;

&lt;p&gt;What changes the math is running large numbers of specialized exploit agents in parallel against a target instead of one consultant working through it linearly, and carrying codebase context forward between engagements instead of starting cold every time, which is closer to how an attacker who's already been inside your codebase (via source access, leaked credentials, or a prior foothold) would actually operate. That combination, source-aware testing plus accumulated context, is what makes a 48-hour full-engagement turnaround and unlimited free retesting operationally sustainable instead of a pricing gimmick.&lt;/p&gt;

&lt;p&gt;If you want to see how that model works end to end: &lt;a href="https://www.codeant.ai/" rel="noopener noreferrer"&gt;CodeAnt AI's continuous penetration testing platform&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The actual takeaway
&lt;/h2&gt;

&lt;p&gt;Annual pentesting isn't wrong. It's a point-in-time answer to a problem that, for most SaaS teams now, is continuous. The question worth asking isn't "did we pass our last pentest," it's "how long does a newly introduced risk sit untested before anyone looks at it." If that number is measured in months, the testing cadence and the deployment cadence have drifted apart, and that drift is the actual attack surface.&lt;/p&gt;

&lt;p&gt;For the cost-side deep dive behind the TCO numbers above, see &lt;a href="https://www.codeant.ai/blog/penetration-testing-cost" rel="noopener noreferrer"&gt;CodeAnt's guide to penetration testing costs&lt;/a&gt;. For the operational checklist on choosing a recurring testing vendor, see &lt;a href="https://www.codeant.ai/blog/ptaas-sla" rel="noopener noreferrer"&gt;PTaaS provider SLAs: what to look for&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>security</category>
      <category>appsec</category>
      <category>devsecops</category>
      <category>pentesting</category>
    </item>
    <item>
      <title>Dissecting FOIS: The Log4j2 Filter That Only Does Half Its Job</title>
      <dc:creator>Amartya Jha</dc:creator>
      <pubDate>Tue, 01 Sep 2026 06:56:17 +0000</pubDate>
      <link>https://dev.to/codeant-security/dissecting-fois-the-log4j2-filter-that-only-does-half-its-job-5c0a</link>
      <guid>https://dev.to/codeant-security/dissecting-fois-the-log4j2-filter-that-only-does-half-its-job-5c0a</guid>
      <description>&lt;h2&gt;
  
  
  __Log4j2's FilteredObjectInputStream checks which classes can be rebuilt, but not how large or deep the data can be. Here's how that gap enables an RCE and two gadget-free crashes on a serialized log receiver.
&lt;/h2&gt;

&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Log4j2's &lt;code&gt;FilteredObjectInputStream&lt;/code&gt; (FOIS) checks the class names in a serialized stream and nothing else. It never limits object size or depth.&lt;/li&gt;
&lt;li&gt;That one gap gives three problems on a single network log receiver: an unauthenticated RCE, a crash that pins a CPU core, and a 44-byte out-of-memory crash. Two of them need no exploit gadget.&lt;/li&gt;
&lt;li&gt;The RCE (Log4j2 #4255) is prior work. CodeAnt's contribution is dissecting the two crashes, the contrast with logback, and the lab numbers.&lt;/li&gt;
&lt;li&gt;Exposure is narrow. You need a serialized log receiver on the network, which is not a Log4j default. High severity, low prevalence.&lt;/li&gt;
&lt;li&gt;Fix: send logs as text, or add the size and depth caps logback already ships.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  At a Glance
&lt;/h2&gt;

&lt;p&gt;A security filter is meant to do one job completely. FOIS does about half of one. It reads the name of every class arriving over the network and rejects anything not on an approved list.&lt;/p&gt;

&lt;p&gt;That is a real check, but it is the only check. FOIS never asks how large an object is or how deeply nested, and it never turns on the size and depth limits Java already ships. Those two missing checks turn one log receiver into three problems.&lt;/p&gt;

&lt;p&gt;The RCE below (Log4j2 #4255) is not our discovery; it is prior work, credited in full at the end.&lt;/p&gt;

&lt;p&gt;What our team adds is the breakdown of the filter itself: the two crashes hiding in the limits FOIS forgot to set, the comparison with the library that got it right, and numbers from a real lab.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://codeant.ai/security-research/dissecting-log4j2-filtered-object-input-stream" rel="noopener noreferrer"&gt;Read the full CodeAnt security research&lt;/a&gt; →&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Lab:&lt;/strong&gt; Log4j 2.26.1 on JDK 17. Run it yourself with one Docker command.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  What FOIS Was Built To Do
&lt;/h2&gt;

&lt;p&gt;Log4j can ship log events between machines as Java objects, and the receiver rebuilds them from raw bytes. Rebuilding attacker-controlled bytes is the flaw behind a decade of Java RCE, so Apache wrapped that step in FOIS in Log4j 2.8.2 (the fix for CVE-2017-5645). FOIS keeps the class names it trusts and throws out the rest.&lt;/p&gt;

&lt;p&gt;Two details of that trusted list matter. It allows a class called &lt;code&gt;java.rmi.MarshalledObject&lt;/code&gt;, and it trusts whole Java packages by prefix, so everything under &lt;code&gt;java.util&lt;/code&gt; and &lt;code&gt;java.lang&lt;/code&gt; is waved through. Each one becomes a way in.&lt;/p&gt;

&lt;h2&gt;
  
  
  The One Mistake
&lt;/h2&gt;

&lt;p&gt;FOIS checks which class is being rebuilt. It never checks how big or how deep. Java can cap array size, nesting depth, and reference count; FOIS turns none of that on. Everything below follows from that.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Three Impacts
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fypegx20yyfe5i4awwcx2.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fypegx20yyfe5i4awwcx2.gif" alt=" " width="660" height="440"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;RCE&lt;/strong&gt; — a &lt;code&gt;java.rmi.MarshalledObject&lt;/code&gt;-wrapped gadget hits the FilteredObjectInputStream receiver; the receiver logs &lt;code&gt;msg=null&lt;/code&gt; while &lt;code&gt;whoami&lt;/code&gt; returns &lt;code&gt;root&lt;/code&gt;, unauthenticated code execution as &lt;code&gt;root&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5g8ubs2wvb0cqpbdmrsl.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5g8ubs2wvb0cqpbdmrsl.gif" alt=" " width="660" height="440"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;DoS&lt;/strong&gt; — the same receiver, no gadget: a 5.7 KB nested &lt;code&gt;HashSet&lt;/code&gt; pins &lt;code&gt;readObject&lt;/code&gt; at 100% CPU permanently, and a 44-byte &lt;code&gt;Object&lt;/code&gt; array triggers &lt;code&gt;OutOfMemoryError: Requested array size exceeds VM limit&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;One: remote code execution, smuggled inside an allowed class.&lt;/strong&gt; &lt;code&gt;MarshalledObject&lt;/code&gt; is trusted, but it is a container: it carries a second batch of bytes inside itself, and when Log4j opens it, those inner bytes are rebuilt on a fresh stream that FOIS is not watching.&lt;/p&gt;

&lt;p&gt;The outer check passes; the real payload rides in behind it. Log4j opens the container on its own during reconstruction, so if the server has a usable gadget library on its classpath, this is unauthenticated RCE.&lt;/p&gt;

&lt;p&gt;In the lab against real Log4j 2.26.1 on JDK 17, a wrapped payload ran &lt;code&gt;whoami&lt;/code&gt; and returned &lt;code&gt;root&lt;/code&gt;, the server stayed up and logged a blank message, and the same payload sent without the container was correctly rejected. (On a modern JDK fewer gadgets survive, but the ones that run commands still work where the right library is present, so the requirement is just a usable gadget on the classpath.)&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Two: a CPU bomb the size of a text message.&lt;/strong&gt; &lt;code&gt;HashSet&lt;/code&gt; is allowed, and rebuilding a set recomputes the hash of everything inside it.&lt;/p&gt;

&lt;p&gt;A 2015 trick called SerialDOS (Wouter Coekaerts) builds a small nested structure where each layer is shared and referenced twice, so a shape a few dozen levels deep is reachable by an exponential number of paths, and the runtime walks every one. With no depth limit, the receiver runs it to completion:&lt;/p&gt;

&lt;p&gt;A few kilobytes freezes the machine, one worker per packet, and it is silent: while the receiver spins there is no log line. The only sign is a thread stuck at 100% CPU.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Three: a 44-byte packet that crashes the process.&lt;/strong&gt; A serialized array states its length before any data, and the runtime reserves that space immediately. With no size limit, that number can be anything.&lt;/p&gt;

&lt;p&gt;A 44-byte packet can claim an array of billions of entries, and the process dies with an out-of-memory error before it even tries to allocate. No gadget, no library, just a lie about size.&lt;/p&gt;

&lt;p&gt;The full research walks through the FOIS implementation, reproduces all three impact paths, and includes the lab setup and measurements.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://codeant.ai/security-research/dissecting-log4j2-filtered-object-input-stream" rel="noopener noreferrer"&gt;Read the complete analysis&lt;/a&gt; →&lt;/p&gt;

&lt;h2&gt;
  
  
  How logback Solved the Same Problem
&lt;/h2&gt;

&lt;p&gt;logback, Spring Boot's default logger, had the identical socket receiver, and its guard shows what the finished job looks like. &lt;code&gt;HardenedObjectInputStream&lt;/code&gt; does everything FOIS does, then the part FOIS skipped: an exact-name allowlist instead of trusting whole packages, a maximum array size and nesting depth (which kills both crashes), and a closed proxy path.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Facopdlcz3ptizt75x9t6.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Facopdlcz3ptizt75x9t6.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;It is immune to all three. The difference between the two libraries is essentially one extra line of filter, the one that turns on the size and depth caps, which FOIS never wrote.&lt;/p&gt;

&lt;h2&gt;
  
  
  How Exposed Is This, Really?
&lt;/h2&gt;

&lt;p&gt;Less than three impacts suggests, and this is the honest part. Every one needs a serialized log receiver exposed on the network, which is not how Log4j runs by default; Apache removed the built-in socket server years ago. When one researcher tested 53 product builds in default configuration, none were exploitable.&lt;/p&gt;

&lt;p&gt;Call it high severity, low prevalence: trivial to hit where a receiver exists, uncommon to have one. Where they do appear, it tends to be logback's own socket receiver, an older Elastic Logstash log4j input, or legacy Log4j 1.x, which has no filter at all.&lt;/p&gt;

&lt;h2&gt;
  
  
  How To Fix It
&lt;/h2&gt;

&lt;p&gt;If you cannot remove a serialized receiver right away, one runtime setting closes the gaps. It denies the gadget package and adds the size and depth caps at once. The catch: only one such filter applies per process and a later setting silently replaces an earlier one, so every rule lives in a single string.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nt"&gt;-Djdk&lt;/span&gt;.serialFilter&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'!org.apache.commons.collections.**;maxarray=10000;maxdepth=16'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is a stopgap. The durable fix is to stop sending Java objects across a trust boundary: carry log events as JSON or plain text, or remove the serialized receiver. Log4j 3.x already dropped this pattern; the 2.x line still in production keeps it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Are You Affected?
&lt;/h2&gt;

&lt;p&gt;Only if you run a serialized log receiver reachable on the network, which is a deliberate setup, not a default. Look for a process listening on a log-ingest port (historically 4560) that rebuilds log events off the socket.&lt;/p&gt;

&lt;p&gt;If you find one, treat it as high severity, apply the filter above, then move that transport to text or retire it. On Log4j 3.x the pattern is already gone.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Original research:&lt;/strong&gt; U-Sec / Wujie Security, first credited with the &lt;code&gt;MarshalledObject&lt;/code&gt; allowlist bypass behind #4255. The original report was later deleted.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Independent PoCs:&lt;/strong&gt; joanbono's &lt;code&gt;log4j2-4255-exploit&lt;/code&gt; and dinosn's &lt;code&gt;log4j-4255&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;CPU-crash technique:&lt;/strong&gt; Wouter Coekaerts, "SerialDOS," 2015.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Prior writeup:&lt;/strong&gt; Jeff McJunkin's post, whose 53-product test and detection guidance inform the exposure and fix sections.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Credit: Prior Work
&lt;/h2&gt;

&lt;p&gt;The RCE (issue #4255), the property that any gadget works once the inner stream is unchecked, and the gadgets themselves are prior work. Our contribution is the breakdown of the filter around that RCE, the two crashes from its missing limits, the logback comparison, and the lab numbers.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Takeaway
&lt;/h2&gt;

&lt;p&gt;A class-name allowlist is not a deserialization boundary. Restricting which objects can be rebuilt may reduce the risk of code execution, but without limits on payload size, object depth, or resource consumption, the same endpoint can remain vulnerable to denial of service.&lt;/p&gt;

&lt;p&gt;The bigger lesson is simple: &lt;strong&gt;a security control is only as strong as the attack surface it actually closes.&lt;/strong&gt; An endpoint can look hardened against one exploit path while still being wide open to another.&lt;/p&gt;

&lt;p&gt;If your application relies on deserialization, APIs, or other security-sensitive boundaries, automated checks alone may not expose these gaps. That is where adversarial testing matters.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Want to know what your application exposes before an attacker does?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;CodeAnt AI helps teams identify security weaknesses across their codebase and infrastructure. For deeper coverage, &lt;strong&gt;CodeAnt's pentesting services&lt;/strong&gt; put those controls to the test from an attacker's perspective, helping uncover exploitable paths that static analysis and automated scanners can miss.&lt;/p&gt;

&lt;p&gt;If you want to see the kind of security research behind that approach, &lt;a href="https://codeant.ai/security-research/dissecting-log4j2-filtered-object-input-stream" rel="noopener noreferrer"&gt;read the full Log4j2 FOIS analysis&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Find the weakness before someone else does.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Apache Log4j2 issue #4255&lt;/li&gt;
&lt;li&gt;logback HardenedObjectInputStream&lt;/li&gt;
&lt;li&gt;ysoserial · log4j2-4255-lab · Jeff McJunkin&lt;/li&gt;
&lt;li&gt;CVEs: CVE-2017-5645, CVE-2017-5929, CVE-2019-17571, CVE-2023-6378&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;This breakdown is part of an ongoing effort by CodeAnt AI Security Research to analyse the trust boundaries in widely deployed infrastructure. The lab ran in isolated local Docker containers with synthetic payloads. The underlying issue and the gadgets are prior work, credited above.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Have a receiver like this, or want a deserialization boundary reviewed? Reach us at &lt;a href="mailto:securityresearch@codeant.ai"&gt;securityresearch@codeant.ai&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>cybersecurity</category>
      <category>security</category>
      <category>pentesting</category>
      <category>vulnerabilities</category>
    </item>
  </channel>
</rss>
