DEV Community

Cover image for Your Next Pen Test Will Be Run by an Agent. Your Hospital Server Probably Will Not Be Ready.
Yano.AI Technologies Inc.
Yano.AI Technologies Inc.

Posted on Originally published at yanoai.tech

Your Next Pen Test Will Be Run by an Agent. Your Hospital Server Probably Will Not Be Ready.

Sixty-four percent of security leaders now formally assess the security of the AI tools their organizations deploy, up from 37 percent a year earlier (Source: World Economic Forum, 2026). That is the defensive side of the ledger moving fast. The offensive side moved faster: one research team catalogued 70 open-source AI penetration testing tools in 18 months (Source: Hadrian, 2026). A Cebu outsourcing firm with two security staff now faces an attacker who reads documentation, writes a probe, parses the result, and tries again without pausing.

Infographic

The asymmetry is not skill, it is patience

A human penetration tester works in shifts. They scope, probe, document, sleep, and resume. An agentic scanner holds the entire thread of an attack in working memory and never hands a finding to the next person on rotation (Source: ISACA, 2026). That is the operational difference that matters, not raw exploit sophistication.

The honest counterpoint: agents are still weak. On CVE-Bench, a benchmark of real-world critical web application vulnerabilities, the state-of-the-art agent framework exploited up to 13 percent of them (Source: CVE-Bench, 2025). Thirteen percent is not a toy number when an attacker points that scanner at your entire perimeter and lets it run.

Why the Philippines feels this first

The National Cybersecurity Plan 2.0 and the National Cybersecurity Inter-Agency Committee rest on a risk-based model for Critical Information Infrastructure (Source: Cyfirma, 2026). That framework assumes defenders can enumerate and prioritize. It does not assume they can absorb continuous machine-speed probing across every internet-facing asset.

The exposure surface here is unusually wide. Internet penetration sits at 83.8 percent, covering more than 98 million individuals, and digital payments account for 52.8 percent of transaction volume (Source: Cyfirma, 2026). Hospitals in Metro Manila and provincial government offices run the same internet-exposed portals as their counterparts anywhere, with a fraction of the security staffing. The Active Cyber Defence law enables inter-agency telemetry sharing and mandatory breach reporting, which helps after an intrusion. It does not tell a mid-sized firm in Cebu which of its 400 exposed assets an agent found yesterday.

Healthcare is under the heaviest pressure. It is the most targeted industry in the country, and medical records trade for 10 to 20 times the price of financial data on underground markets (Source: Cyfirma, 2026). Over 60 percent of healthcare breaches lead to operational disruption, and Qilin, which leads Southeast Asian ransomware activity with 48 tracked incidents, has hit Philippine hospitals alongside Medusa (Source: Cyfirma, 2026).

The AI you deployed is also an attack surface

Most boards in this region spent 2024 and 2025 asking whether to adopt AI. The security question has quietly inverted. Across more than 1,200 cloud environments, 81 percent of organizations running AI packages had at least one known vulnerability with an average CVSS score of 8.79, and 99.9 percent of AI vulnerabilities with an available fix remained unpatched (Source: Orca Security, 2026).

That does not mean AI is uniquely broken. It means the supply chain discipline applied to web applications did not travel with the new packages, and the fix rate is near zero. A payroll assistant that reads invoices and a chatbot on your public website now sit inside your trust boundary, and neither has a patch SLA.

The deeper exposure is architectural. When an agent can call your agent, the blast radius is defined by what the tool is permitted to do, not what the model is capable of doing (Source: Synack, 2026). That shift is now visible in the 2026 CyberSecurity Breakthrough Awards, which added a dedicated agentic AI security category (Source: Synack, 2026).

What a proportionate response looks like for a 40-person firm

Start with the perimeter inventory, not the vendor conversation. You cannot patch what you have not listed, and an agentic scanner finds forgotten assets faster than a consultant's spreadsheet (Source: Cyfirma, 2026). Pull every internet-facing hostname, API, and remote access endpoint into one list.

Then decide which assets carry the crown jewels, and treat that subset as needing agentic-grade testing (Source: Hadrian, 2026). You do not need continuous testing on 400 assets. You need it on the payment gateway, the HR system holding employee data, and the VPN into the network. Everything else can live on a quarterly cycle.

Third, bound what your AI tools can reach. Give the invoice assistant read access to one folder, not the shared drive, and rotate its credentials on the same schedule as a human account.

Fourth, budget for continuous testing as a line item, not a project. The global penetration testing market reached 2.72 billion US dollars in 2026 and is growing at 15.29 percent annually (Source: Mordor Intelligence, 2026). That growth is funded by buyers who stopped treating the annual pen test as sufficient.

The uncomfortable middle

There is a version of this story where AI levels the field and a small firm with a good managed provider competes with a bank. There is another where the gap widens, because the tools are cheap and distributed to everyone, including actors who only want to resell access.

Right now both are true. The 13 percent exploit rate on real-world CVEs means a determined human operator with an agent still beats the agent (Source: CVE-Bench, 2025). But the buyer of a stolen credential does not care. They will use the 13 percent, resell it, and move on. That is the argument for not waiting, and the 78.3 percent of dark web threats targeting the Philippines that are domestic-focused or state-linked is a reminder that patience is not a strategy anyone else is following (Source: Cyfirma, 2026).

FAQ

Is AI pentesting already better than a human tester?
No. On CVE-Bench, state-of-the-art agents exploited up to 13 percent of critical real-world web application CVEs (Source: CVE-Bench, 2025). They win on breadth, speed, and continuous coverage. A skilled human still outperforms them on depth and business-context reasoning.

Should a small business buy an AI pentesting tool?
Only after you have a clean asset inventory. If nobody knows what you expose to the internet, an automated scanner hands you 400 findings and no prioritization (Source: Hadrian, 2026).

How do I secure the AI tools I already deployed?
Inventory them, restrict each tool's permissions to the minimum data it needs, rotate credentials on a human schedule, and apply patches. Only 0.1 percent of fixable AI vulnerabilities get patched across the cloud estate (Source: Orca Security, 2026).

Why does this matter more in the Philippines?
Internet penetration of 83.8 percent and digital payments at 52.8 percent of transaction volume mean more systems exposed, while most organizations run lean security teams (Source: Cyfirma, 2026). Healthcare is the most targeted sector here, with medical records valued 10 to 20 times higher than financial data underground (Source: Cyfirma, 2026).

Is continuous testing worth the cost for a mid-sized firm?
For your highest-value assets, yes. The market grows at 15.29 percent annually because buyers stopped treating an annual test as sufficient (Source: Mordor Intelligence, 2026). For low-risk internal assets, a quarterly cycle is defensible.

Key Takeaway

The question is no longer whether an AI agent will test your defenses. It is whether you find the gaps it finds first, or read about them afterward in a breach notification from the National Privacy Commission.

Your perimeter inventory is probably wrong today. Would you rather fund one agentic scan of your highest-value assets this quarter, or explain to a hospital board, a bank auditor, or an LGU mayor why the finding was already public?

Sources

Top comments (0)