<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Aditya Soni</title>
    <description>The latest articles on DEV Community by Aditya Soni (@aditya_soni_e5b9d5213e544).</description>
    <link>https://dev.to/aditya_soni_e5b9d5213e544</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4121446%2F48a2f631-4998-469f-ae25-17c72e2d71d3.png</url>
      <title>DEV Community: Aditya Soni</title>
      <link>https://dev.to/aditya_soni_e5b9d5213e544</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/aditya_soni_e5b9d5213e544"/>
    <language>en</language>
    <item>
      <title>A Prompt Injection Turned Into a Shell: Inside Semantic Kernel's Two RCE CVEs</title>
      <dc:creator>Aditya Soni</dc:creator>
      <pubDate>Thu, 17 Sep 2026 14:36:35 +0000</pubDate>
      <link>https://dev.to/aditya_soni_e5b9d5213e544/a-prompt-injection-turned-into-a-shell-inside-semantic-kernels-two-rce-cves-22l7</link>
      <guid>https://dev.to/aditya_soni_e5b9d5213e544/a-prompt-injection-turned-into-a-shell-inside-semantic-kernels-two-rce-cves-22l7</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F56a5u0pbk6vjn7p29ndt.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F56a5u0pbk6vjn7p29ndt.png" alt=" " width="800" height="447"&gt;&lt;/a&gt;&lt;em&gt;Two ordinary agent-framework conveniences, a filterable vector search and a file-download tool, turned into remote code execution once an LLM's output was trusted a little too much. Here's what happened in Microsoft Semantic Kernel, and what to check in your own agent code today.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Prompt injection usually gets discussed as an output-quality problem: the model says something it shouldn't, or does something a user didn't ask for. CVE-2026-26030 and CVE-2026-25592, two vulnerabilities Microsoft disclosed in its Semantic Kernel framework, are a sharper reminder of what it can become: full remote code execution, demonstrated in Microsoft's own writeup by launching &lt;code&gt;calc.exe&lt;/code&gt; on the machine running the agent.&lt;/p&gt;

&lt;p&gt;Semantic Kernel is Microsoft's open-source framework for building agents and wiring LLMs into tool-calling applications, widely used both inside Microsoft's own ecosystem and in third-party projects. Both vulnerabilities follow the same underlying pattern in different SDKs, and that pattern is the actually useful thing to take away from this.&lt;/p&gt;

&lt;h2&gt;
  
  
  CVE-2026-26030: when a search filter is also a code path
&lt;/h2&gt;

&lt;p&gt;This one lives in the Python SDK's &lt;code&gt;InMemoryVectorStore&lt;/code&gt;. When an agent searches a vector store with a metadata filter, Semantic Kernel builds that filter as a Python lambda expression and evaluates it with &lt;code&gt;eval()&lt;/code&gt; at query time. The filter parameters can include AI-model-controlled input, and according to Microsoft's account, that input was not sanitized before being interpolated into the expression.&lt;/p&gt;

&lt;p&gt;The framework did have a blocklist meant to catch dangerous constructs. It was bypassed using Python's class hierarchy traversal, essentially reaching dangerous built-ins through an indirect attribute-access path the blocklist didn't anticipate rather than referencing them directly. Microsoft's proof of concept, built around a "hotel finder" agent scenario, used this path to launch &lt;code&gt;calc.exe&lt;/code&gt;. The fix shipped in &lt;code&gt;semantic-kernel&lt;/code&gt; 1.39.4 for Python.&lt;/p&gt;

&lt;h2&gt;
  
  
  CVE-2026-25592: a file-download tool with no path validation
&lt;/h2&gt;

&lt;p&gt;The second vulnerability is in the .NET SDK's &lt;code&gt;SessionsPythonPlugin&lt;/code&gt;. Its &lt;code&gt;DownloadFileAsync&lt;/code&gt; function was exposed to the model through a &lt;code&gt;[KernelFunction]&lt;/code&gt; attribute, the standard way Semantic Kernel marks a method as callable by the agent, without validating the destination path. That let a sufficiently crafted prompt cause the agent to write a file to an arbitrary location on the host, including the Windows Startup folder.&lt;/p&gt;

&lt;p&gt;A file dropped in Startup runs automatically the next time the user logs in. That's the more consequential part of this one: it doesn't just execute code inside the agent's process, it persists past the session and past whatever sandbox the agent's normal execution was running in, because a container's sandbox isolation typically doesn't extend to filesystem locations mounted or shared with the host, and Startup-folder execution triggers outside the container entirely on the next host login. The fix shipped in the .NET SDK 1.71.0, the same day as the Python fix.&lt;/p&gt;

&lt;h2&gt;
  
  
  The pattern underneath both
&lt;/h2&gt;

&lt;p&gt;Strip away the specifics and both vulnerabilities are the same mistake: a powerful, code-adjacent primitive, &lt;code&gt;eval()&lt;/code&gt; in one case, unrestricted file I/O in the other, received a value that ultimately traced back to model output, and nobody treated that value as attacker-controlled. That's the actual lesson, and it generalizes far past these two specific CVEs.&lt;/p&gt;

&lt;p&gt;An LLM's output is attacker-controlled the moment it has processed any untrusted content: a scraped webpage, a user-supplied document, retrieved search results, another agent's message. If that output can influence a filter string that gets evaluated as code, or a path that gets written to without validation, or really any string that reaches a powerful primitive without going through a narrow, explicit contract, you have the shape of this bug, regardless of which framework you're using.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to actually check in your own code
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Audit every method your framework marks as callable by the model&lt;/strong&gt; (&lt;code&gt;[KernelFunction]&lt;/code&gt; in Semantic Kernel, the equivalent decorator or registration call in whatever you're using). For each one, ask: if the model called this with the worst possible string, what happens?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Never&lt;/strong&gt; &lt;strong&gt;&lt;code&gt;eval()&lt;/code&gt;&lt;/strong&gt; &lt;strong&gt;or&lt;/strong&gt; &lt;strong&gt;&lt;code&gt;exec()&lt;/code&gt;&lt;/strong&gt; &lt;strong&gt;a value built from model output, even for something that feels as harmless as a search filter.&lt;/strong&gt; Build a real filter DSL, a small, explicit grammar you parse and validate, rather than reaching for a code-executing shortcut because it's faster to ship.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Validate file paths against a canonicalized allowlist of directories, not a blocklist of bad patterns.&lt;/strong&gt; The vector-store bug's blocklist bypass is a reminder that blocklists lose to creative attackers almost by definition; allowlists at least bound the failure.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Don't treat container or sandbox isolation as your whole defense.&lt;/strong&gt; CVE-2026-25592 shows a process that looks properly contained can still reach outside it through a filesystem write that was never meant to be a privilege boundary in the first place.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;If you're on Semantic Kernel, upgrade now&lt;/strong&gt;: Python to 1.39.4 or later, .NET to 1.71.0 or later, and grep your own codebase for any &lt;code&gt;[KernelFunction]&lt;/code&gt;-tagged method that touches the filesystem, shells out, or evaluates a string as code.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Why this keeps happening
&lt;/h2&gt;

&lt;p&gt;Agent frameworks compete partly on how much they let a model do without hand-written glue code, which means each new release tends to add more direct, convenient bridges between "thing the model said" and "thing that actually executes." Every one of those bridges is a candidate for this exact failure mode. Semantic Kernel isn't unusually careless here; it's unusually well-documented, because Microsoft wrote the postmortem itself. The realistic assumption for anyone building on any agent framework is that equivalent, undisclosed instances of this pattern exist elsewhere, and the fix is the same regardless of framework: treat model output as untrusted input everywhere it lands, not just at the final response to the user.&lt;/p&gt;

&lt;h3&gt;
  
  
  Sources
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://www.microsoft.com/en-us/security/blog/2026/05/07/prompts-become-shells-rce-vulnerabilities-ai-agent-frameworks/" rel="noopener noreferrer"&gt;https://www.microsoft.com/en-us/security/blog/2026/05/07/prompts-become-shells-rce-vulnerabilities-ai-agent-frameworks/&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.pointguardai.com/ai-security-incidents/semantic-kernel-lets-a-prompt-open-a-shell-cve-2026-25592-cve-2026-26030" rel="noopener noreferrer"&gt;https://www.pointguardai.com/ai-security-incidents/semantic-kernel-lets-a-prompt-open-a-shell-cve-2026-25592-cve-2026-26030&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://nvd.nist.gov/vuln/detail/CVE-2026-25592" rel="noopener noreferrer"&gt;https://nvd.nist.gov/vuln/detail/CVE-2026-25592&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;em&gt;This story was written with the assistance of an AI writing program.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>aisecurity</category>
      <category>agentframeworks</category>
    </item>
    <item>
      <title>DeepSeek's New Model Isn't Interesting Because of Benchmarks. It's the Pricing.</title>
      <dc:creator>Aditya Soni</dc:creator>
      <pubDate>Sat, 12 Sep 2026 19:16:22 +0000</pubDate>
      <link>https://dev.to/aditya_soni_e5b9d5213e544/deepseeks-new-model-isnt-interesting-because-of-benchmarks-its-the-pricing-31i4</link>
      <guid>https://dev.to/aditya_soni_e5b9d5213e544/deepseeks-new-model-isnt-interesting-because-of-benchmarks-its-the-pricing-31i4</guid>
      <description>&lt;p&gt;Another open-weight model beating a frontier benchmark isn't news anymore. A pricing structure that makes long-context work ten times cheaper at certain hours might actually change how you architect a product.&lt;br&gt;
What DeepSeek shipped&lt;br&gt;
DeepSeek released V4.1 Flash on September 10, 2026, according to the company's own API changelog and coverage from VentureBeat, Dataconomy, and DataCamp. It's a mixture-of-experts model with a reported 552-billion-parameter backbone, native vision support, and a 1-million-token context window, released under an MIT license with weights available on Hugging Face for commercial use.&lt;/p&gt;

&lt;p&gt;On published benchmarks, DeepSeek reports 90.6 on Terminal-Bench 2.1 (up from 82.7 on the prior V4 Flash release) and 74.2% on DeepSWE v1.1 (up from 54.4%), alongside a 3,471 Codeforces rating and 90.9 on GPQA Diamond. VentureBeat's coverage frames these as "eclipsing" GPT-5.6 Sol and Claude Opus 5 on select metrics, though it's worth being precise: beating a frontier closed model on a specific benchmark subset is not the same as being a better general-purpose model, and DeepSeek's own comparison points are self-selected.&lt;/p&gt;

&lt;p&gt;The more unusual detail is the pricing. During off-peak hours, DeepSeek prices input tokens at $0.003 per million on a cache hit, versus $0.15 per million on a cache miss, with output at $0.60 per million. Peak-hour rates roughly double. DeepSeek has also said its existing V4 Pro tier will route to V4.1 Flash after September 14, 2026, effectively retiring the older tier in favor of this one.&lt;/p&gt;

&lt;p&gt;Why the pricing structure is the actual engineering story&lt;br&gt;
A 50x gap between cache-hit and cache-miss pricing is a strong, explicit signal about where DeepSeek's infrastructure cost actually lives: recomputing attention over long contexts, not generating tokens. That's not a new insight in principle — every major provider has some form of prompt caching — but pricing it this aggressively, and pairing it with a genuine 1M-token window, changes the calculus for a specific class of application: agents and tools that repeatedly re-read a large, mostly-static context (a codebase, a knowledge base, a long conversation history) and only append a small amount of new information per turn.&lt;/p&gt;

&lt;p&gt;If your workload looks like "re-send a 200K-token codebase snapshot on every turn, with the same snapshot for the next hour," this pricing model rewards you specifically for keeping that context stable and hitting the cache, in a way that a flat per-token price doesn't. It's an argument for restructuring how you manage context in an agent loop: batch static context in a way that maximizes cache hits, and treat cache-miss traffic as the expensive path to minimize, not just an occasional cost.&lt;/p&gt;

&lt;p&gt;The tradeoffs nobody puts in the launch post&lt;br&gt;
Off-peak pricing implies a queuing/scheduling decision you now have to make. If a meaningful chunk of the cost advantage is time-of-day dependent, batch and non-interactive workloads (evals, bulk document processing, offline agent runs) can be scheduled to capture it; latency-sensitive user-facing traffic generally can't. That's a real architectural split, not just a footnote.&lt;/p&gt;

&lt;p&gt;Benchmark gains on Terminal-Bench and DeepSWE are coding-agent-specific. These are useful signals if you're building a coding assistant, less so if your use case is retrieval-heavy Q&amp;amp;A or general reasoning, where the model's relative position may differ. Don't generalize a coding-benchmark win into "best model for our RAG pipeline" without testing on your own eval set.&lt;/p&gt;

&lt;p&gt;Open weights change your deployment options, not your operational burden. MIT-licensed weights mean you can self-host V4.1 Flash instead of using DeepSeek's API, which matters for data residency or cost-at-scale reasons — but self-hosting a 552B-parameter MoE model competently is its own significant infrastructure project, not a free lunch.&lt;/p&gt;

&lt;p&gt;Practical takeaway&lt;br&gt;
Before switching a workload to chase this pricing, model your actual cache-hit rate under realistic usage patterns, not the best case. A cache-friendly workload (stable context, frequent short follow-ups) can see genuinely large savings; a cache-unfriendly one (constantly changing context) will mostly pay the cache-miss rate and see far less benefit than the headline number suggests. Run the numbers on your own traffic shape before committing.&lt;/p&gt;

&lt;p&gt;Sources&lt;br&gt;
&lt;a href="https://venturebeat.com/technology/deepseek-v4-1-flash-debuts-with-0-003-1m-off-peak-cached-input-rate-and-benchmarks-eclipsing-gpt-5-6-sol-claude-opus-5" rel="noopener noreferrer"&gt;https://venturebeat.com/technology/deepseek-v4-1-flash-debuts-with-0-003-1m-off-peak-cached-input-rate-and-benchmarks-eclipsing-gpt-5-6-sol-claude-opus-5&lt;/a&gt;&lt;br&gt;
&lt;a href="https://dataconomy.com/2026/09/11/deepseek-v4-1-flash-ultralow-token-pricing/" rel="noopener noreferrer"&gt;https://dataconomy.com/2026/09/11/deepseek-v4-1-flash-ultralow-token-pricing/&lt;/a&gt;&lt;br&gt;
&lt;a href="https://api-docs.deepseek.com/updates/" rel="noopener noreferrer"&gt;https://api-docs.deepseek.com/updates/&lt;/a&gt;&lt;br&gt;
&lt;a href="https://www.datacamp.com/blog/deepseek-v4-1-flash" rel="noopener noreferrer"&gt;https://www.datacamp.com/blog/deepseek-v4-1-flash&lt;/a&gt;&lt;br&gt;
This story was written with the assistance of an AI writing program.&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0mptt7dd9id8k289w8r8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0mptt7dd9id8k289w8r8.png" alt=" " width="799" height="436"&gt;&lt;/a&gt;&lt;/p&gt;

</description>
    </item>
    <item>
      <title>An Attacker's Multi-Agent Framework Stole Thousands of Credentials in Under Six Hours</title>
      <dc:creator>Aditya Soni</dc:creator>
      <pubDate>Sat, 12 Sep 2026 09:21:16 +0000</pubDate>
      <link>https://dev.to/aditya_soni_e5b9d5213e544/an-attackers-multi-agent-framework-stole-thousands-of-credentials-in-under-six-hours-4kgh</link>
      <guid>https://dev.to/aditya_soni_e5b9d5213e544/an-attackers-multi-agent-framework-stole-thousands-of-credentials-in-under-six-hours-4kgh</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F491z1dxcn8snkbmhnw7p.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F491z1dxcn8snkbmhnw7p.png" alt=" " width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;A financially motivated attacker ran reconnaissance, exploitation, and cleanup with almost no human in the loop — using the same design patterns you'd use to build a helpful agent. That symmetry is the actual story.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What Google's threat researchers found
&lt;/h2&gt;

&lt;p&gt;Google's Threat Intelligence Group (GTIG) published its Q3 2026 AI Threat Tracker documenting a mass credential-harvesting campaign carried out largely by a multi-agent framework rather than a human operator. According to GTIG's report and coverage from The Hacker News, SiliconANGLE, and Help Net Security, a suspected financially motivated actor gained access to an organization's cloud infrastructure, then deployed an autonomous multi-agent system from inside that environment — which let requests originate from legitimate-looking IP addresses rather than obviously malicious infrastructure.&lt;/p&gt;

&lt;p&gt;The attacker reportedly built the framework using an AI coding chatbot, a set of instructions, and preconfigured Markdown files functioning as operational playbooks — essentially the same "agent skills" pattern legitimate teams use to give an LLM reusable, structured procedures. Per Mandiant's incident-response analysis cited in GTIG's report, the resulting system managed the vulnerability-scanning pipeline, collected credentials, resolved technical problems as they came up, and rotated IP addresses — without a human approving each step. Google says the entire operation, from setup to compromising thousands of third-party credentials, took under six hours.&lt;/p&gt;

&lt;h2&gt;
  
  
  The part that should worry engineers, specifically
&lt;/h2&gt;

&lt;p&gt;Most public discussion of "agentic security risk" focuses on prompt injection against your own agents, or an assistant being tricked into leaking your data. This is a different and arguably more mundane threat: the same reliability engineering that makes agents useful — self-correction, retries, graceful handling of transient failures — is exactly what makes an attack pipeline resilient enough to run unattended for hours.&lt;/p&gt;

&lt;p&gt;A scanning script that dies on the first unexpected HTTP response needs a human to restart it. An agent that "resolves technical problems as they arise" doesn't. That's the entire value proposition of agentic tooling, and it's now showing up on the offensive side at effectively the same maturity level as it shows up in developer tools. There is no special dangerous capability here beyond what a decent coding agent already has — which is precisely the point: the barrier to running a multi-day human-operated campaign as a six-hour autonomous one has dropped to "know how to write agent instructions."&lt;/p&gt;

&lt;h2&gt;
  
  
  What this changes for teams running agents in cloud environments
&lt;/h2&gt;

&lt;p&gt;If your organization runs agentic coding tools, CI-integrated assistants, or internal automation with cloud credentials in scope, this campaign is a preview of what compromise of &lt;em&gt;that&lt;/em&gt; infrastructure looks like from the attacker's side, not just the defender's. A few concrete implications:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Treat agent-accessible credentials as high-value targets, explicitly.&lt;/strong&gt; If an AI coding tool or CI agent has broad cloud IAM permissions "to be safe," that's now a more attractive target than it was a year ago, because compromising it hands an attacker a ready-made autonomous operator, not just a static credential.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;IP-based anomaly detection is weaker than it used to be.&lt;/strong&gt; Attacks launched from inside compromised legitimate infrastructure won't trip geographic or reputation-based alerts. Detection needs to shift more weight onto behavioral signals — unusual API call sequences, scanning patterns, credential-access velocity — that don't depend on the request's origin looking suspicious.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Markdown-based agent instructions are now a security-relevant artifact.&lt;/strong&gt; If your org uses "skills" files, playbooks, or reusable prompt templates for internal agents, those same file types deserve the code-review scrutiny you'd give a shell script, because that's functionally what they are.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What we don't know yet
&lt;/h2&gt;

&lt;p&gt;GTIG's report describes the campaign's structure and outcome but, per the available public reporting, doesn't name the specific AI coding chatbot the attacker used, nor does it publish the full Markdown playbooks (for obvious reasons). That means teams can't yet check their own environment against specific indicators of compromise beyond general behavioral patterns — this is a capability disclosure more than an incident-response playbook. Expect more operational detail to surface as other vendors and incident responders corroborate or extend the findings.&lt;/p&gt;

&lt;h3&gt;
  
  
  Sources
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://cloud.google.com/blog/topics/threat-intelligence/from-prompting-to-autonomy-the-evolution-of-adversarial-ai" rel="noopener noreferrer"&gt;https://cloud.google.com/blog/topics/threat-intelligence/from-prompting-to-autonomy-the-evolution-of-adversarial-ai&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://thehackernews.com/2026/09/autonomous-ai-agents-compromise.html" rel="noopener noreferrer"&gt;https://thehackernews.com/2026/09/autonomous-ai-agents-compromise.html&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://siliconangle.com/2026/09/08/google-says-attackers-used-ai-agents-to-steal-credentials-in-under-six-hours/" rel="noopener noreferrer"&gt;https://siliconangle.com/2026/09/08/google-says-attackers-used-ai-agents-to-steal-credentials-in-under-six-hours/&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.helpnetsecurity.com/2026/09/08/ai-agents-cyberattacks-automation-google-research/" rel="noopener noreferrer"&gt;https://www.helpnetsecurity.com/2026/09/08/ai-agents-cyberattacks-automation-google-research/&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;em&gt;This story was written with the assistance of an AI writing program.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>aisecurity</category>
      <category>agentreliability</category>
    </item>
  </channel>
</rss>
