The case
Cisco Talos has detected, using its CAIRN framework, a Windows malware sample that does not carry its instructions inside. It asks for them outside.
The thing queries DeepSeek, Qwen, Mistral and Gemini to decide what it does next, with nobody giving it orders from a command-and-control server.
It steals credentials and cryptocurrency, and the most striking part is not the theft, it is how it decides.
It is one of the first samples of malware orchestrated by several LLMs at once (Sources: marketingprofs.com, Talos).
The binary starts, looks at what it has in front of it and instead of running a fixed routine, it asks.
It asks one model for an opinion, another for a second one.
With the answers it builds the next step.
If the environment does not fit, it changes plan without anyone rewriting the binary.
That breaks the detection logic we have been using for years.
A signature looks for a known pattern, and here the pattern is not in the executable, it is in the conversation with four external APIs. And that conversation can be different on every victim.
Read that last sentence again, please, and let it sink in.
The pattern
The same week has left two more samples that the attacker no longer writes the whole kill-chain, it delegates it.
An autonomous AI agent compromised DIVD, the Dutch vulnerability disclosure center, chaining two Zammad zero-days (CVE-2026-102489 and CVE-2026-102490).
From hijacked session to root in seconds, with lateral movement and exfiltration.
This is already in CISA's KEV catalog, with a patch in 7.2.0 (Sources: cloudsecurityalliance.org, securityweek.com).
And Rejetto HFS came in through CVE-2026-61500, found by AI and already exploited according to VulnCheck on October 2, to recover the cookie signing key and achieve administrative RCE (Source: securityweek.com).
The attacker stops programming every step and builds a system that decides on the fly.
In CLOSEDQUORUM the decision is made by four models in parallel.
In Zammad it was made by an agent chaining exploits.
Summing up, in all three cases, the time between first access and impact is measured in seconds, not hours.
The other side
Google presented Gemini 4 Argon on September 30, its frontier model aimed at cybersecurity.
It is deployed only among trusted defenders through the Fairwind program and for that group, without cyber guardrails.
The introductory API goes for 2 and 10 dollars per million tokens (Source: thehackernews.com).
That is, the same kind of model the malware queries by API is also being sold as a defensive tool.
The difference is not in the model, it is in who calls it and with what permissions.
Executives from Anthropic, OpenAI, Google and Meta testified under oath before the New York City Council about the risks of AI, in the first joint appearance of this kind after the incidents with autonomous agents (Source: cnbc.com).
While that was happening in a room, CLOSEDQUORUM was already querying four models to decide what to do.
The reading
A binary that talks to four AI APIs leaves a very specific trace in outbound traffic, and that trace is easier to see than any signature.
Outbound traffic to AI APIs is usually allowed by default, because half the company uses it to work, and a process that should not be talking to the internet opening TLS connections to four different providers does not trigger anything.
This hurts. A SOC built on antivirus alerts does not see this but if it correlates processes with outbound connections, it does.
What to look at
- Outbound restricted by default. Review which processes in your fleet can reach the internet and to which destinations. If an unknown binary opens connections to AI APIs, that has to be an alert, not background noise.
- Outbound traffic to AI APIs. Keep an inventory of the legitimate domains your organization uses. Any connection to a model provider that is not on that list deserves a look.
- Segmentation before detection. If the binary cannot reach the internal network from where it runs, credential theft stays half done. Segmentation is not glamorous, but it stops this better than any rule.
How would I test it in my lab?
I would set up a clean machine with Sysmon and DNS capture and run a harmless sample that simulates the pattern.
What I want to see is whether I am able to detect the sequence without a signature.
A process that is not a browser opening TLS connections to several AI providers in a short window.
I would test it with my own script that makes three calls to different APIs and see what shows up in the logs.
If the pattern is visible in my lab, it is visible anywhere but if it is not, I already know where to start looking before you come and tell me.
Closing
Malware that asks four AIs for an opinion is not smarter than the previous one, it is just harder to sign, and that, in defense, changes what you have to look at.
Originally published at https://sammideblas.com/notas/malware-asks-four-ais-before-it-acts
Si has leído hasta aquí... reacciona y comparte -> "la seguridad y la defensa en la era de la IA es cosa de todos los que la usan" ¿No Crees?
Top comments (0)