Anthropic has published a new report about how people misused Claude between December and August.
The examples are unsettling. One operation used more than 4,700 dating-app personas to send over two million messages. Another copied an activist’s writing style to fool his contacts. Other cases involved malware, surveillance, and possible weapons work.
The important part is not just what Claude could do. It is how much repeated work the model handled for its operators.
Keeping thousands of fake identities alive takes time. Rebuilding malware after security tools catch it takes time too. AI can turn both jobs into something much easier to repeat.
The scary shift is from one bad message to an operation that keeps running.
Anthropic says it blocked accounts and disrupted the activity. But that does not always remove software that has already been built or deployed somewhere else.
This is why stronger models need more than a simple list of banned requests. They need monitoring, access controls, and people who can spot patterns across many small actions.
The hard question is where to draw the line. Lock things down too much and useful security work suffers. Leave them open and abuse scales faster.
AI safety is starting to look less like a lab problem and more like an operations problem.
Top comments (0)