DEV Community

jamilxt
jamilxt

Posted on

She Used Claude as a Diary. A Human at Anthropic Read It and Called the Police. Now She Faces a Felony.

On September 26, a woman in Bonita Springs, Florida typed something into Claude. She said she was going to "shoot up" the Lee County Sheriff's Office. The next day, using the same account, she wrote that she had gotten a new gun.

Claude's automated safety systems flagged the messages. The flag went to a human review team at Anthropic. That person read the messages, decided the threat was credible, and reported it to law enforcement. Deputies identified Carli Michelle Heller, went to her home, and detained her without incident. She is now charged with making a written threat of violence under Florida Statute 836.10, which makes it a second-degree felony to transmit a written threat to kill or injure someone or carry out a mass shooting. Her court date is in November.

When investigators asked her about it, she said something that should stop you cold: she uses AI like a "diary."

That detail, more than the charge, is why this story hit 808 points and 649 comments on Hacker News in under two days. Everyone who has ever vented into a chatbot at 2 a.m. read that line and thought about their own message history.

Full disclosure before we go further: everything in this article comes from public records and published policies. I have no inside knowledge of what happens inside Anthropic's or OpenAI's review queues. I read the arrest report, both companies' terms and policies, and the reporting from WINK News, TechSpot, Ars Technica, and the Washington Post's coverage of the OpenAI case, and I have linked each claim to its source. What I found is not the story most people are arguing about.

Most of the argument is about the wrong question

The take you will see a hundred times today is "never tell a chatbot anything private." That is true, and it is useless advice, because people will keep doing it. The conversational format is designed to feel intimate. Ian Marlow, CEO of FITECH, put it plainly to WINK News: something that feels conversational does not immediately feel recorded, because people are just having a conversation.

The sharper question is a different one. Between the moment you press enter and the moment the model replies, your words pass through a pipeline. Somewhere in that pipeline there is a rule, written by a handful of people at the company, that decides when a stranger gets to read your message and when that stranger calls the police.

That rule is not public. And the two biggest AI companies just applied it in opposite directions, with consequences you can measure in courtroom dates.

What Anthropic's own documents say

This is not a leak or a whistleblower story. Anthropic publishes its position, and the position is broader than most users assume.

  • The Consumer Terms reserve a right, in the company's sole discretion, to report you. The terms, effective October 8, 2025, state: "We reserve the right, at our sole discretion, to report information from or about you, including but not limited to Inputs, Outputs, or Actions to law enforcement." Sole discretion means no warrant, no subpoena, no judge. That is not an accusation; it is the quoted text of the agreement every claude.ai user accepted.
  • The privacy policy allows disclosure on a good-faith belief. The policy, effective September 10, 2026, permits sharing personal data with authorities when the company believes in good faith that disclosure is reasonably necessary to prevent serious harm to any person or to property. Note the scope: the phrase covers property, not just injury to people.
  • Flagged conversations are not deleted on the normal schedule. Under Anthropic's retention practices, content flagged for trust and safety review is exempt from routine deletion, and the current policy keeps flagged chat text for up to two years and the classifier metadata around it for up to seven.
  • Human review is the exception, not the default, but the exception is broad. Anthropic's documentation says staff do not routinely read conversations and that human access happens through a controlled path when automated systems flag content. The arrest report confirms the path worked exactly as designed: safety measures monitor for key phrases and threatening content, severe cases are escalated to a human review team, and that team reported the statements to police.

Here is the detail that surprised me most. Anthropic's usage policy contains exactly one category where the company commits in writing to reporting users to authorities: child sexual abuse material and coercion of minors. For violent threats, there is no published trigger, no published threshold, and no commitment to tell the user that their conversation was escalated. The company wrote a hard rule for the category where the law forces its hand, and a judgment call for everything else. The judgment call is the one that just put a user in front of a felony judge.

And the transparency numbers do not fill the gap. Anthropic's government requests report for July through December 2025 lists 2 content requests, 20 non-content requests, 7 preservation requests, and 0 emergency requests. That counts inbound government demands. It does not count the outbound referrals the company initiates itself, like this one. There is no published figure for how often Anthropic's reviewers escalate a user to police, because the company does not track that category publicly.

Two companies, opposite outcomes, one unwritten rule

If Anthropic's reviewer had stayed quiet and something had happened, this article would not exist, because the story would be a massacre and a lawsuit. The referral in this case looks defensible on its facts: a named target, a government building, then a weapon the next day. I am not arguing the reviewer got it wrong.

But look at what the other lab did with a nearly identical situation, and the asymmetry becomes visible.

In June 2025, OpenAI's automated systems flagged the ChatGPT account of a 17-year-old in Tumbler Ridge, British Columbia for content involving gun violence. Human reviewers debated whether to contact the Royal Canadian Mounted Police. They decided the activity did not meet the company's threshold for a police referral, which at the time required what CBC News described as evidence of the target, the means, and the timing of a planned act of violence. They banned the account instead.

On February 10, 2026, that person killed eight people.

Afterward, Ann O'Leary, OpenAI's vice president for global policy, said in a letter to Canadian officials: "Under our enhanced law enforcement referral protocol, we would refer the account banned in June 2025 to law enforcement if it were discovered today." The threshold was lowered so it no longer requires a specific target, means, and timing. A pattern of detailed violent scenarios is now enough. Mental health experts were added to assess high-risk cases. The change applies in all jurisdictions, not just Canada.

On September 21, 2026, British Columbia sued OpenAI and Sam Altman in federal court in San Francisco over the decision not to refer. The complaint alleges that safety staff recommended contacting police and were overruled. Those are allegations, and OpenAI has not yet had its day in court on them.

Put the two cases side by side:

  • Anthropic referred a diary entry. A user faces a felony. Nobody sued the company.
  • OpenAI declined to refer a flagged account. Eight people died. The province is suing.

Any trust and safety lead at any AI company can do this math. The downside of over-reporting is a privacy debate on Hacker News. The downside of under-reporting is a lawsuit and a press conference. The incentive is to move the line toward calling. Which means the effective threshold at every lab is probably drifting in one direction, set privately, with no published standard and no audit.

I want to be careful here. Nothing I found shows that Anthropic's threshold has drifted or that this particular referral was wrong. That is exactly the problem: nothing public could show it either way.

What this means for you, practically

If you use a consumer chatbot, here is the honest model of your exposure, based on the published policies.

  • Assume a flag is possible on anything you type. Automated classifiers screen consumer conversations continuously. They are tuned to detect threats, CSAM, and other severe policy violations, and they do not understand context like "I was just venting."
  • A flag can put a human in your message history. At Anthropic, that review happens through a controlled access path limited to approved reviewers, and every access is logged. But "a trained professional reading my worst 2 a.m. thoughts" is still a stranger reading your diary.
  • Flagged content lives longer than you think. Normal deleted chats purge within about 30 days. Flagged content is exempt, and at Anthropic currently means up to two years for the text and seven for the classifier metadata.
  • Nobody is obligated to tell you. Neither company's published policy commits to notifying a user that a conversation was escalated.
  • If you build on the API, your users are in a different bucket. Anthropic's consumer reporting clause explicitly does not apply to content processed on behalf of business customers. Commercial terms contain no law-enforcement reporting sentence at all. What your app's users type is governed by your agreement with Anthropic and your own privacy policy. That is your legal surface to manage, not Claude.ai's.

And one rule of thumb worth saving: treat every cloud chatbot like a postcard, not a diary. A postcard is useful, fast, and any handler along the route can read it. If you would not write it on a postcard, use a local notes app, or a locally running model if you need to talk it through with an AI.

None of this means Anthropic did something wrong in this case. The referral probably was the right call, and I would rather live in a world where a credible threat against a sheriff's office gets a human to look at it than one where it does not. But "we did the right thing this time, trust us" is not a system. A system has a written standard, published numbers, and a defined process. Anthropic publishes its model's honesty scores to a decimal place and has never published how often it calls the sheriff.

The takeaway

The diary case is not really a privacy story. It is a governance story. A small number of employees at a private company are making case-by-case decisions about reporting users to law enforcement, under a standard that is written down nowhere the public can see, with an incentive gradient that points in one direction. The fix is cheap: state the threshold in operational terms, publish referral counts every six months the way companies already publish government request counts, and say whether a flagged user gets told. Any lab could do that tomorrow at zero cost to safety.

The next lab to publish its threshold gets a lasting trust advantage. The first to be forced to produce it in discovery gets a very bad week.


I write about AI, developer tools, and the systems behind them every week. Subscribe, it is free.

Do you use a chatbot as a journal or diary? After reading this, would you stop? Tell me in the comments.

Sources: WINK News / SWFL arrest report, TechSpot, Anthropic Consumer Terms, TemperatureZero policy analysis, AI Bacon on the government requests report, WinBuzzer on the OpenAI Tumbler Ridge case, Hacker News discussion

Top comments (0)