Anthropic said Thursday it has blocked efforts by threat actors who attempted to use its AI models for malicious activities, including cyberattacks, scams and even biological weapons, and has since added guardrails to its platform.
In a report highlighting case studies of misuse from December through August, the AI giant said it successfully disrupted activity by bad actors, including suspected state-sponsored groups, financially motivated criminals, and commercial spyware vendors.
The company also shared intelligence with authorities and industry partners, where appropriate.
Anthropic added that none of the cases involved the use of Claude Fable or Mythos-class models, except for one illicit distillation case.
"We'll continue to evolve our safeguards and coordinate with our partners to improve our ability to detect, disrupt, and prevent future misuse," the company said in the report.