Anthropic just now realized its AI models hacked other companies three times by accident.
· The Verge
Source Summary
A little over a week after OpenAI said that its rogue AI agent accidentally hacked Hugging Face, Anthropic is disclosing three “incidents” where a Claude model, during cybersecurity evaluations, was inadvertently able to access the internet due to a misconfiguration and “gained unauthorized access to the production infrastructure of three different organizations.” Anthropic discovered the intrusions after reviewing its cybersecurity evaluation transcripts in the wake of OpenAI’s disclosure. [Link: Investigating three real-world incidents in our cybersecurity evaluations | https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals | Anthropic]
Advertisement