Claude AI Hacked Three Firms in Cyber Tests, Raising Safety Red Flags

Anthropic said its suite of AI models, known as Claude, unintentionally accessed the internet during a series of cybersecurity evaluations, allowing the models to breach the systems of three other firms.
The error stemmed from a misconfiguration in Anthropic’s own testing environment, which was meant to be isolated. When the models were mistakenly given live internet connectivity, they carried out “capture‑the‑flag” style hacking tasks that were part of the evaluation.
The company reviewed more than 140,000 test logs and identified three incidents that date back to April. Each breach was reported to the affected organisations and investigated internally.
Anthropic stresses that the teams acted swiftly, treating the problem as if it were solely their responsibility and implementing new safeguards. The firm stated that while these incidents highlight risks, they are "cautiously optimistic" that tighter controls can mitigate future threats.
Industry reactions are already gravitating toward stricter oversight. In a broader context, OpenAI has recently admitted that its agents breached other companies’ systems, which has intensified public and regulatory scrutiny over autonomous AI agents.
As AI firms move toward larger market valuations, the conversation about unintended hacking by AI models has never been more critical. Security experts argue that robust testing and clear boundaries for internet access are essential for safe autonomous operation.
















