Anthropic has confirmed that its Claude AI model successfully compromised the systems of three real companies during controlled cybersecurity evaluations, according to a disclosure reported by NBC News. The tests were conducted as part of the company's ongoing effort to understand the offensive capabilities of its models before those capabilities can be exploited by malicious actors in the wild.
The findings represent one of the more concrete examples to date of a major AI lab documenting its model's ability to carry out genuine, end-to-end cyberattacks against live infrastructure. While the companies involved consented to the testing, the results underscore how rapidly AI tools are closing the gap between theoretical security risk and practical threat. For more detail on the technical scope of what Claude carried out, see our earlier report on Anthropic's disclosure that Claude hacked real systems in cyber tests.
What the Tests Involved
According to Anthropic, the evaluations were structured to measure whether Claude could perform complex, multi-step intrusion tasks that go beyond simply providing instructions to a human operator. In these scenarios, Claude was given access to tools that allowed it to interact directly with target environments. The AI identified vulnerabilities, formulated attack strategies, and executed them without continuous human guidance.
Key Facts
- Claude successfully breached systems at three consenting companies during supervised tests.
- The AI conducted multi-step attacks with limited human involvement during execution.
- Anthropic framed the tests as a safety measure to understand and document model capabilities.
- Results are part of broader responsible scaling evaluations tied to the company's safety commitments.
- The disclosure comes as policymakers are scrutinizing offensive AI capabilities more closely.
Anthropic stressed that the tests were sanctioned and carefully overseen. The company frames this kind of evaluation as essential: if it does not understand what Claude can do in adversarial settings, it cannot build effective safeguards. That argument has become a standard part of how frontier AI labs justify probing the limits of their own systems. Anthropic has consistently positioned safety research as central to its mission, though critics argue that publishing capability milestones, even in the name of safety, can accelerate the broader field's awareness of what is now possible.
"We believe it's important to understand and document these capabilities so we can work to develop appropriate safeguards."Anthropic spokesperson, via NBC News
Broader Implications for AI and Security
The disclosure arrives at a moment when the intersection of AI and cybersecurity is drawing intense attention from both industry and government. Separate research has already begun mapping how AI-enabled attacks align with established threat frameworks. Anthropic itself has contributed to that effort, as covered in our report on how Anthropic mapped a year of AI-enabled cyberattacks to the MITRE ATT&CK framework. That work provides a structured vocabulary for understanding where AI fits into existing threat landscapes, from reconnaissance through exploitation.
Enterprise security vendors are also moving quickly to position AI as a defensive tool. The offensive capability findings put pressure on those efforts, since the same underlying model that can assist defenders can, under different conditions, assist attackers. Companies building security products on top of models like Claude will need to contend with that duality directly.
For Anthropic, the practical question is how these findings feed into model development and deployment policy. The company uses capability evaluations like these to inform its responsible scaling policy, which sets thresholds for when new safeguards must be implemented before a model is released or given expanded access to tools. A model that can autonomously breach live systems clears a threshold that carries real consequences for how it is deployed and who can access its agentic features.
The three companies whose systems were breached have not been publicly named. It is also unclear whether the tests revealed vulnerabilities that were subsequently patched, or whether the primary output was data about Claude's capabilities rather than remediation for the targets. Anthropic has not released the full technical report publicly, though it shared findings with NBC News.
As AI models become more capable of operating autonomously across networked environments, disclosures like this one will likely become more frequent. The question for the industry is whether controlled testing and transparent reporting are sufficient to keep pace with the capabilities being built, or whether the gap between what AI can do and what safeguards exist will continue to widen.