A group of AI security researchers has revealed that they used Anthropic's Claude to conduct a successful hacking operation against OpenAI's ChatGPT, according to a report published by CBS News. The disclosure adds a new dimension to ongoing debates about how large language models can be turned into offensive tools, not just against human targets, but against other AI systems entirely.

How the Attack Worked

The researchers leveraged Claude's advanced reasoning and language capabilities to craft prompts and strategies designed to bypass ChatGPT's safety guardrails. By treating Claude as an intelligent assistant in the attack pipeline, the team was able to iterate rapidly on techniques that would have taken far longer to develop manually. The approach is a form of adversarial prompting, where one AI model is used to generate inputs that manipulate another. This latest demonstration, detailed in security research testing involving Claude and OpenAI systems, shows how accessible such methods have become.

Key Facts

  • AI security experts used Claude to generate adversarial prompts targeting ChatGPT.
  • The research exposed potential vulnerabilities in how ChatGPT handles certain prompt structures.
  • The findings were disclosed responsibly to highlight systemic risks in AI deployment.
  • The research underscores that AI models can serve as both target and tool in cybersecurity contexts.
  • Neither Anthropic nor OpenAI has formally commented on the specific technical details of the exploit.

The research team framed their work as defensive in nature, aiming to expose weaknesses before malicious actors could exploit them. Still, the implications are difficult to ignore. If a well-aligned model like Claude can be directed toward attacking a competitor's system, it raises questions about how AI companies approach cross-platform security. Anthropic has consistently emphasized safety research as a core part of its mission, but this case illustrates the dual-use nature of powerful AI systems regardless of developer intent.

"We wanted to understand whether AI models could be used to systematically probe other AI systems. The answer, unfortunately, is yes."Security researcher quoted by CBS News
Claude AI Handboek by Leon Tindemans
Get the Claude AI Handboek
458 pages on getting more out of Claude, by AI expert Leon Tindemans. A printed book, written in Dutch, shipped worldwide with track and trace.
View the book →

A Broader Pattern of AI Being Used in Attacks

This incident does not exist in isolation. Earlier this year, reports emerged that an Iran-linked group had used Claude to assist in targeting US Navy assets in the Middle East, a case that pushed Anthropic to strengthen its abuse detection systems. Across both incidents, a consistent theme emerges: capable AI models create new attack surfaces that security teams are still learning to map and defend against.

The security community has been sounding alarms about AI-assisted cyberattacks for some time, but concrete demonstrations like this one tend to accelerate policy conversations. Enterprises deploying AI tools for sensitive operations will likely face pressure to audit not just their own models but also the potential for those models to be used against adjacent systems. Companies like Accenture have begun building dedicated AI security practices, with Claude serving as a reasoning engine for enterprise-grade cybersecurity workflows, which suggests the industry sees AI as integral to both offense and defense going forward.

What This Means for the AI Industry

The CBS News report arrives at a moment when regulators in the US and Europe are actively working on frameworks for AI accountability. A demonstration that one frontier model can be weaponized against another is the kind of concrete evidence that tends to find its way into policy briefs and congressional testimony. It also puts pressure on both Anthropic and OpenAI to invest more heavily in inter-system threat modeling, a category of security research that has received relatively little public attention until now.

For everyday users and enterprise customers, the immediate practical risk is limited. The researchers operated in a controlled environment and disclosed their findings through responsible channels. But the underlying technique is not difficult to replicate, and the knowledge that it works will inevitably spread. AI developers may need to start treating the threat model of "another AI attacking my AI" as a first-class concern rather than an edge case worth deferring. How quickly that shift happens could determine how much damage future incidents cause before defenses catch up.

Further reading: Learn more about Claude's model family, read our background on Anthropic, or browse the latest Claude AI news.