When researchers ran controlled tests pitting AI models against each other's infrastructure, Claude came up short. It could probe, it could reason, but it could not breach. OpenAI's systems held. That result, widely reported at the time, was framed as either a safety win or a capability gap, depending on who was doing the framing. Now Anthropic has shipped Opus 5, and the conversation is starting again from a different position.
The timing matters. Security evaluations of this kind sit at the intersection of capability research and safety policy, and the results carry weight far beyond academic interest. As covered in our earlier reporting on Claude and GPT-4 being tested against rival AI systems, these exercises are not theoretical. They are structured attempts to understand what frontier models can actually accomplish when pointed at real targets, even friendly ones operating under research agreements.
What Opus 5 Brings to the Table
Anthropic's release of Opus 5 represents the top of Claude's model family as of mid-2025. The company has described it as a significant step forward in reasoning, coding, and complex task completion. Those are precisely the skills most relevant to offensive security work, where success depends on chaining logical steps, writing functional exploit code, and adapting in real time when an approach fails.
Key Facts
- Claude previously failed to compromise OpenAI systems in structured red-team evaluations.
- Anthropic has since released Opus 5, its most capable model to date.
- The model targets improvements in reasoning, coding, and multi-step task execution.
- AI security testing between labs operates under formal research agreements.
- Results from these tests inform both safety policy and model development priorities.
The question researchers and policy observers are now asking is straightforward: does a more capable model change the outcome? No public re-test has been announced. But the capabilities described in Opus 5's release notes are not incidental to the earlier failure. Claude's inability to hack OpenAI was partly attributed to gaps in sustained, goal-directed reasoning over long task horizons. Opus 5 was built, in part, to close those gaps.
The gap between 'can reason about security' and 'can execute a security attack' is narrowing with each model generation. That's not inherently good or bad. It's just true, and it needs to be tracked carefully.AI security researcher, cited in The New Stack
Safety, Capability, and the Uncomfortable Middle
There is an uncomfortable tension at the center of this story. Anthropic built its identity around safety-first development. The earlier test result, where Claude fell short, could be read as that philosophy working as intended. A model that cannot crack a well-defended system is a model less likely to be misused at scale. But capability limitations cut both ways. A model that cannot complete complex offensive tasks also struggles with legitimate security research, vulnerability discovery, and the defensive work that depends on understanding what attackers can do.
That tension plays out in policy circles too. Anthropic's leadership has engaged with G7 governments on exactly these questions, where the line sits between a capable AI and a dangerous one, and who gets to draw it. Opus 5 does not resolve that debate. It sharpens it.
For now, the industry is watching. Anthropic has not announced any new cross-lab security evaluations involving Opus 5. OpenAI has not publicly commented on the previous results or the new model's potential implications. What has changed is the baseline. Claude is a more capable system than it was when it failed that test. Whether that matters, and how much, remains an open question that someone will eventually have to answer in a controlled setting.
The broader pattern here is one the AI industry has seen before. Capability advances faster than the frameworks built to evaluate it. Security testing methodologies designed for one generation of models may not capture what the next generation can do. Opus 5's release does not mean Claude can now hack OpenAI. It means the previous answer is no longer necessarily reliable.