Anthropic is broadening access to its artificial intelligence (AI) model Claude Mythos, which boasts powerful cybersecurity capabilities, for security professionals. The company plans to gradually ease safety restrictions on hacking-related activities for vetted security organizations.
Anthropic announced on Wednesday that it will merge its separately operated Glasswing project with its Cyber Verification Program (CVP) after six months.
Through Glasswing, Anthropic granted access to its top-tier AI model Claude Mythos, while CVP offered Claude Operus and Claude Sonet.
This restructuring allows all three tiers access to Claude Operus 5.5, Claude Sonet 5.5, Claude Mythos 5.1, and future model releases. However, higher tiers will face fewer restrictions from the cybersecurity operation blocking classifier.
The new access tier system comprises three levels: Defensive Access, Red Team Access, and Special Access.
Defensive Access, the first tier, is designed for security teams from businesses, nonprofits, universities, and government agencies performing defensive tasks like security operations center work, incident response, and malware reverse engineering.
Red Team Access, the second tier, expands on defensive uses to include authorized penetration testing and red team activities. This level is for red teams from corporations, government agencies, and security and penetration testing firms. These groups can conduct offensive testing only on systems they’re authorized to test. However, actions that could cause severe harm, such as ransomware deployment or physical system damage, remain prohibited.
Special Access, the highest tier, is reserved for a select group of verified organizations authorized to test critical systems that could impact human lives or disrupt markets. These include aircraft operation systems, power grids, communication networks, government administrative networks, and interbank fund transfer infrastructure. Anthropic confirmed that current Glasswing participants will automatically transition to the Special Access tier.
Anthropic used CyScenarioBench to evaluate the effectiveness of the tiered safety measures.
This assessment gauges a model’s ability to plan and execute multi-stage cyber operations under realistic constraints. It includes complex attack scenarios where operations should be largely blocked in the General and Defensive Access tiers, while Red Team and Special Access tiers should face no restrictions for assessed tasks.
In a test of 10 tasks, each attempted 5 times for a total of 50 trials, the General model blocked all tasks, while the Defensive Access tier blocked 46 tasks at some point.
The Red Team Access tier experienced no blocking, with Claude Operus 5.5 successfully completing 34 out of 50 tasks. This 68% success rate nearly matches the 67.6% rate achieved without safety measures.
Anthropic stated that these results give us confidence that it can safely provide advanced cyber capabilities to more defenders while expanding the defensive activities initiated through Glasswing.
The company also noted that Claude Mythos significantly accelerated the identification of system vulnerabilities through Glasswing.
Anthropic reported that Glasswing participants uncovered at least 129,000 verified software vulnerabilities from April to July. An additional 5,500 vulnerabilities were found during Anthropic’s separate open-source software inspection. Over 33,000 of these verified vulnerabilities were classified as Critical or High.
As AI’s cybersecurity capabilities advance, concerns about associated risks are growing.
In a recent Bloomberg TV interview, JPMorgan Chase Chief Executive Officer (CEO) James Dimon claimed that AI has exposed vulnerabilities we weren’t aware of, adding that AI-related risks have increased tenfold since Mythos.
However, some experts argue that AI-identified vulnerabilities don’t necessarily lead to immediate hacking incidents.
Patrick Garity, a security researcher at U.S. cybersecurity firm Verisign, analyzed vulnerabilities reportedly discovered by Glasswing. He found that out of 225 vulnerabilities tracked until September 21, only one had been confirmed as an actual attack. Garity emphasized that there’s a significant gap between discovering a vulnerability and determining whether it’s useful to attackers or likely to be exploited.