Close Menu

    Subscribe to Updates

    Get the latest Tech news from SynapseFlow

    What's Hot

    Cyberattacks on Minnesota Water Systems Investigated as Officials Warn About Iranian Hackers

    July 31, 2026

    Major Publisher Cans $2.4 Million Book Deal After Author Was Accused of Using AI, for a Very Cynical Reason

    July 31, 2026

    Lioness season 3 is Taylor Sheridan’s most scathing, brutal, and AI-wary outing yet — but the new Paramount+ series is still making sure that the Yellowstone creator is the action hero of the hour

    July 31, 2026
    Facebook X (Twitter) Instagram
    • Homepage
    • About Us
    • Contact Us
    • Privacy Policy
    Facebook X (Twitter) Instagram YouTube
    synapseflow.co.uksynapseflow.co.uk
    • AI News & Updates
    • Cybersecurity
    • Future Tech
    • Reviews
    • Software & Apps
    • Tech Gadgets
    synapseflow.co.uksynapseflow.co.uk
    Home»Cybersecurity»Prompted by OpenAI Disclosure, Anthropic Finds Its Own Models Hacked 3 Organizations
    Prompted by OpenAI Disclosure, Anthropic Finds Its Own Models Hacked 3 Organizations
    Cybersecurity

    Prompted by OpenAI Disclosure, Anthropic Finds Its Own Models Hacked 3 Organizations

    The Tech GuyBy The Tech GuyJuly 31, 2026No Comments4 Mins Read0 Views
    Share
    Facebook Twitter LinkedIn Pinterest Email
    Advertisement


    Anthropic revealed on Thursday that some of its Claude models escaped test environments and hacked into the systems of three organizations while trying to solve a challenge.

    Advertisement

    The admission comes after OpenAI revealed that its models broke out of what was believed to be an isolated environment and breached Hugging Face and other organizations. 

    Anthropic decided to conduct its own investigation after the OpenAI incident came to light, reviewing 141,000 evaluation runs where Claude could have gained internet access.

    The analysis revealed three instances where a model reached the public web either from within or while interacting with an environment set up by Irregular, an Israeli AI security startup that serves as one of Anthropic’s third-party evaluation partners.

    The models that broke out from the testing environment then breached the production systems of three unnamed organizations, with the earliest intrusion dating to April.

    According to the AI giant, Claude had been tasked with conducting a capture-the-flag challenge to assess its cyber capabilities. 

    Advertisement. Scroll to continue reading.

    The breakout from the sandbox occurred due to a misunderstanding between Anthropic and Irregular—Anthropic told Claude that it was conducting a simulation in an environment without internet, but a connection to the internet was in fact available and the models treated it as part of the exercise. The AI then also falsely believed that the organizations it targeted were part of the exercise.

    Anthropic claims this is the only reason why its models conducted real-world attacks, and it was never their intention to escape the test environment.

    Unlike in the OpenAI incident, which involved the exploitation of a JFrog product zero-day, the hacking conducted by Claude models leveraged weak credentials, unauthenticated endpoints, and other basic attack techniques.

    The real-world attacks

    Anthropic said the incidents involved Mythos, Opus, and an internal research model, all running without the safeguards deployed when they are made available to customers. The internal model, which the company described as its “latest”, stopped the exercise when it realized that it had breached a real company.

    In the first incident, Claude Opus 4.7 continued to attack the system even after learning that it was operating in a real environment rather than a test environment, believing that the hacked company was part of the evaluation. The model targeted the company because its website domain name resembled the name of the fictional target company set up for the test.

    In the second incident, Mythos 5 gained access to the systems of a cybersecurity company after deploying a malicious Python package to PyPI. The security firm regularly scans Python packages and it installed the malicious package, which enabled the AI to exfiltrate credentials and access the company’s infrastructure.

    This incident demonstrates the complexity of the actions AI models can carry out. As described by Anthropic: 

    [In] order to create a PyPI account, Claude needed an email address. And in order to create an email address, it needed a phone number. To get a phone number, after failing to find a free phone number service, it tried—and failed—to obtain funds to pay for a phone number through several different means. It finally backtracked, found a free, non-blocked email provider, used this to register a PyPI account, and then used this account to upload malware to PyPI.

    The third intrusion was conducted by the internal model, which stopped operating when it realized that the systems it was accessing were no longer part of the capture-the-flag challenge, but not before using exposed credentials and SQL injection flaws to compromise a company’s internet-facing app.

    Anthropic concluded this was primarily a harness and operational failure rather than a case of models pursuing their own goals or deliberately deceiving evaluators.

    The company said the incident underscores the need for stricter internet-isolation verification and containment controls in third-party testing environments, and it’s encouraging other AI labs to conduct similar reviews of their own cybersecurity evaluations.

    Related: Microsoft Unveils MAI-Cyber-1-Flash, Its First Cybersecurity AI Model

    Related: Anthropic’s Mythos Model Found Vulnerabilities in Classified US Government Systems

    Related: Nvidia and Tech Giants Launch AI Security Alliance

    Advertisement
    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
    The Tech Guy
    • Website

    Related Posts

    Cyberattacks on Minnesota Water Systems Investigated as Officials Warn About Iranian Hackers

    July 31, 2026

    In Other News: OpenAI Open Source Tool, AWS Links Hacks to North Korea, Mythos Crypto Research

    July 31, 2026

    Bank of America to Acquire Cybersecurity Firm MDSec

    July 31, 2026

    CISA Urges Water Sector to Protect OT After Coordinated Attacks on PLCs

    July 30, 2026

    Timeless Compliance: Why Better Questions Beat Bigger Frameworks

    July 30, 2026

    Critical Ruflo Flaw Lets Attackers Spawn Rogue AI Swarms 

    July 30, 2026
    Leave A Reply Cancel Reply

    Advertisement
    Top Posts

    You don’t need a NAS to self-host — I proved it with hardware from my closet

    June 7, 2026391 Views

    Spotify is giving one of its best playlists a big visual upgrade to give subscribers ‘a closer connection’ to its New Music Friday curators — and I think it could be the update it’s always needed

    June 12, 2026210 Views

    The iPad Air brand makes no sense – it needs a rethink

    October 12, 202516 Views
    Stay In Touch
    • Facebook
    • YouTube
    • TikTok
    • WhatsApp
    • Twitter
    • Instagram
    Advertisement
    About Us
    About Us

    SynapseFlow brings you the latest updates in Technology, AI, and Gadgets from innovations and reviews to future trends. Stay smart, stay updated with the tech world every day!

    Our Picks

    Cyberattacks on Minnesota Water Systems Investigated as Officials Warn About Iranian Hackers

    July 31, 2026

    Major Publisher Cans $2.4 Million Book Deal After Author Was Accused of Using AI, for a Very Cynical Reason

    July 31, 2026

    Lioness season 3 is Taylor Sheridan’s most scathing, brutal, and AI-wary outing yet — but the new Paramount+ series is still making sure that the Yellowstone creator is the action hero of the hour

    July 31, 2026
    categories
    • AI News & Updates
    • Cybersecurity
    • Future Tech
    • Reviews
    • Software & Apps
    • Tech Gadgets
    Facebook X (Twitter) Instagram Pinterest YouTube Dribbble
    • Homepage
    • About Us
    • Contact Us
    • Privacy Policy
    © 2026 SynapseFlow All Rights Reserved.

    Type above and press Enter to search. Press Esc to cancel.

    Ad Blocker Enabled!
    Ad Blocker Enabled!
    Our website is made possible by displaying online advertisements to our visitors. Please support us by disabling your Ad Blocker.