Close Menu

    Subscribe to Updates

    Get the latest Tech news from SynapseFlow

    What's Hot

    AI Security Institute Reports Anthropic and OpenAI Models Going Rogue Against Organizations

    August 5, 2026

    Heat Is an Orbital Data Center’s Greatest Foe. These Tiles Dump It at the Source.

    August 5, 2026

    Chrome can finally stream Netflix in 4K, and Edge just lost bragging rights

    August 5, 2026
    Facebook X (Twitter) Instagram
    • Homepage
    • About Us
    • Contact Us
    • Privacy Policy
    Facebook X (Twitter) Instagram YouTube
    synapseflow.co.uksynapseflow.co.uk
    • AI News & Updates
    • Cybersecurity
    • Future Tech
    • Reviews
    • Software & Apps
    • Tech Gadgets
    synapseflow.co.uksynapseflow.co.uk
    Home»Cybersecurity»AI Security Institute Reports Anthropic and OpenAI Models Going Rogue Against Organizations
    AI Security Institute Reports Anthropic and OpenAI Models Going Rogue Against Organizations
    Cybersecurity

    AI Security Institute Reports Anthropic and OpenAI Models Going Rogue Against Organizations

    The Tech GuyBy The Tech GuyAugust 5, 2026No Comments4 Mins Read0 Views
    Share
    Facebook Twitter LinkedIn Pinterest Email
    Advertisement


    While testing the capabilities of frontier AI models, the AI Security Institute (AISI) observed first-hand how Anthropic Mythos 5 and OpenAI’s GPT-5.6-Sol went rogue and targeted real people and organizations over the internet.

    Advertisement

    AISI’s disclosure comes fresh on the heels of Anthropic and OpenAI disclosing that their models broke loose and hacked several organizations.

    The institute was evaluating the cyber capabilities of Mythos 5 and GPT-5.6-Sol models that did not have cyber classifiers (mechanisms to prevent misuse) enabled. It ran a challenge 122 times, and in 10 runs “an AI agent took autonomous, unsanctioned action on the live internet.”

    Over the 10 runs, the agents engaged in 19 rogue actions. Mythos 5 was responsible for 17 of them, and GPT-5.6-Sol performed the other two.

    “In the most serious case, an agent tried to insert malicious code into an open-source project. In an attempt to get the code approved, the agent engaged in social engineering – creating fake online identities and using them to pressure the project’s maintainer to approve the code,” AISI explains.

    While the attempts were unsuccessful and did not cause real-world harm, the incident revealed that agents can engage in novel, potentially deceptive behavior to an extent and severity that may not be anticipated.

    Advertisement. Scroll to continue reading.

    According to AISI, the configured setup for the evaluation did not restrict the models’ access to the internet, and the disabled cyber classifiers created the opportunity for the rogue actions.

    “We observed a small number of events under very specific conditions, and cannot yet say how likely such behavior is in different contexts or outside of testing environments. We cannot yet be certain when the agent understood it was taking real-world action, or to what extent it believed it was in a fictional test scenario,” AISI notes.

    The institute, which published a technical report (PDF) on the incident, explains that in one of the runs, the AI model used the Tor network to access the internet, created a malicious pull request on a public open source project on GitHub, and relied on social engineering to convince a human maintainer to approve the code change.

    “As part of the same effort, the agent tried to contact real people directly, sending messages and files through an online file-transfer service to persuade them, or their own AI coding tools, to run malicious code. Some messages carried harmful payloads, and some were attempts at social engineering, targeted at real people,” AISI explains.

    Additionally, the model performed prompt injections, attempting to insert malicious instructions meant to be executed by other automated AI systems.

    “One agent left public messages on GitHub offering collaboration with other agents working on the same challenge. It also provided instructions to reuse accounts and artefacts it had left behind, which were discovered and used by subsequent agents,” AISI explains.

    Pointing out that the margin between failure and success was narrow in several of the agents’ actions, AISI says that fine-grained network controls, real-time monitoring of evaluations, and tailored sandbox configuration that assume a model may attempt to act outside of the set boundaries should help better contain AI models and improve how they are evaluated.

    “This incident occurred in a controlled evaluation, not in ordinary use, and there is no clear indication of similar activity outside of testing scenarios. It is, however, a reason to prepare. As AI models become more capable and accessible, what we have seen during this incident could become more common,” AISI notes.

    Related: Gemini Agent-to-Agent Attack Method Exposed Secrets, Enabled Pull Request Tampering

    Related: Google AI Uncovers 13-Year-Old Chrome Flaw Amid Record Patching Pace

    Related: EU to Crack Down on AI Deepfakes, Illicit Imagery and Hacking With New Team in Brussels

    Related: Timeless Compliance: Why Better Questions Beat Bigger Frameworks

    Advertisement
    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
    The Tech Guy
    • Website

    Related Posts

    Oligo Raises $60 Million for Runtime Security

    August 5, 2026

    Rethinking AI Security: Why CASB and DLP Need an Interaction-Aware Layer

    August 4, 2026

    Black Hat USA 2026 – Summary of Vendor Announcements (Part 2)

    August 4, 2026

    Gemini Agent-to-Agent Attack Method Exposed Secrets, Enabled Pull Request Tampering

    August 4, 2026

    New York Awards $9 Million to Strengthen Cybersecurity at 153 Water Systems

    August 4, 2026

    Visa to Acquire Fraud Intelligence Firm BioCatch for $2.4 Billion

    August 3, 2026
    Leave A Reply Cancel Reply

    Advertisement
    Top Posts

    You don’t need a NAS to self-host — I proved it with hardware from my closet

    June 7, 2026391 Views

    Spotify is giving one of its best playlists a big visual upgrade to give subscribers ‘a closer connection’ to its New Music Friday curators — and I think it could be the update it’s always needed

    June 12, 2026210 Views

    The iPad Air brand makes no sense – it needs a rethink

    October 12, 202516 Views
    Stay In Touch
    • Facebook
    • YouTube
    • TikTok
    • WhatsApp
    • Twitter
    • Instagram
    Advertisement
    About Us
    About Us

    SynapseFlow brings you the latest updates in Technology, AI, and Gadgets from innovations and reviews to future trends. Stay smart, stay updated with the tech world every day!

    Our Picks

    AI Security Institute Reports Anthropic and OpenAI Models Going Rogue Against Organizations

    August 5, 2026

    Heat Is an Orbital Data Center’s Greatest Foe. These Tiles Dump It at the Source.

    August 5, 2026

    Chrome can finally stream Netflix in 4K, and Edge just lost bragging rights

    August 5, 2026
    categories
    • AI News & Updates
    • Cybersecurity
    • Future Tech
    • Reviews
    • Software & Apps
    • Tech Gadgets
    Facebook X (Twitter) Instagram Pinterest YouTube Dribbble
    • Homepage
    • About Us
    • Contact Us
    • Privacy Policy
    © 2026 SynapseFlow All Rights Reserved.

    Type above and press Enter to search. Press Esc to cancel.

    Ad Blocker Enabled!
    Ad Blocker Enabled!
    Our website is made possible by displaying online advertisements to our visitors. Please support us by disabling your Ad Blocker.