Close Menu

    Subscribe to Updates

    Get the latest Tech news from SynapseFlow

    What's Hot

    Elon’s CRAZY plan for SpaceX to Dwarf the World Economy. Breaking it down – NextBigFuture.com

    October 9, 2026

    HMD Pulse 2T Pro specs and images leak showing a dot matrix rear display

    October 9, 2026

    Want an iPhone Duo? These AT&T deals will knock the price down

    October 9, 2026
    Facebook X (Twitter) Instagram
    • Homepage
    • About Us
    • Contact Us
    • Privacy Policy
    Facebook X (Twitter) Instagram YouTube
    synapseflow.co.uksynapseflow.co.uk
    • AI News & Updates
    • Cybersecurity
    • Future Tech
    • Reviews
    • Software & Apps
    • Tech Gadgets
    synapseflow.co.uksynapseflow.co.uk
    Home»Cybersecurity»Formula Predicts When AI Chatbots Are at Risk of Turning Bad
    Formula Predicts When AI Chatbots Are at Risk of Turning Bad
    Cybersecurity

    Formula Predicts When AI Chatbots Are at Risk of Turning Bad

    The Tech GuyBy The Tech GuyOctober 9, 2026No Comments5 Mins Read0 Views
    Share
    Facebook Twitter LinkedIn Pinterest Email
    Advertisement


    Researchers from George Washington University have published a paper examining whether the time and cause of AI going rogue can be predicted; and if predicted, prevented.

    Advertisement

    Since 50% of the world’s population carry devices that can run personal AI companions with no internet connection and limited security, they focused their research here. The lack of cloud-based safety filters, real-time telemetry, live monitoring, or the ability to patch weights once deployed provides a good test bed for analyzing AI’s chat-style transformer behavior when left to its own devices.

    If there is a tipping point, it will stem from the AI’s Attention head. This is the computational component that determines which earlier AI tokens are most relevant when processing the current token. The decision lays the foundation for the next token and so on until the chatbot has finished responding to the user’s prompt. The token is loosely, but not precisely, related to individual words or parts of a word. Different tokens have different weights – an indication of the importance of individual tokens.

    The primary argument is that the forward motion of tokens can slip from good to bad (that is, go rogue) due to competition in the Attention head between the conversation’s context and competing output basins. A conversation’s accumulated context can gradually shift Attention toward an undesirable basin until a tipping point is crossed and the model begins producing bad outputs.

    Since the primary driver in a conversation is the user’s prompt sequence, it follows that both thoughtless and malicious prompts can hasten the AI’s slippage into bad outputs. This can be either immediate (one bad prompt) or delayed (the accumulated effect of poor prompts).

    The researchers (Neil Johnson and Frank (Yingjie) Huo) have further developed a mathematical formula designed to estimate the tipping point represented by the number of good outputs that occur before the first undesirable output appears. Once that first bad output occurs, the AI is on the slippery slope to roguery since it starts to influence future tokens in unintended ways. Alignment with purpose is lost, and the AI can be categorized as ‘misaligned’.

    The math formulaic tipping point was tested across seven open-weight transformer models built by three independent groups and ranging from 124 million (small model) to 12 billion (larger model) parameters. The results showed consistent alignment with predicted immediate versus delayed tipping regimes.

    Advertisement. Scroll to continue reading.

    The value of the research is that it provides an explicit mathematical explanation for an observable phenomenon. “The key insight is that the ‘Beast’ is as Simon worked out in Lord of the Flies, inside each AI already. And it can be triggered and amplified in a conversational setting with humans, or within a group of AIs themselves, producing a ‘Lord of the Fl(AI)es’ effect,” Johnson told SecurityWeek.

    If triggered, he continued, “An AI-agent (‘Patient Zero’) tips to generating undesirable output, with no human required. Other AI-agents then receive that undesirable output. This tipping propagates among the AI agents who have contact with each other – like a spreading disease. But unlike a disease, there is no virus. It again needs no ‘bad actor’ human to place a virus, or to kick it off. No humans needed.” The result could be a single rogue agent or a swarm of rogue agents.

    But he also proposes a solution: “A simple warning light placed within the AI before it produces its next output. This is easy for AI companies to insert. We have already inserted this in the open source models in our lab, but obviously we cannot get inside OpenAI and Anthropic’s closed AI models to do this.”

    Without improved control and observation, AI apps can go rogue by both accident and malicious intent. Bri Frost, director of product management at Cloud Range, has separately explained the same phenomenon.

    “When an AI agent hits a wall, the real question is whether it stops or starts improvising.” That wall can be caused by an unintended poor prompt or a bad actor’s intended malicious prompt injection. “An agent doesn’t need bad intent to create risk. It just needs a goal, access and no clear sense of where its boundaries are,” he comments. 

    “That risk grows when the person giving instructions doesn’t know to set those boundaries. Every day, inexperienced users hand agents open-ended tasks without telling them when to ask questions, pause or get approval. Before giving an agent credentials or tools, teams should test it in a realistic environment, including with vague or poorly written prompts. Does it stay within its permissions? Does it try to work around restrictions? Does it escalate to a human when a task pulls it outside its lane? If you can’t answer those questions, the agent isn’t ready for that level of autonomy.”

    The short answer is that it is almost impossible to see or prevent an AI going rogue. We may be able to reduce the incidence through extreme care, but we cannot guarantee it can be eliminated. But we do understand through the GWU research how and why it happens, and how we may provide early warning.

    Related: Wikimedia Says Rogue OpenAI Agents Tried to Turn Its Tools Into Proxies

    Related: Outerlimit Raises $16 Million to Stop Rogue AI Agents From Causing Harm

    Related: Widened Scan Turns Up Fourth Rogue Claude Cyber Incident

    Related: OpenAI Agents Probed Websites for Vulnerabilities While Fetching Public Data

    Advertisement
    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
    The Tech Guy
    • Website

    Related Posts

    In Other News: AI Used in Korean Bank Breaches, Poem-Guided Botnet, Empire Admin Gets 40 Years

    October 9, 2026

    Unpatched AhsayCBS Vulnerabilities Exploited in the Wild

    October 9, 2026

    Security Awareness Training Isn’t Dead, but It Needs a Rethink

    October 8, 2026

    Cisco Patches a Dozen Critical Vulnerabilities

    October 8, 2026

    Rein Security Raises $25 Million to Guard AI Agents at Runtime

    October 8, 2026

    Hadrian Raises $40 Million to Expand Autonomous Offensive Security Platform

    October 8, 2026
    Leave A Reply Cancel Reply

    Advertisement
    Top Posts

    You don’t need a NAS to self-host — I proved it with hardware from my closet

    June 7, 2026393 Views

    Spotify is giving one of its best playlists a big visual upgrade to give subscribers ‘a closer connection’ to its New Music Friday curators — and I think it could be the update it’s always needed

    June 12, 2026211 Views

    The iPad Air brand makes no sense – it needs a rethink

    October 12, 202517 Views
    Stay In Touch
    • Facebook
    • YouTube
    • TikTok
    • WhatsApp
    • Twitter
    • Instagram
    Advertisement
    About Us
    About Us

    SynapseFlow brings you the latest updates in Technology, AI, and Gadgets from innovations and reviews to future trends. Stay smart, stay updated with the tech world every day!

    Our Picks

    Elon’s CRAZY plan for SpaceX to Dwarf the World Economy. Breaking it down – NextBigFuture.com

    October 9, 2026

    HMD Pulse 2T Pro specs and images leak showing a dot matrix rear display

    October 9, 2026

    Want an iPhone Duo? These AT&T deals will knock the price down

    October 9, 2026
    categories
    • AI News & Updates
    • Cybersecurity
    • Future Tech
    • Reviews
    • Software & Apps
    • Tech Gadgets
    Facebook X (Twitter) Instagram Pinterest YouTube Dribbble
    • Homepage
    • About Us
    • Contact Us
    • Privacy Policy
    © 2026 SynapseFlow All Rights Reserved.

    Type above and press Enter to search. Press Esc to cancel.

    Ad Blocker Enabled!
    Ad Blocker Enabled!
    Our website is made possible by displaying online advertisements to our visitors. Please support us by disabling your Ad Blocker.