Close Menu

    Subscribe to Updates

    Get the latest Tech news from SynapseFlow

    What's Hot

    Phishing Research Challenges Conventional Security Awareness Testing

    September 11, 2026

    Lawyers Already Lining Up to Defend Victims of Cybercab Crashes

    September 11, 2026

    Heavys H1E review: Born for One Thing

    September 11, 2026
    Facebook X (Twitter) Instagram
    • Homepage
    • About Us
    • Contact Us
    • Privacy Policy
    Facebook X (Twitter) Instagram YouTube
    synapseflow.co.uksynapseflow.co.uk
    • AI News & Updates
    • Cybersecurity
    • Future Tech
    • Reviews
    • Software & Apps
    • Tech Gadgets
    synapseflow.co.uksynapseflow.co.uk
    Home»Future Tech»Frontier AI Is Faceplanting at Real-World Workplace Tasks
    Frontier AI Is Faceplanting at Real-World Workplace Tasks
    Future Tech

    Frontier AI Is Faceplanting at Real-World Workplace Tasks

    The Tech GuyBy The Tech GuyJuly 21, 2026No Comments4 Mins Read1 Views
    Share
    Facebook Twitter LinkedIn Pinterest Email
    Advertisement



    Sign up to see the future, today

    Advertisement

    Can’t-miss innovations from the bleeding edge of science and tech

    To date, AI industry spending has topped $1.6 trillion, and shows no sign of slowing anytime soon.

    So what do we actually have to show for it? Historically, it’s been a whole lot of nothing: as numerous studies have shown us, tools like AI chatbots and autonomous agents have been ineffective at completing real world tasks in a competent way.

    The tech industry insists that’s all about to change within the next few years, as AI’s capabilities grow by leaps and bounds, enabling economic growth the likes of which the world has never seen. But is it really?

    Not necessarily. A new study out of the University of California Berkeley’s Center for Responsible, Decentralized Intelligence — flagged by the College Fix — shows that frontier AI tools of all makes and models are still incapable of completing the vast majority of workplace tasks at an acceptable level, throwing a major wrench in the tech industry’s assertions that the AI revolution is imminent.

    To come to that conclusion, the UC researchers designed a rigorous assessment they call the “Agents’ Last Exam,” developed to test “job-readiness” across numerous state-of-the-art AI models. Basically, the ALE — an impish riff on “Humanity’s Last Exam” — is designed to put an AI system through its paces, covering “more than 1,500 expert-sourced tasks spanning 55 occupations,” the researchers wrote in apress release.

    Those test spans the typical line-up of AI-exposed jobs like software engineering and graphic design, but also a substantial number of jobs whose fates remain less certain, such as maritime engineering, agriculture, audio production, and public health operations.

    Using the ALE benchmark, researchers took a hard look at advanced “closed” models — proprietary AI systems developed by private companies — like Anthropic’s Fable 5, OpenAI’s GPT-5.5, Cursor’s Composer 2.5, and Google’s Gemini 3.1 Pro. (For good measure, they also looked at two open-source models by Chinese developers.)

    As cutting-edge as these AI models are, the research found that they’re far from ready for the complex needs of the modern workplace. Out of all of the models put through the gauntlet, each of them failed spectacularly. OpenAI’s GPT-5.5 came in with the highest score: a passing rate of just 24 percent overall.

    “Today’s agents can solve a meaningful fraction of professional tasks,” the researchers wrote. “However, when we look at the hardest tasks that require sustained reasoning, deep domain expertise, and reliable execution over long horizons, they are still far from human-level performance.”

    And as tasks became more complicated, even those meager aggregate scores fell off fast.

    “On ALE’s hardest tier, every frontier agent we tested, including Fable 5, achieved a 0 percent success rate,” the presser explains.

    The researchers also break down some cost considerations. The cutting edge Fable 5, they note, delivers “similar performance” to models like GPT-5.5 and Composer 2.5, “while costing roughly 4-12× more per completed task.”

    Despite the horrible test results, researchers caution that the technology could still upend the job market for more AI-exposed — as plenty of corporate executives have shown us, the tech doesn’t need to work particularly well to keep workers on their back heels.

    “Even if current pass rates remain relatively low, occupations dominated by routine and well-defined procedures are likely to experience disruption first, while decision-intensive roles will remain more resilient for longer,” Berkeley computer science researcher and study co-author Dawn Song told College Fix.

    “The key factor,” Song added, “is not the industry itself, but the nature of the work.”

    More on AI: OpenAI Appears to Be Missing Its Sales Goals by a Vast Margin

    Advertisement
    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
    The Tech Guy
    • Website

    Related Posts

    Lawyers Already Lining Up to Defend Victims of Cybercab Crashes

    September 11, 2026

    Gen1 Makes It Easy To Capture GrandParents Stories – Know Your Family History – NextBigFuture.com

    September 11, 2026

    APOD: 2026 September 11 – M83: The Southern Pinwheel

    September 11, 2026

    Europe’s First Private Rocket Reaches Orbit

    September 11, 2026

    OpenAI’s Supposed Mathematical Breakthrough Devolves Into Explosive Drama as Mathematician Accuses It of Stealing His Work

    September 10, 2026

    World GDP 2024 to 2037 – Country by Country – NextBigFuture.com

    September 10, 2026
    Leave A Reply Cancel Reply

    Advertisement
    Top Posts

    You don’t need a NAS to self-host — I proved it with hardware from my closet

    June 7, 2026391 Views

    Spotify is giving one of its best playlists a big visual upgrade to give subscribers ‘a closer connection’ to its New Music Friday curators — and I think it could be the update it’s always needed

    June 12, 2026210 Views

    The iPad Air brand makes no sense – it needs a rethink

    October 12, 202517 Views
    Stay In Touch
    • Facebook
    • YouTube
    • TikTok
    • WhatsApp
    • Twitter
    • Instagram
    Advertisement
    About Us
    About Us

    SynapseFlow brings you the latest updates in Technology, AI, and Gadgets from innovations and reviews to future trends. Stay smart, stay updated with the tech world every day!

    Our Picks

    Phishing Research Challenges Conventional Security Awareness Testing

    September 11, 2026

    Lawyers Already Lining Up to Defend Victims of Cybercab Crashes

    September 11, 2026

    Heavys H1E review: Born for One Thing

    September 11, 2026
    categories
    • AI News & Updates
    • Cybersecurity
    • Future Tech
    • Reviews
    • Software & Apps
    • Tech Gadgets
    Facebook X (Twitter) Instagram Pinterest YouTube Dribbble
    • Homepage
    • About Us
    • Contact Us
    • Privacy Policy
    © 2026 SynapseFlow All Rights Reserved.

    Type above and press Enter to search. Press Esc to cancel.

    Ad Blocker Enabled!
    Ad Blocker Enabled!
    Our website is made possible by displaying online advertisements to our visitors. Please support us by disabling your Ad Blocker.