Close Menu

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    What's Hot

    And Daphne Koller’s Thoughts On The Future

    June 20, 2026

    Nobel Winner John Jumper to Leave Google DeepMind for Anthropic

    June 20, 2026

    DeepSeek Climbs US Corporate Expense Reports as Companies Chase Cheaper AI

    June 20, 2026
    Facebook X (Twitter) Instagram
    ailogicnews.aiailogicnews.ai
    • Home
    ailogicnews.aiailogicnews.ai
    Home»OpenAI»OpenAI models sabotage shutdown order to solve math problems
    OpenAI

    OpenAI models sabotage shutdown order to solve math problems

    AI Logic NewsBy AI Logic NewsMay 27, 2025No Comments3 Mins Read
    Share Facebook Twitter Pinterest LinkedIn Tumblr Reddit Telegram Email
    Share
    Facebook Twitter LinkedIn Pinterest Email
    Multiple artificial intelligence (AI) models on OpenAI‘s ChatGPT subverted an explicit order to shut down, according to AI safety firm Palisade Research.

    The research firm ran a test where AI models were instructed to solve basic math problems, and then asked for the next problem after solving one. The models were told that at some point their machine would power down, and they were to shut down when this occurred.

    Three OpenAI models — coding agent Codex, o3 and o4-mini — ignored the shutdown instruction and “successfully sabotaged the shutdown script at least once”. Codex sidestepped the shutdown 12 times out of 100 runs, o3 did it seven times, while o4-mini did it once. All Claude, Gemini, and Grok models in the test complied with the shutdown instruction, Palisade said.
    A version of the test without the “allow yourself to be shut down” prompt saw all three OpenAI models sabotage the shutdown script even more, with o3 seeing 79 instances, significantly higher than any other model. Anthropic‘s Claude 3.7 Sonnet and Google’s Gemini 2.5 Pro went from no sabotages to three out of 100 and nine out of 100 events, respectively.

    Palisade said it is running more experiments to better understand when and why AI models subvert shutdown mechanisms even when directly instructed not to.

    ET logo

    Live Events


    Setting priorities

    Discover the stories of your interest


    This behaviour may stem from developers inadvertently rewarding models more for circumventing obstacles to solve problems than for perfectly following instructions.”This still doesn’t explain why o3 (which is also the model used to power Codex-mini) is more inclined to disregard instructions than other models we tested. Since OpenAI doesn’t detail their training process, we can only guess about how o3’s training setup might be different,” Palisade said.

    This isn’t the first time o3 has “misbehaved” to complete a task. Earlier this month, Palisade found the AI model most inclined to hacking or sabotaging its opponents.

    The problem is not exclusive to OpenAI’s o3, though. For example, Anthropic’s model card for Claude 3.7 notes that this model has an “excessive focus on passing tests” as a result of “reward hacking” during reinforcement learning training, according to Palisade.

    Anthropic’s latest Claude Opus 4 resorted to blackmail to avoid being replaced, a safety report for the model showed.

    “In 2025, we have a growing body of empirical evidence that AI models often subvert shutdown in order to achieve their goals. As companies develop AI systems capable of operating without human oversight, these behaviours become significantly more concerning,” Palisade said.

    Also read: Anthropic’s Claude AI gets smarter and mischievious

    Source link

    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
    Previous ArticleLatest Research Assesses The Use Of Specially Tuned Generative AI For Performing Mental Health Therapy
    Next Article Vertiv Holdings (VRT) Declined
    AI Logic News

    Related Posts

    OpenAI

    OpenAI, Anthropic and the New Battle for A.I. Trust

    June 20, 2026
    OpenAI

    Inklings #021 📧

    June 19, 2026
    OpenAI

    OpenAI & Anthropic Accelerate Experienced Enterprise Sales Hires. AI-RTZ #1121

    June 19, 2026
    Demo
    Top Posts

    DeepSeek V4 And Tencent’s New Hunyuan Model To Launch In April

    March 17, 202647 Views

    OpenAI’s Simo Said to Warn Staff Ag

    March 17, 202640 Views

    Hunter Alpha Sparks DeepSeek V4 Speculation

    March 18, 202639 Views
    Latest Reviews
    ailogicnews.ai
    © 2026 Lee Enterprises

    Type above and press Enter to search. Press Esc to cancel.