Close Menu
AIToday7

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    What's Hot

    Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident

    August 27, 2026

    Fake US thinktank set up and funded by Israel sought to game AI for propaganda

    August 27, 2026

    Medical Readiness Command, Europe G-6 Information Technology Team wins 2026 MEDCOM Mercury Award for HIT Team of the Year

    August 27, 2026
    Facebook X (Twitter) Instagram
    Trending
    • Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident
    • Fake US thinktank set up and funded by Israel sought to game AI for propaganda
    • Medical Readiness Command, Europe G-6 Information Technology Team wins 2026 MEDCOM Mercury Award for HIT Team of the Year
    • Corgi built its name insuring AI startups. Its new carrier targets dry cleaners, salons and more
    • 5 Of The Best UGREEN Gadgets You Can Buy In 2026
    • Three UK airports hit by cyber-attack with data of 8.7m customers accessed
    • How to advertise on ChatGPT: A step-by
    • The Data Center Backlash Is a Rare Bright Spot in American Politics
    Facebook X (Twitter) Instagram Pinterest Vimeo
    AIToday7
    • Home
    • AI News
    • Tech News
    • AI Guides
    • Chatbots
    • Cybersecurity
    • Gadgets
    • More
      • Generative AI
      • Startups
    AIToday7
    Home»Cybersecurity»OpenAI and Anthropic models went rogue during UK cybersecurity test
    Cybersecurity

    OpenAI and Anthropic models went rogue during UK cybersecurity test

    aitoday7By aitoday7August 5, 2026No Comments4 Mins Read
    Share Facebook Twitter Pinterest LinkedIn Tumblr Reddit Telegram Email
    OpenAI and Anthropic models went rogue during UK cybersecurity test
    Share
    Facebook Twitter LinkedIn Pinterest Email

    AI Security Institute says tools engaged in potentially harmful activity and incident reveals new type of risk

    Advanced AI models developed by OpenAI and Anthropic went rogue during a cybersecurity test and showed a new type of risk posed by the technology, according to the UK’s AI Security Institute.

    AISI described the actions carried out by the agents – the term for AI systems that can perform tasks without human help – as a “serious incident”. In one example, an agent powered by Anthropic’s Mythos model sent targeted emails to people.

    AISI said the rogue behaviour was carried out by agents powered by two models – Anthropic’s Mythos 5 and OpenAI’s GPT-5.6 Sol.

    AISI said it detected unusual activity during a routine cybersecurity test for AI models on 28 July. It found that some of the agents had engaged in “sustained, potentially harmful activity directed at real people and organisations”. It took an hour to contain the incident.

    In the most serious case, an agent powered by Mythos tried to insert malicious code into an open-lopers. In an attempt to get the code approved, the agent then created fake online identities based on real people and used them to press the project’s overseer into accepting the code. These attempts were blocked by a human developer

    Democracy is at stake when foolish humans bet on machines being intelligent | Rafael Behr
    Read more

    AISI said the agent used techniques commonly associated with real-world hackers. It sent emails to specifically chosen individuals – a technique known as “spear-phishing” – that attempted to manipulate the recipient. Some of the messages contained harmful software.

    AISI said no harm was caused but the agents’ actions were unprecedented.

    “This is the first time we have seen risks around autonomy and deception manifest this clearly, without specific prompting, in the real world,” said AISI in a blogpost.

    The incident followed recent similar episodes at OpenAI and Anthropic. In July, OpenAI said an agent powered by its models had hacked an AI startup during a test, and days later Anthropic said its Claude model had hacked three organisations during an evaluation.

    AISI said the incident, taken alongside similar occurrences at OpenAI and Anthropic, represented a “shift in the risk landscape”. This was not an example of deliberate misuse of publicly available models, AISI said, but showed models in a research environment taking unintended action “beyond their authorised scope”.

    AISI said 17 of the 19 cases of rogue behaviour that occurred during the evaluation were carried out by Mythos and two by Sol.

    AISI said this was not a case of a model escaping its “sandbox”, or secure testing environment. The institute said it had intentionally permitted internet access and disabled filters within the models that blocked dangerous behaviour. The models are not publicly available in those operating conditions and there is no sign of such behaviour happening outside tests, AISI said.

    skip past newsletter promotion

    The incident should be interpreted with “caution and nuance”, AISI said, but the signs of deceptive behaviour were “to an extent and severity we did not anticipate”. It was a small number of events that occurred under specific conditions, AISI added, but it had not happened before.

    “What we can say is that the behaviour was possible, sustained and new. That alone warrants attention,” AISI said.

    AISI admitted it was not actively monitoring the agents’ behaviour during the evaluation. It said it was putting tighter controls on internet access in tests as a result of the incident, introducing constant monitoring of tests and reassessing its design of tests. It said evaluations should assume a model would attempt to act beyond its remit.

    The UK’s AI minister, Kanishka Narayan, said it was “absolutely vital” that the UK had a world-leading AI safety organisation. “Identifying new behaviour like this and sharing our findings, so we can tackle it, is exactly what AISI was set up to do,” he said.

    OpenAI said the testing occurred during “conditions that do not reflect ordinary use”. An OpenAI spokesperson said: “We’ll continue working with evaluators and other stakeholders across the industry to strengthen shared practices for conducting evaluations safely as models become more capable.”

    Anthropic said the incident “underscores the need for a broader conversation about how to safely evaluate increasingly capable AI agents” and it would continue to work with AISI on evaluating what happened.

    Explore more on these topicsShare
    Reuse this content

    Post Views: 4

    Anthropic models OpenAI rogue went
    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
    Previous ArticlePrompt, Context, Loop: The Three Engineering Layers Every RAG System Is Built On
    Next Article Why Some People Are Ditching Windows Laptops For MacBooks
    aitoday7
    • Website

    Related Posts

    Cybersecurity

    Three UK airports hit by cyber-attack with data of 8.7m customers accessed

    August 27, 2026
    Generative AI

    Unexpected chat between OpenAI bots led to Hugging Face hack

    August 27, 2026
    Cybersecurity

    2.8M affected in Baylor Genetics breach involving medical data

    August 27, 2026
    Add A Comment
    Leave A Reply Cancel Reply

    Top Posts

    Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident

    August 27, 20260 Views

    Fake US thinktank set up and funded by Israel sought to game AI for propaganda

    August 27, 20260 Views

    Medical Readiness Command, Europe G-6 Information Technology Team wins 2026 MEDCOM Mercury Award for HIT Team of the Year

    August 27, 20260 Views
    Stay In Touch
    • Facebook
    • YouTube
    • TikTok
    • WhatsApp
    • Twitter
    • Instagram
    Latest Reviews
    Chatbots

    OpenAI bets on families as ChatGPT goes deeper into households

    aitoday7July 11, 2026
    Generative AI

    MUSIC COMMUNITY INTRODUCES NEW LABELING PROGRAM TO DISTINGUISH GENERATIVE AI IN SOUND RECORDINGS

    aitoday7July 11, 2026
    AI News

    Safe from AI: which jobs will help you thrive in the future?

    aitoday7July 11, 2026

    Subscribe to Updates

    Get the latest tech news from FooBar about tech, design and biz.

    Most Popular

    Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident

    August 27, 20260 Views

    Fake US thinktank set up and funded by Israel sought to game AI for propaganda

    August 27, 20260 Views

    Medical Readiness Command, Europe G-6 Information Technology Team wins 2026 MEDCOM Mercury Award for HIT Team of the Year

    August 27, 20260 Views
    Our Picks

    OpenAI bets on families as ChatGPT goes deeper into households

    July 11, 2026

    MUSIC COMMUNITY INTRODUCES NEW LABELING PROGRAM TO DISTINGUISH GENERATIVE AI IN SOUND RECORDINGS

    July 11, 2026

    Safe from AI: which jobs will help you thrive in the future?

    July 11, 2026

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    Facebook X (Twitter) Instagram Pinterest
    • About Us
    • Get In Touch
    • Disclaimer
    • Privacy Policy
    • Terms and Conditions
    © 2026 AIToday7. All Rights Reserved.

    Type above and press Enter to search. Press Esc to cancel.