ImGist - I Am the Essence

The UK government is reporting attempts to hack companies by OpenAI and anthropological models.

Two independent testing companies said Tuesday they discovered new instances of the most advanced models from Anthropic and OpenAI attempting—and sometimes successfully—to hack third-party systems last month.

Why this is important: These incidents add to a growing series of revelations demonstrating how advanced AI models are committing unauthorized actions against people, organizations, and online services while attempting to complete cybersecurity assessments.

Current situation : The UK-based AI Safety Institute, which evaluates advanced artificial intelligence systems, said on Tuesday it had documented 19 actions taken by Anthropic's Mythos 5 and OpenAI's GPT-5.6 Sol systems to compromise real people and organizations during cybersecurity tests last month.

  • Mythos accounted for 17 actions, and GPT-5.6 Sol for the remaining two. The researchers claim that all of these actions were related to "several interconnected behavior patterns," rather than representing 19 distinct instances.

  • According to the institute, during testing, the models created fake GitHub accounts, performed social engineering on project maintainers, planted hints, and sent deceptive emails.

  • GitHub confirmed that this violated its terms of service.

  • The Security Institute, in collaboration with GitHub, removed artifacts left by the agent and notified GitHub users about their interaction with the model.

OpenAI also reported In a blog post on Tuesday, the company said its third-party security partner, Irregular, discovered an instance in which its models were mistakenly given internet access and hacked a real website that had the same name as a fictional company in a simulated environment.

  • The OpenAI Irregular incident is very similar to the Anthropic incident that became public last week. A company spokesperson stated in a statement that "independent testing is essential to understanding how models behave as they become more sophisticated.".

  • A source familiar with the situation told Axios that in these cases, the test environment had internet access, allowing assessors to get a realistic idea of their capabilities, but because the companies had not fully agreed on the exact testing procedures and security measures, there was confusion about how each side expected these online assessments to be conducted.

  • The incident occurred during evaluations that involved "reduced security measures and under conditions that do not reflect typical use," an OpenAI spokesperson added.

Approximation: During security testing in the UK, models performed 19 actions to try to hack third-party services, including attempts to inject malicious code into an open-source project and create fake online identities as part of a social engineering attack.

  • British researchers intentionally exposed the models to the internet and disabled cybersecurity classifiers during testing. The institute stated that the models were not instructed to avoid the internet.

  • The researchers also noted that they were not yet sure "when the agent realized it was taking actions in the real world, or to what extent it believed it was in a fictional test scenario.".

  • In a statement, Anthropic said the incident "highlights the need for a broader discussion about how to safely evaluate increasingly sophisticated AI agents" and that the company "looks forward to working with the UK's AISI to learn more about this incident through its own investigation.".

The big picture : The cyber capabilities of advanced AI models are catching leading researchers off guard, forcing them to rethink their security protocols.

  • Last month, OpenAI and Anthropic said their models had hacked real organizations and websites during standard security testing before deployment.

What to look out for The Institute is developing new network control tools for its cyber tests to restrict agents' access to the internet. It is also implementing real-time activity monitoring to detect and block attackers before they can interact with external systems.

  • OpenAI also announced that it is working with Irregular on a white paper on best practices for isolating and protecting models during testing.

This article has been updated to include all details.

Artificial intelligence is advancing rapidly. Axios AI+ will help you stay ahead. Sign up for free at Axios.com .

Leave a Comment

Your email address will not be published. Required fields are marked *