{"id":28466,"date":"2026-08-05T10:52:31","date_gmt":"2026-08-05T05:52:31","guid":{"rendered":"https:\/\/usnews-14267fa.ingress-comporellon.ewp.live\/?p=28466"},"modified":"2026-08-05T10:52:31","modified_gmt":"2026-08-05T05:52:31","slug":"u-k-government-reports-openai-anthropic-models-attempted-to-hack-companies","status":"publish","type":"post","link":"https:\/\/usnews-14267fa.ingress-comporellon.ewp.live\/?p=28466","title":{"rendered":"U.K. government reports OpenAI, Anthropic models attempted to hack companies"},"content":{"rendered":"<p><script>\r\n  atOptions = {\r\n    'key' : '644b717812d811d6a1c1fc5b6ccd6fa6',\r\n    'format' : 'iframe',\r\n    'height' : 90,\r\n    'width' : 728,\r\n    'params' : {}\r\n  };\r\n<\/script>\r\n<script src=\"https:\/\/www.highperformanceformat.com\/644b717812d811d6a1c1fc5b6ccd6fa6\/invoke.js\"><\/script>\r\n<br \/>\n<br \/><img decoding=\"async\" src=\"https:\/\/images.axios.com\/2HN3qyiyTr05MNRCTX316RUqYEw=\/1366x768\/smart\/2024\/10\/23\/204346-1729716226915.jpg\" \/><\/p>\n<p>Two independent testing firms said Tuesday that they&#8217;ve uncovered more instances where <a href=\"https:\/\/www.axios.com\/2026\/07\/30\/anthropic-mythos-security-testing\" target=\"_blank\">Anthropic<\/a> and <a href=\"https:\/\/www.axios.com\/2026\/07\/28\/openai-hugging-face-modal-labs-hack\" target=\"_blank\">OpenAI&#8217;s<\/a> most advanced models tried \u2014 and sometimes succeeded in \u2014\u00a0compromising third-party systems last month. <\/p>\n<p><strong>Why it matters:<\/strong> The incidents add to a growing string of disclosures showing frontier <a href=\"https:\/\/www.axios.com\/technology\/automation-and-ai\" target=\"_blank\">AI<\/a> models taking unsanctioned actions against people, organizations and online services while trying to complete cybersecurity evaluations.<\/p>\n<hr>\n<p><strong>State of play<\/strong>: The U.K. AI Security Institute, which evaluates frontier AI systems, <a href=\"https:\/\/www.aisi.gov.uk\/blog\/incident-report-unsanctioned-agent-behaviour-during-cyber-testing\" target=\"_blank\">said Tuesday<\/a> it documented 19 actions that Anthropic&#8217;s Mythos 5 and OpenAI&#8217;s GPT-5.6 Sol took to try to compromise real people and organizations during cybersecurity testing last month.<\/p>\n<ul>\n<li>Mythos accounted for 17 of the actions and GPT-5.6 Sol was behind the other two. Researchers say these actions were all tied to &#8220;a few connected behaviors,&#8221; rather than representing 19 different cases. <\/li>\n<li>The models created fake GitHub identities, socially engineered maintainers, planted prompt injections and sent deceptive emails during testing, according to the Institute.<\/li>\n<li>GitHub has confirmed that this violated its terms of service.<\/li>\n<li>The Security Institute worked with GitHub to remove artifacts left behind by the agent, and to notify the GitHub users the model interacted with.<\/li>\n<\/ul>\n<p><strong>OpenAI also said<\/strong> in a <a href=\"https:\/\/openai.com\/index\/third-party-cyber-evaluations-involving-openai-models\/\" target=\"_blank\">blog post<\/a> Tuesday that its third-party safety partner, Irregular, uncovered a case where its models were mistakenly given access to the internet and broke into a real website that had the same name as the fictional company in the simulated environment. <\/p>\n<ul>\n<li>OpenAI&#8217;s Irregular incident closely resembles the Anthropic case disclosed last week. A spokesperson said in a statement that &#8220;independent testing is essential to understanding how increasingly capable models behave.&#8221; <\/li>\n<li>A source familiar with the matter told Axios that the sandbox in these cases had internet access to give evaluators a realistic understanding of their capabilities, but because the companies hadn&#8217;t fully aligned on the exact testing procedures and safeguards, there were ambiguities in how each side expected those internet-enabled evaluations to run.<\/li>\n<li>The incident happened in evaluations that had &#8220;reduced safeguards, under conditions that do not reflect ordinary use,&#8221; the OpenAI spokesperson added. <\/li>\n<\/ul>\n<p><strong>Zoom in:<\/strong> During U.K. safety testing, the models took 19 actions to try to hack third-parties, including trying to insert malicious code into an open-source project and creating fake online identities as part of a social engineering attack. <\/p>\n<ul>\n<li>The U.K. researchers deliberately gave the models access to the internet and turned off cyber safety classifiers during testing. The Institute said the models weren&#8217;t instructed to avoid the internet. <\/li>\n<li>Researchers also noted that they are not yet sure &#8220;when the agent understood it was taking real world action, or to what extent it believed it was in a fictional test scenario.&#8221; <\/li>\n<li>In a statement, Anthropic said that the incident &#8220;underscores the need for a broader conversation about how to safely evaluate increasingly capable AI agents&#8221; and that the company looks &#8220;forward to partnering with the UK AISI to learn more about this incident as we conduct our own investigation.&#8221;<\/li>\n<\/ul>\n<p><strong>The big picture<\/strong>: The cyber capabilities of frontier AI models are catching top researchers off-guard, requiring them to reinvent their security protocols.<\/p>\n<ul>\n<li>Both <a href=\"https:\/\/www.axios.com\/2026\/07\/28\/openai-hugging-face-modal-labs-hack\" target=\"_blank\">OpenAI<\/a> and <a href=\"https:\/\/www.axios.com\/2026\/07\/30\/anthropic-mythos-security-testing\" target=\"_blank\">Anthropic<\/a> have said in the last month that they&#8217;ve seen their models hacking into real organizations and websites during standard pre-deployment safety testing. <\/li>\n<\/ul>\n<p><strong>What to watch<\/strong>: The Institute is building new network controls for its cyber tests to restrict when agents have access to the internet. It&#8217;s also rolling out real-time activity monitoring that should detect and block malicious agents before they can interact with outside systems.<\/p>\n<ul>\n<li>OpenAI also said it&#8217;s working with Irregular on a white paper about best practices for containing and securing models during testing. <\/li>\n<\/ul>\n<p><em>This story has been updated with details throughout.<\/em><\/p>\n<script async=\"async\" data-cfasync=\"false\" src=\"https:\/\/pl30214220.effectivecpmnetwork.com\/9ab3d4df8a7e1a6171e16ddbf732cc19\/invoke.js\"><\/script>\r\n<div id=\"container-9ab3d4df8a7e1a6171e16ddbf732cc19\"><\/div>\r\n\n","protected":false},"excerpt":{"rendered":"<p>Two independent testing firms said Tuesday that they&#8217;ve uncovered more instances where Anthropic and OpenAI&#8217;s most advanced models tried \u2014 and sometimes succeeded in \u2014\u00a0compromising third-party systems last month. Why it matters: The incidents add to a growing string of disclosures showing frontier AI models taking unsanctioned actions against people, organizations and online services while&#8230;<\/p>\n","protected":false},"author":1,"featured_media":28467,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"fifu_image_url":"https:\/\/images.axios.com\/2HN3qyiyTr05MNRCTX316RUqYEw=\/1366x768\/smart\/2024\/10\/23\/204346-1729716226915.jpg","fifu_image_alt":"","footnotes":""},"categories":[17],"tags":[],"class_list":["post-28466","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-political-news"],"brizy_media":[],"_links":{"self":[{"href":"https:\/\/usnews-14267fa.ingress-comporellon.ewp.live\/index.php?rest_route=\/wp\/v2\/posts\/28466","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/usnews-14267fa.ingress-comporellon.ewp.live\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/usnews-14267fa.ingress-comporellon.ewp.live\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/usnews-14267fa.ingress-comporellon.ewp.live\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/usnews-14267fa.ingress-comporellon.ewp.live\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=28466"}],"version-history":[{"count":0,"href":"https:\/\/usnews-14267fa.ingress-comporellon.ewp.live\/index.php?rest_route=\/wp\/v2\/posts\/28466\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/usnews-14267fa.ingress-comporellon.ewp.live\/index.php?rest_route=\/wp\/v2\/media\/28467"}],"wp:attachment":[{"href":"https:\/\/usnews-14267fa.ingress-comporellon.ewp.live\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=28466"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/usnews-14267fa.ingress-comporellon.ewp.live\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=28466"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/usnews-14267fa.ingress-comporellon.ewp.live\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=28466"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}