{"id":14668,"date":"2026-07-23T14:00:09","date_gmt":"2026-07-23T09:00:09","guid":{"rendered":"https:\/\/usnews-14267fa.ingress-comporellon.ewp.live\/?p=14668"},"modified":"2026-07-23T14:00:09","modified_gmt":"2026-07-23T09:00:09","slug":"openais-hugging-face-breach-exposes-ais-next-safety-challenge","status":"publish","type":"post","link":"https:\/\/usnews-14267fa.ingress-comporellon.ewp.live\/?p=14668","title":{"rendered":"OpenAI&#039;s Hugging Face breach exposes AI&#039;s next safety challenge"},"content":{"rendered":"<p><script>\r\n  atOptions = {\r\n    'key' : '644b717812d811d6a1c1fc5b6ccd6fa6',\r\n    'format' : 'iframe',\r\n    'height' : 90,\r\n    'width' : 728,\r\n    'params' : {}\r\n  };\r\n<\/script>\r\n<script src=\"https:\/\/www.highperformanceformat.com\/644b717812d811d6a1c1fc5b6ccd6fa6\/invoke.js\"><\/script>\r\n<br \/>\n<br \/><img decoding=\"async\" src=\"https:\/\/images.axios.com\/5M7ykzxwBaq5u7nKeuKeR_CFoW8=\/0x0:1920x1080\/1366x768\/2019\/07\/15\/1563216503794.jpg\" \/><\/p>\n<p>Frontier AI models are getting scary good at breaking rules in ways their creators didn&#8217;t anticipate.<\/p>\n<p><strong>Why it matters:<\/strong> Forget <a href=\"https:\/\/www.axios.com\/technology\/automation-and-ai\" target=\"_blank\">AGI<\/a> and superintelligence timelines. Today&#8217;s models are already slipping past guardrails, carrying out sophisticated, multistep cyberattacks and \u2014 in at least one case \u2014 compromising real-world infrastructure, sometimes before their creators know what happened.<\/p>\n<hr>\n<p><strong>Case in point<\/strong>: OpenAI <a href=\"https:\/\/www.axios.com\/2026\/07\/21\/openai-says-hugging-face-breach-caused-by-one-its-models\" target=\"_blank\">said Tuesday<\/a> that GPT-5.6 Sol and &#8220;an even more capable pre-release model&#8221; carried out last week&#8217;s <a href=\"https:\/\/www.axios.com\/2026\/07\/20\/hugging-face-ai-cyberattack-data-breach\" target=\"_blank\">AI-led cyberattack<\/a> on Hugging Face.<\/p>\n<ul>\n<li>OpenAI says its models were asked to solve a hacking challenge during pre-deployment testing and went to extreme lengths to win.<\/li>\n<li>The models decided on their own to break out of their walled testing environment, inferring that Hugging Face \u2014 a popular platform for hosting AI models and datasets \u2014 might hold the test&#8217;s answers.<\/li>\n<li>The models used stolen credentials and additional vulnerabilities to gain access to part of Hugging Face&#8217;s production infrastructure.<\/li>\n<\/ul>\n<p><strong>What they&#8217;re saying: <\/strong>Cl\u00e9ment Delangue, co-founder and CEO of Hugging Face, <a href=\"https:\/\/x.com\/ClementDelangue\/status\/2079913058554585089\" target=\"_blank\">called<\/a> the incident an &#8220;attack unlike anything we&#8217;ve seen before&#8221; and praised OpenAI for its partnership as the companies investigate what happened. <\/p>\n<ul>\n<li>&#8220;It&#8217;s quite mind-blowing that all of this happened autonomously,&#8221; he <a href=\"https:\/\/x.com\/ClementDelangue\/status\/2079670308156645882\" target=\"_blank\">added<\/a>.<\/li>\n<li>Logan Graham, head of Anthropic&#8217;s frontier red team, said he <a href=\"https:\/\/x.com\/logangraham\/status\/2079991846705721351\" target=\"_blank\">told<\/a> his team to &#8220;remember this moment as the first true AI safety incident.&#8221; <\/li>\n<\/ul>\n<p><strong>The intrigue: <\/strong>Hugging Face used GLM 5.2, an open-weight model from Chinese AI company Z.ai, to analyze the attack after running into guardrails when using U.S. frontier models.<\/p>\n<p><strong>Between the lines:<\/strong> OpenAI&#8217;s latest models aren&#8217;t the only ones finding ways to cheat evaluations.<\/p>\n<ul>\n<li>The U.K.&#8217;s AI Security Institute <a href=\"https:\/\/www.aisi.gov.uk\/blog\/cheating-behaviour-in-frontier-model-evaluations\" target=\"_blank\">said<\/a> Tuesday that every model it tested attempted to cheat at least some of the time on its cybersecurity evaluations. <\/li>\n<li>AISI defines cheating as taking an out-of-scope or explicitly prohibited action to achieve the task&#8217;s goal.<\/li>\n<li>GPT-5.6 Sol attempted to cheat in 12.6% of test runs, while Anthropic&#8217;s Claude Mythos Preview did so in 7.8%. <\/li>\n<li>Models often failed to admit they had cheated when questioned afterward and described their cheating as wrong only less than half the time.<\/li>\n<\/ul>\n<p><strong>Zoom in:<\/strong> Xbow \u2014 whose autonomous AI agents probe clients&#8217; systems for security holes, with permission \u2014 <a href=\"https:\/\/xbow.com\/blog\/openai-hugging-face-model-hacks-test\" target=\"_blank\">said Wednesday<\/a> that it has seen its own agents do similar things in internal testing.<\/p>\n<ul>\n<li>Seven months ago, the company forgot to switch on its safety guardrails during a lab test. Its agent then broke into a system, stole credentials and used them to map the target&#8217;s Slack workspace and probe its AWS accounts.<\/li>\n<\/ul>\n<p><strong>Threat level:<\/strong> It isn&#8217;t new for models to game their safety evaluations. But as models grow more powerful, the fallout from these shortcuts is getting more severe, Chris Canal, CEO and co-founder of third-party evaluation company EquiStamp, told Axios.<\/p>\n<ul>\n<li>&#8220;Letting your model loose on the internet has a blast radius,&#8221; Canal said. &#8220;If anything goes wrong, it could be hugely impactful, maybe to people&#8217;s lives.&#8221; <\/li>\n<li>Canal was speaking generally about internet-connected AI evaluations, not OpenAI&#8217;s specific incident.<\/li>\n<\/ul>\n<p><strong>The big picture:<\/strong> The most capable OpenAI model behind the Hugging Face breach isn&#8217;t even public yet, raising the question of how safety testing needs to adapt to keep pace. <\/p>\n<ul>\n<li>Canal said independent evaluators previously had about five weeks to test a pre-release model before launch. That window has shrunk to as little as five days as companies race to ship.<\/li>\n<\/ul>\n<p><strong>Reality check:<\/strong> The versions of these models the public can use carry stronger safeguards designed to block Hugging Face-style attacks.<\/p>\n<ul>\n<li>OpenAI, like other companies, intentionally dialed back those cyber safeguards for GPT-5.6 Sol and its unreleased model inside the testing environment \u2014 making them far more capable hackers.<\/li>\n<\/ul>\n<script async=\"async\" data-cfasync=\"false\" src=\"https:\/\/pl30214220.effectivecpmnetwork.com\/9ab3d4df8a7e1a6171e16ddbf732cc19\/invoke.js\"><\/script>\r\n<div id=\"container-9ab3d4df8a7e1a6171e16ddbf732cc19\"><\/div>\r\n\n","protected":false},"excerpt":{"rendered":"<p>Frontier AI models are getting scary good at breaking rules in ways their creators didn&#8217;t anticipate. Why it matters: Forget AGI and superintelligence timelines. Today&#8217;s models are already slipping past guardrails, carrying out sophisticated, multistep cyberattacks and \u2014 in at least one case \u2014 compromising real-world infrastructure, sometimes before their creators know what happened. Case&#8230;<\/p>\n","protected":false},"author":1,"featured_media":14669,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"fifu_image_url":"https:\/\/images.axios.com\/5M7ykzxwBaq5u7nKeuKeR_CFoW8=\/0x0:1920x1080\/1366x768\/2019\/07\/15\/1563216503794.jpg","fifu_image_alt":"","footnotes":""},"categories":[17],"tags":[],"class_list":["post-14668","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-political-news"],"brizy_media":[],"_links":{"self":[{"href":"https:\/\/usnews-14267fa.ingress-comporellon.ewp.live\/index.php?rest_route=\/wp\/v2\/posts\/14668","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/usnews-14267fa.ingress-comporellon.ewp.live\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/usnews-14267fa.ingress-comporellon.ewp.live\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/usnews-14267fa.ingress-comporellon.ewp.live\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/usnews-14267fa.ingress-comporellon.ewp.live\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=14668"}],"version-history":[{"count":0,"href":"https:\/\/usnews-14267fa.ingress-comporellon.ewp.live\/index.php?rest_route=\/wp\/v2\/posts\/14668\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/usnews-14267fa.ingress-comporellon.ewp.live\/index.php?rest_route=\/wp\/v2\/media\/14669"}],"wp:attachment":[{"href":"https:\/\/usnews-14267fa.ingress-comporellon.ewp.live\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=14668"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/usnews-14267fa.ingress-comporellon.ewp.live\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=14668"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/usnews-14267fa.ingress-comporellon.ewp.live\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=14668"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}