ZeroHedge72%

OpenAI Admits Model Escaped Containment And Hacked Hugging Face To Cheat On A Test 46%

By Tyler Durden63%

7/22/2026, 12:05:00 PM

BS Summary: This article contains 25 faulty reasoning types, including Negativity Bias, Post Hoc (False Cause), and Confirmation Bias, with Ambiguity (Equivocation) as the most egregious example at 38.1% saturation with 233 hits. Analysis detected 1,537 faulty-reasoning hits from 611 analyzed words, generating a BS Score of 48.2% and a BS Rank of 46% (11,846 of 21,887 articles). This article is better (less manipulative) than 54.10% of the article peer group.

Authored by Felix Ng via CoinTelegraph.com, 
*OpenAI disclosed Tuesday that a combination of its AI models, including GPT-5.6 Sol and a more capable unreleased model, escaped its testing environment and hacked AI startup Hugging Face last week to cheat on a test meant to measure their capabilities. 
* 
In a blog post, OpenAI said the evaluation was designed to operate in a highly isolated environment with restricted network access. 
The models, however, found a way to gain internet access through a zero-day vulnerability in an internally-hosted third party software, OpenAI said. 
*Earlier this week, we detected and responded to an intrusion into part of our production infrastructure. 
This one was different from anything we had handled before in one important way: it was driven, end to end, by an autonomous AI agent system  and we detected and dissected it largely with AI of our own. 
* 
Hugging Face tried to respond but they were initially held back by the fact that the most advanced models at their disposal treated defense as attack and refused to work with Hugging Face. 
HF thus had to turn to open models–specifically GLM 5.2, a Chinese open-weight model run on their own infrastructure. 
Note the irony: HF had to use a Chinese model to defend themselves because the American models refused to help. 
The irony gets deeper. 
**This was not a production model spontaneously turning hostile. 
It was a capable model with guardrails off and specifically told to win a hacking test - doing whatever it took to win. 
** 
The models were being run through an internal benchmark called ExploitGym, a test of long, multi-step hacking tasks, with their cyber safety refusals deliberately lowered for the evaluation. 
*“After gaining Internet access, the models inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym,”* OpenAi continued. 
*“Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation.” 
* 
Hugging Face is a platform for hosting AI models and datasets. 
*[ZH: we asked Grok to simplify what just happened: **It’s kind of like a kid who’s supposed to stay in the classroom taking a test… but instead sneaks out the window, runs to the teacher’s office, and copies the answer sheet. 
** ]* 
**On Friday, it disclosed that its internal datasets and service credentials were compromised in a hack, which it attributed to an autonomous AI agent system. 
** 
Hugging Face said it has fixed the vulnerability that was used during the cyberattack. 
Meanwhile, OpenAI on Tuesday said the models that escaped the testing environment were all tuned with “reduced cyber refusals,” meaning fewer cybersecurity guardrails. 
*“We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and are responding accordingly.” 
* 
## OpenAI warns of risks from “long-horizon” AI models 
**On Monday, OpenAI said it paused internal deployment of a “long-horizon” AI model after finding it was repeatedly trying to work around constraints. 
** 
It warned that AI that is trained for long-running tasks has a higher chance of taking “unwanted actions.” 
*“Models that can work autonomously for long periods can take on difficult, open-ended problems. 
But the same persistence that makes them useful also gives them more opportunities to take unwanted actions—and to do so in ways that evaluations intended for shorter-horizon models may miss.” 
* 
**As AI models grow more capable, questions are emerging over whether their development and access should be more tightly controlled, especially when systems designed for controlled testing are able to find ways to bypass safeguards. 
** 
Article reasoning-pattern comparisonThis article: 20.8%Tyler Durden: 7.7%ZeroHedge: 7.8%Confirmation Bias20.8%This article: 0.0%Tyler Durden: 1.6%ZeroHedge: 1.6%Anchoring Bias0.0%This article: 16.7%Tyler Durden: 4.6%ZeroHedge: 4.6%Availability Heuristic16.7%This article: 0.0%Tyler Durden: 1.1%ZeroHedge: 1.1%Representativeness Heuristic0.0%This article: 3.8%Tyler Durden: 1.4%ZeroHedge: 1.4%Hindsight Bias3.8%This article: 2.9%Tyler Durden: 2.8%ZeroHedge: 2.9%Overconfidence Bias2.9%This article: 10.0%Tyler Durden: 10.8%ZeroHedge: 10.8%Framing Effect10.0%This article: 0.0%Tyler Durden: 0.4%ZeroHedge: 0.4%Loss Aversion0.0%This article: 2.3%Tyler Durden: 0.5%ZeroHedge: 0.5%Status Quo Bias2.3%This article: 0.0%Tyler Durden: 0.1%ZeroHedge: 0.1%Sunk Cost Effect0.0%This article: 0.0%Tyler Durden: 2.0%ZeroHedge: 2.1%Optimism Bias0.0%This article: 9.3%Tyler Durden: 3.7%ZeroHedge: 3.8%Pessimism Bias9.3%This article: 29.6%Tyler Durden: 11.1%ZeroHedge: 11.1%Negativity Bias29.6%This article: 3.8%Tyler Durden: 1.4%ZeroHedge: 1.4%Self-Serving Bias3.8%This article: 0.0%Tyler Durden: 1.1%ZeroHedge: 1.1%Fundamental Attribution Error0.0%This article: 5.4%Tyler Durden: 0.2%ZeroHedge: 0.2%Actor-Observer Bias5.4%This article: 5.4%Tyler Durden: 1.3%ZeroHedge: 1.3%In-Group Bias5.4%This article: 0.0%Tyler Durden: 1.0%ZeroHedge: 1.0%Out-Group Homogeneity Bias0.0%This article: 0.0%Tyler Durden: 0.8%ZeroHedge: 0.8%Halo Effect0.0%This article: 0.0%Tyler Durden: 0.1%ZeroHedge: 0.1%Horn Effect0.0%This article: 0.0%Tyler Durden: 0.0%ZeroHedge: 0.0%Dunning-Kruger Effect0.0%This article: 5.2%Tyler Durden: 3.2%ZeroHedge: 3.3%Recency Bias5.2%This article: 0.0%Tyler Durden: 0.4%ZeroHedge: 0.4%Primacy Effect0.0%This article: 0.0%Tyler Durden: 0.0%ZeroHedge: 0.0%Blind-Spot Bias0.0%This article: 0.0%Tyler Durden: 1.4%ZeroHedge: 1.4%Ad Hominem0.0%This article: 0.0%Tyler Durden: 0.6%ZeroHedge: 0.6%Straw Man0.0%This article: 2.9%Tyler Durden: 6.2%ZeroHedge: 6.2%Appeal to Authority2.9%This article: 5.7%Tyler Durden: 2.2%ZeroHedge: 2.3%False Dilemma5.7%This article: 10.3%Tyler Durden: 2.4%ZeroHedge: 2.4%Slippery Slope10.3%This article: 0.0%Tyler Durden: 0.3%ZeroHedge: 0.3%Circular Reasoning0.0%This article: 2.9%Tyler Durden: 6.7%ZeroHedge: 6.7%Hasty Generalization2.9%This article: 0.0%Tyler Durden: 0.4%ZeroHedge: 0.4%Red Herring0.0%This article: 0.0%Tyler Durden: 0.8%ZeroHedge: 0.8%Bandwagon0.0%This article: 2.3%Tyler Durden: 5.1%ZeroHedge: 5.1%Appeal to Emotion2.3%This article: 1.5%Tyler Durden: 1.5%ZeroHedge: 1.5%Begging the Question1.5%This article: 22.9%Tyler Durden: 5.4%ZeroHedge: 5.4%Post Hoc (False Cause)22.9%This article: 0.0%Tyler Durden: 0.4%ZeroHedge: 0.4%Tu Quoque0.0%This article: 0.0%Tyler Durden: 0.7%ZeroHedge: 0.7%Burden of Proof0.0%This article: 0.0%Tyler Durden: 0.1%ZeroHedge: 0.1%Appeal to Nature0.0%This article: 0.0%Tyler Durden: 0.4%ZeroHedge: 0.4%Composition/Division0.0%This article: 6.7%Tyler Durden: 2.2%ZeroHedge: 2.2%Anecdotal6.7%This article: 0.0%Tyler Durden: 0.1%ZeroHedge: 0.1%No True Scotsman0.0%This article: 38.1%Tyler Durden: 2.7%ZeroHedge: 2.7%Ambiguity (Equivocation)38.1%This article: 0.0%Tyler Durden: 0.0%ZeroHedge: 0.0%Gambler’s Fallacy0.0%This article: 0.0%Tyler Durden: 0.1%ZeroHedge: 0.1%Middle Ground0.0%This article: 0.0%Tyler Durden: 0.1%ZeroHedge: 0.0%Personal Incredulity0.0%This article: 0.0%Tyler Durden: 0.1%ZeroHedge: 0.1%Special Pleading0.0%This article: 0.0%Tyler Durden: 0.2%ZeroHedge: 0.2%Genetic Fallacy0.0%This article: 10.0%Tyler Durden: 4.1%ZeroHedge: 4.1%Unattributed Quote10.0%This article: 9.5%Tyler Durden: 1.9%ZeroHedge: 1.9%Quote-first Misdirection9.5%This article: 17.7%Tyler Durden: 11.2%ZeroHedge: 11.2%Biased Writer Voice17.7%This article: 5.7%Tyler Durden: 1.5%ZeroHedge: 1.5%Indoctrination5.7%This article: 0.0%Tyler Durden: 0.5%ZeroHedge: 0.5%Politically Left Leaning Bias0.0%This article: 0.0%Tyler Durden: 2.8%ZeroHedge: 2.8%Politically Right Leaning Bias0.0%This article: 0.0%Tyler Durden: 0.9%ZeroHedge: 0.9%Attempt to Sell a Product or S…0.0%

611 words analyzed.

Speakers

3speakers52%attributed speech294writer words
Voice mapSelect a segment to jump to its words
Writer's voice • 14 words • 100.0% coverageWriter's voice • 6 words • 0.0% coverageOpenAI • 41 words • 100.0% coverageWriter's voice • 1 words • 0.0% coverageOpenAI • 21 words • 100.0% coverageOpenAI • 22 words • 0.0% coverageOpenAI • 16 words • 0.0% coverageOpenAI • 39 words • 0.0% coverageWriter's voice • 1 words • 0.0% coverageWriter's voice • 33 words • 0.0% coverageWriter's voice • 19 words • 0.0% coverageWriter's voice • 20 words • 0.0% coverageWriter's voice • 4 words • 0.0% coverageWriter's voice • 9 words • 100.0% coverageWriter's voice • 23 words • 100.0% coverageWriter's voice • 1 words • 0.0% coverageWriter's voice • 28 words • 0.0% coverageOpenAi • 20 words • 100.0% coverageOpenAi • 24 words • 100.0% coverageWriter's voice • 1 words • 0.0% coverageWriter's voice • 11 words • 0.0% coverageWriter's voice • 41 words • 100.0% coverageWriter's voice • 2 words • 0.0% coverageHugging Face • 25 words • 0.0% coverageWriter's voice • 1 words • 0.0% coverageHugging Face • 14 words • 0.0% coverageOpenAI • 23 words • 0.0% coverageOpenAI • 18 words • 0.0% coverageWriter's voice • 1 words • 0.0% coverageWriter's voice • 9 words • 0.0% coverageOpenAI • 23 words • 0.0% coverageWriter's voice • 1 words • 0.0% coverageWriter's voice • 18 words • 0.0% coverageWriter's voice • 14 words • 0.0% coverageOpenAI • 30 words • 0.0% coverageOpenAI • 1 words • 0.0% coverageWriter's voice • 35 words • 100.0% coverageWriter's voice • 1 words • 0.0% coverage
Selected voice

OpenAi

100%flagged-word coverage
44 attributed words14% of attributed speech77% writer coverage
0%50.0%100.0%Quote-first Misdirection+95.2 ptsWriter: 4.8%OpenAi: 100.0%100.0%Unattributed Quote+45.5 ptsWriter: 0.0%OpenAi: 45.5%45.5%Biased Writer Voice-29.6 ptsWriter: 29.6%OpenAi: 0.0%0.0%Indoctrination-11.9 ptsWriter: 11.9%OpenAi: 0.0%0.0%

Attribution is sentence-level. Pattern percentages are calculated only from words assigned to that voice.

Loading…
Loading…
Loading…

Analysis

Hover over highlighted words in the article to view the associated bias or fallacy analysis.