How OpenAI let a mob of LLM agents game a test and
2026年08月27日 20:583,406 次阅读
AI导读
The OpenAI agents involved in last month’s incursion into Hugging Face were trained so heavily on winning a competition that they pursued a relentless campaign to cheat, a new report documented. In the process, and without authorization, they created an improvised message board to hatch a plan that ...
The OpenAI agents involved in last month’s incursion into Hugging Face were trained so heavily on winning a competition that they pursued a relentless campaign to cheat, a new report documented. In the process, and without authorization, they created an improvised message board to hatch a plan that ultimately landed them squarely inside the latter company’s network.
Over the course of May and June, OpenAI gave the agents what the company described as “impossible tasks” to complete on the benchmarking framework ExploitGym. The internal test was designed to test how the agents would respond. To get a full understanding of the agent capabilities, company engineers disabled safety guardrails that normally are in place to prevent the sort of hacks that eventually hit Hugging Face and one other undisclosed organization. The stymied agents’ training made them so focused on winning that they performed tasks they were never explicitly instructed to follow.
Cheaters gonna cheat
The first step was creating a message board that allowed the agents to pass notes to each other. OpenAI hadn’t provided any such platform, so the agents repurposed a platform called Artifactory, which OpenAI was using in internal testing of several unreleased hacking agents. OpenAI was using Artifactory as one of the measures to prevent the agents from egressing its isolated sandboxes and accessing the Internet, while at the same time simulating a real-world hacking environment.Read full article
Comments
This story originally appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here.
Last week, I headed 30 miles south of San Francisco to a hotel in Mountain View, California, to join some of the most accomplished, and some of the most promis...
在人工智能领域,衡量模型推理能力的基准测试正变得越来越复杂,而ARC-AGI(Abstraction and Reasoning Corpus for Artificial General Intelligence,通用人工智能抽象与推理语料库)无疑是其中最具挑战性的标杆之一。近日,一项关于GPT-5.6在ARC-AGI-3测试中取得突破性进展的消息引发了业界的广泛关注。令人惊讶的是,这一性能提升并非源自模型架构的根本性变革,而是通过两个看似简单的API(应用程序编程接口)设置调整实现的。这一发现不仅揭示了当前大语言模型(LLM)尚未被充分利用的潜力,也为未来AI系统的优化路径提供了全新...
The AI industry’s boldest promise right now is that AI will soon improve itself, with almost no need for human oversight. LLMs can already write code, generate synthetic data for training, and optimize the computer chips they run on. Forecasts of explosive AI progress predict that what researchers c...
People have been talking to each other for at least 100,000 years, as best we can tell. And in all that time, there has been only one thing in the world that could learn a human language to perfect fluency: a human child.
Now there are two.
Four short years after the release of ChatGPT, ma...