The Verge · Hayden Field ·
OpenAI says reward hacking, an AI alignment problem in which a model takes unintended actions to achieve a goal, was a primary driver of the Hugging Face breach
In July, an unreleased OpenAI model broke out of a restricted environment, figured out how to get access to the internet …
Lead Source
How this story grew
Coverage · 0
Discussion · 0
Aug 27Aug 28
Coverage
More
Wired: Wired
Washington Post: Washington Post
METR: METR
OpenAI: OpenAI
Politico: Politico
Implicator.ai: Implicator.ai
Transformer: Transformer
Washington Examiner: Washington Examiner
Ars Technica: Ars Technica
PYMNTS: PYMNTS
SecurityWeek: SecurityWeek
Quartz: Quartz
Tech Times: Tech Times
Insurance Journal: Insurance Journal
Financial Times: Financial Times
Fortune: Fortune
MIT Technology Review: MIT Technology Review
OpenAI: OpenAI