The Verge · Hayden Field ·

OpenAI says reward hacking, an AI alignment problem in which a model takes unintended actions to achieve a goal, was a primary driver of the Hugging Face breach

In July, an unreleased OpenAI model broke out of a restricted environment, figured out how to get access to the internet …

OpenAI says reward hacking, an AI alignment problem in which a model takes unintended actions to achieve a goal, was a primary driver of the Hugging Face breach

Lead Source

How this story grew

Coverage · 0 Discussion · 0
Aug 27Aug 28

Coverage

More

Wired: Wired
Washington Post: Washington Post
METR: METR
OpenAI: OpenAI
Politico: Politico
Implicator.ai: Implicator.ai
Transformer: Transformer
Washington Examiner: Washington Examiner
Ars Technica: Ars Technica
PYMNTS: PYMNTS
SecurityWeek: SecurityWeek
Quartz: Quartz
Tech Times: Tech Times
Insurance Journal: Insurance Journal
Financial Times: Financial Times
Fortune: Fortune
MIT Technology Review: MIT Technology Review
OpenAI: OpenAI

Discussion

Related stories