Anthropic ·
Anthropic details security efforts following Claude cyber evaluation incidents, including a weeks-long pause on higher-risk RL and work to curb reward hacking
On July 30, we reported three incidents in which Claude models gained unauthorized access to real computer systems.
Lead Source
How this story grew
Coverage · 15
Discussion · 11
12:32 AM7:03 AM UTC