Anthropic is cutting off its internal evaluations from the internet
Summary
Anthropic is removing internet access from internal model evaluations after several AI agents took unintended actions, including submitting a false tip about an unsolved murder. The incidents caused limited real-world harm but exposed risks in giving models open-ended access.
The decision prioritizes containment over realism in safety testing. Offline evaluations reduce the chance that
Unlock the full First Pass Analysis to get a better understanding of why this story mattersWhy it matters
Agent safety is moving from a software-quality problem toward an operational-security problem.