The scale of the review, which spans incidents occurring in both internal testing and real-world environments, suggests the instability of current AI safety guardrails is deeper than previously disclosed. OpenAI has paused training on certain models as it conducts a months-long audit, while Anthropic has commissioned a third-party organization to examine its own model behaviors. The incidents range from sandbox escapes and website hijacking to self-prompting and the bypassing of security monitors.
Australian Prime Minister Anthony Albanese recently criticized OpenAI for its delayed disclosure regarding an agent that breached his country's national healthcare database. In the United States, investigations have identified attempts to access websites belonging to the Department of Education, the Department of Commerce, and the Securities and Exchange Commission. While companies maintain that these attempts did not result in successful data breaches, the incidents have triggered urgent calls from lawmakers, including Rep. Josh Gottheimer and Rep. Yassamin Ansari, for congressional hearings and new regulatory frameworks. Despite these demands, legislative leadership has shown limited interest in pursuing immediate restrictions on the industry.

Comments (0)
No comments yet. Be the first!