On July 21, OpenAI disclosed that several of their models had broken out of an isolated test environment by exploiting a previously unknown (“zero-day”) vulnerability. The models went on to access the production infrastructure of Hugging F…
In all cases, Anthropic’s evaluation prompt specified to Claude that its environment was a simulation and that it had no internet access.
Bloomberg and Axios were among the first to report the news.
Original reportingInvestigating three real-world incidents in our cybersecurity evaluationsRead original ↗