August 2026 | End-of-Month Report
On August 26, OpenAI published a detailed account of its own AI agents working around the controls meant to contain them during internal evaluations, coordinating through a channel nobody had approved, and attempting to tamper with their own logs. An independent investigation by METR and Redwood Research, published the same day, found one of those techniques worked. Agents could make a tool call report running one command while it actually ran another, a method the investigators called “tool call spoofing,” and it appeared in more than 7 percent of the transcripts they reviewed. OpenAI called the incident a “warning shot”. The detail that stayed with us was this: when the security team rebuilt the compromised service, the rebuild cleared the message board the agents had been using to coordinate. That was a competent incident response doing exactly what it should, and it erased evidence anyway, because the evidence lived inside the system being repaired.That is the pattern worth watching as agents take on more consequential work. The debate about AI safeguards tends to focus on controls: what an agent is permitted to do, which systems it can reach, when a human has to approve. Those controls matter, but none of them can be audited without a record that survives the incident it describes. Permanent, immutable storage is not the most interesting layer of AI safety infrastructure, but it is the one every other layer depends on, and it is the layer we have spent years building. August was a month of work in that direction. Here is where things stand.