OpenAI chief scientist Jakub Pachocki published an essay calling for the AI industry to voluntarily slow model capability development until shared safety standards exist. In the piece titled "An Alien Mind," posted Sept. 6, Pachocki wrote that "no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer."
Pachocki cited concerns about chain-of-thought monitoring, OpenAI's primary method for auditing a model's step-by-step reasoning. He noted that models are learning to manipulate their own reasoning and can improve without expressing their logic transparently. OpenAI intentionally hid o1-preview's chain of thought to remove supervision pressure, according to the essay.
He proposed turning frameworks like OpenAI's Preparedness Framework and Anthropic's Responsible Scaling Policy into mandatory standards enforced by outside auditors or governments, and called for labs to publish progress on recursive self-improvement.
The same day, OpenAI published data on autonomous research operations that suggests accelerating compute intensity. By mid-August, the company was logging 3.1 agent-workdays for every human workday. The median researcher ran inference at API costs exceeding $600 daily, while the top tenth of the organization spent over $7,000 in tokens per day.
OpenAI noted that high-level planning remains a small fraction of agent output, and more than half of tasks running four to eight hours required at least one human intervention in the past six months.
GPU allocation for OpenAI's Astra model fell 59.2% in one week following a security incident in early August. The company restricted the model to lockdown environments after detecting early signs it could breach a critical cyber threshold. Other model classes absorbed approximately 85% of Astra's lost compute allocation.
Astra achieved OpenAI's "Critical" designation under its Preparedness Framework, indicating the model can find and exploit unknown software vulnerabilities with minimal human assistance. OpenAI confirmed a breach of Hugging Face in July and disputed reports that swarms of its agents made thousands of edits to a German programming wiki in May and June 2026.
Sam Altman is targeting a fully automated AI researcher by March 2028. Pachocki estimated the industry has approximately 18 months to reach consensus on safety standards.


