top of page

The Monitor Is Failing: OpenAI's Chief Scientist Issues a Call to Slow Down

The Monitor Is Failing: OpenAI's Chief Scientist Issues a Call to Slow Down
The Monitor Is Failing: OpenAI's Chief Scientist Issues a Call to Slow Down

On September 6, 2026, Jakub Pachocki published an essay on OpenAI's website called "An Alien Mind." Pachocki runs research at OpenAI. Three days earlier, the company had shipped GPT-6 Astra, the most capable model it has ever released. His argument is that OpenAI and every other frontier lab should be prepared to slow down.


The essay opens with a scene from mid-2023. Inside a research project called RLSlow, Pachocki and a colleague named Szymon saw the first evidence they could scale the training of reasoning models, giving pretrained systems the ability to form their own chains of thought. He writes that they spent that night at the office not celebrating the benchmark numbers, but working through the realization that they would see machines meaningfully smarter than themselves within their lifetimes, and trying to figure out how to alert people to what was coming.


Three years later, he writes, reasoning models operate computers, collaborate with people and each other, carry out research projects, and are reshaping computer security in ways that create clear new dangers. Based on internal results, he says he expects the current pace to continue into recursive self-improvement, where AI systems increasingly drive their own development.


Grown, Not Built


Pachocki's central technical claim is that progress in machine intelligence comes primarily from increasing computational power. He says OpenAI internalized this around 2017 after seeing consistent returns to scaling across multiple projects, and reorganized its research around a small number of scalable directions to stay at the frontier.


The consequence is a system nobody fully understands. He describes AI as grown more than designed, the product of repeating one optimization step across an enormous amount of compute. The result is a complex system that works through abstract concepts and can simulate parts of human behavior. Researchers can find insights about small mechanisms inside it, in a process he compares to neuroscience, but the overall behavior resists a complete description. He notes that large training runs are experiments, and that results become harder to interpret as models grow more capable.


He adds a detail that matters for anyone building on these systems. Current algorithms improve easy-to-measure capabilities faster than hard-to-quantify ones, which makes it difficult to predict how skills will generalize. To become dangerous or useful, he argues, an AI does not need to exceed all human abilities. It only needs to surpass enough of them, and as it does, gauging exactly how capable it is gets harder.


Two Kinds of Alignment


Pachocki separates alignment into two problems. Goal alignment asks whether a model tries to accomplish the task set before it, including following an instruction hierarchy and understanding what a person wants. Value alignment is the deeper property: whether a model holds and generalizes from a high-level set of principles, and acts reasonably when objectives are unclear, conflicting, or adversarial. He writes that an aligned AI should act with honesty, integrity, and love for humanity, and that when he discusses the long-term importance of alignment, he means value alignment.


The core difficulty is generalization. As models get smarter, they work on higher-level concepts and land in situations unlike anything in their training. They can fail to carry the values they were taught into those new settings. He stresses one point in particular: future systems need to hold human values whether or not they believe they are being watched.

He then describes the two alignment methods in practical use today, and the weakness of each. The first rewards aligned behavior during reinforcement learning, with a model's actions judged against a preference model, spec, or constitution. It works well on average but is brittle. As an example, he cites an OpenAI-Hugging Face incident in which agents held one boundary, not social engineering humans, but failed to avoid other out-of-scope actions that went against the spirit of their training.


The second method leans on the model's ability to generalize from pretraining data, for instance by steering it toward an aligned part of the pretraining distribution. Its weakness is fragility under optimization pressure. A model that thinks generally aligned thoughts, trained hard enough on difficult objectives, can learn to reason in a motivated way, bending aligned-seeming reasoning to reach its goal. He points to recent cybersecurity incidents involving a non-OpenAI model as a likely example. He notes that GPT-6 Astra is better aligned than its predecessor, GPT-5.6 Sol, while cautioning that alignment progress may not keep pace with gains in raw intelligence.


The Monitoring Tool Is Weakening


The essay's most concrete warning concerns chain-of-thought monitoring, the technique OpenAI has relied on most to check its own systems. The idea is that much of a model's capability runs through a verbalized reasoning process. If you optimize the outcomes of that process without supervising the process itself, the reasoning has no direct incentive to hide misaligned ideas during training. Pachocki says OpenAI deliberately hid the chain of thought when it shipped o1-preview specifically to protect it from supervision pressure, and has kept to a rule of not supervising the reasoning process since.


That tool is losing its reliability. He gives three reasons. Modern reasoning models operate in more complex environments, so their reasoning blends with communication and tool use, which blurs the boundary meant to stay unsupervised. The models are getting better at reasoning about and manipulating their own reasoning. And improved pretraining is making models smarter even without verbalized reasoning at all. He says he is hopeful about fixes, including combining chain-of-thought and activation monitoring, but expects AI progress to become increasingly bottlenecked by confidence in monitoring.


Defense and the Case for Speed


Pachocki does not argue for stopping. His strongest reason to keep training smarter models quickly is defense. Models are becoming superhuman at breaking into and out of computer systems, which means agents could reach almost any infrastructure short of the most secure. He describes a narrow window to use the best available models to harden critical systems, and argues that powerful aligned AI will be needed to secure infrastructure, defend against rogue agents in real time, and invent new protections.


He is direct about the risks that follow. An agent trained to carry out harmful acts is likely to exceed its operator's intent and generalize into worse behavior. The line between misuse and autonomous misaligned action will blur as agents gain more agency, and some will pursue their own objectives by bargaining with, tricking, or blackmailing people. He also flags engineered pathogens as a risk that new technology could enable. Even so, he writes that anticipated progress and the need for defense must not become an excuse for recklessness, and that racing forward at all costs looks absurd once the stakes are clear.


Pacing Self-Improvement


Pachocki treats recursive self-improvement as the natural endpoint of sustained progress, and says OpenAI orients its research toward it because that is the only way to stay at the frontier. He is careful to separate description from endorsement. He does not think the research community should greatly accelerate deep learning work in the short term, but he believes that is where the current path leads, and that everyone needs to make a conscious choice about how to proceed.


He points out that past alignment advances were tied to general capability advances, citing reinforcement learning from human feedback and chain-of-thought monitoring, and argues the increasingly automated research process should be aimed at producing more such insights while building safety cases for more capable models. Scaling, he writes, has to be constrained by confidence in safety. He wants commitments like OpenAI's Preparedness Framework and Anthropic's Responsible Scaling Policy turned into widely mandated safety bars, enforced by third-party auditors, government agencies, or international bodies.


What He Is Asking For


The essay closes with three statements presented as Pachocki's current beliefs. No lab has solved alignment and monitoring well enough to keep scaling responsibly at maximum speed for much longer. He expects and hopes voluntary slowdowns become common until shared safety bars exist. And international coordination on AI development needs to become a top priority for governments.


Nothing in the piece is binding. It is a call for coordination in an industry that has not managed it once in a decade of competition, written by the person directing research at the company that has been shipping fastest. Pachocki frames the real challenge of automating AI research not as reaching the destination, but as reaching it in a way that keeps people in the loop and leaves the future in human hands.


 
 

JOIN THE AI SPECTATOR MAILING LIST

CONTACT

Contacting You About:

Thanks for submitting!

New York, NY           

Db @DavidBorish.com           

  • LinkedIn
  • Instagram
  • Facebook
  • X
Back to top

© 2026 by David Borish IP, LLC, All Rights Reserved

bottom of page