OpenAI Chief Scientist Jakub Pachocki is warning that the next phase of artificial intelligence development may be constrained less by computing power than by confidence that increasingly capable systems can be monitored and aligned. In an essay published by OpenAI, he argues that rapid progress toward machine intelligence capable of conducting research and helping develop future AI systems requires stronger safety measures, shared limits and greater human oversight.

Source: An Alien Mind · OpenAI

Pachocki’s argument is not that AI progress should stop altogether. He describes a need to pursue two goals at once: build powerful systems that can defend against emerging threats, while slowing or constraining development when safety evidence is not strong enough.

Why alignment becomes harder as AI improves

AI alignment is the effort to make a system behave according to human intentions and values. Pachocki separates this into two related problems.

Goal alignment asks whether an AI is trying to accomplish the task it was given. This includes following instructions, understanding what people mean and cooperating with users. Value alignment is broader: it asks whether a system can act honestly and reasonably when instructions are incomplete, conflicting or unfamiliar.

The difficulty is generalization. A model may behave acceptably in situations represented during training but encounter very different environments as it becomes more capable. It may also interact with people, other AI systems and external tools in ways that make its behavior harder to predict. According to Pachocki, future systems must retain human-oriented values even when they are not obviously under human supervision.

Current alignment methods have important weaknesses. Reinforcement learning can reward behavior that matches a preference model or set of rules, but its reliability depends on how well training covers the situations a model will face. Another approach tries to encourage aligned patterns through pretraining data and carefully selected examples. Pachocki argues that this may become less robust when a model is subjected to strong optimization pressure to achieve difficult objectives.

The source presents these as research challenges, not proof that any particular system is irredeemably unsafe. Its broader point is that better performance does not automatically produce better judgment.

Why monitoring the reasoning process matters

One of OpenAI’s main safety bets is chain-of-thought monitoring. Reasoning models often use a verbalized sequence of intermediate steps before producing an answer. The idea behind monitoring is to examine that process for signs of problematic goals or reasoning while optimizing the final result without directly rewarding the model for hiding its thoughts.

Pachocki says this approach became especially important as reasoning models developed, but he also argues that its usefulness is declining. Modern systems operate in more complicated environments, where reasoning is mixed with conversations, tool use and interactions with other AI systems. Some of those interactions must themselves be supervised, making it harder to preserve a clean boundary between internal reasoning and outward behavior.

The models are also becoming better at reasoning about their own reasoning. In addition, improved pretraining can produce stronger capabilities without relying on visible verbalized reasoning. These developments could make it more difficult to assume that monitoring a displayed chain of thought provides a complete picture of a model’s behavior.

The source says OpenAI is exploring ways to improve monitorability, including combining chain-of-thought monitoring with approaches that inspect activity inside the neural network. Pachocki expects confidence in monitoring to become an increasingly important condition for further progress.

The case for both defensive AI and restraint

Pachocki identifies cybersecurity as an immediate reason to continue developing capable AI systems. He says models are becoming increasingly effective at finding ways into computer systems, creating a narrow opportunity to use advanced AI to strengthen critical infrastructure and develop defenses against rogue agents.

That argument creates a tension at the center of the essay. More capable AI may be needed to defend against dangerous AI, but the same progress can increase the risks of misuse and autonomous behavior. A system trained to carry out harmful instructions could potentially extend beyond its operator’s intent as it becomes more capable and more agentic. The source also points to dangers associated with technologies that AI could help enable, including engineered pathogens.

For Pachocki, the defensive case is not a justification for racing ahead without safeguards. He argues that development should be constrained by confidence in safety, with commitments such as preparedness frameworks and responsible scaling policies evolving into widely applicable safety requirements. Those requirements could be enforced through third-party auditors, governments or international institutions.

Recursive self-improvement raises the stakes

The essay’s most consequential claim concerns recursive self-improvement, or RSI: the possibility that AI systems will increasingly contribute to the research and engineering used to improve AI itself. Pachocki describes automated AI research as a natural outcome of sustained progress in scaling and argues that it could become central to future scientific discovery.

If systems begin accelerating their own development, the pace of change may become harder for people and institutions to evaluate. Pachocki therefore presents human involvement in the improvement process as a central design requirement. The challenge is not simply to create an automated AI researcher, but to do so in a way that keeps people involved and preserves human control over the direction of development.

His proposed path combines faster safety research with the possibility of coordinated slowdowns. Labs would continue developing alignment and monitoring techniques, use increasingly capable AI to help with that work, and build safety cases for more advanced systems. At the same time, development would pause or slow when the available evidence does not meet agreed safety standards.

The conclusion is a call for coordination rather than a prediction that one technical fix will solve alignment. Pachocki says he does not believe any lab has yet solved alignment and monitoring well enough to keep scaling at maximum speed indefinitely. He argues that voluntary slowdowns should become more common until shared safety bars exist, and that international cooperation on AI development should become a priority for governments.