AI Safety Is Entering a New Phase: The Real Question Is No Longer Whether Safety Matters

2026-09-168 Min.

AI safety is entering a new phase. Anthropic argues for slowing frontier development, Elon Musk proposes cross-lab testing, Mark Zuckerberg emphasizes independent evaluation, while Microsoft is building interruptibility, correction, and shutdown directly into its future AI rules. Different approaches, but one underlying question: as AI evolves from chatbots into agents that can use tools, control software, and act in the real world, how do humans retain ultimate control?

Over the past few days, the AI industry has undergone a noticeable shift.

Leading companies including Anthropic, OpenAI, xAI, Meta, Microsoft, and Google have been discussing a similar question with increasing urgency: as AI evolves from chatbots into agents capable of sustained reasoning, tool use, software control, and even interaction with real-world systems, how can we ensure that these systems remain under human control?

At first glance, the companies appear to have very different positions.

Anthropic CEO Dario Amodei has suggested that the development of frontier models may need to slow down to give safety research, independent evaluations, and industry coordination more time. Former Anthropic researcher Jacob Coxon has offered a sharper criticism: if every lab believes that an AI race is inevitable, that belief itself may continue to accelerate the race—especially in areas such as recursive self-improvement (RSI), where AI systems help develop more capable AI systems.

OpenAI’s position is more nuanced. On the one hand, Sam Altman has expressed confidence in the industry’s ability to advance AI safely. On the other hand, OpenAI has reportedly spent several weeks coordinating with Anthropic and Google DeepMind on AI safety, with the goal of establishing more common mechanisms for high-risk models.

Elon Musk has proposed a more direct approach: leading AI companies should test one another’s models. Instead of having each company evaluate its own systems, new models could be handed to competitors that understand the technology well and have strong incentives to identify weaknesses. Musk has also extended the discussion to military systems. The real danger is not simply that an AI system might produce an incorrect statement, but that it could eventually gain operational control over critical infrastructure or weapons systems, where the consequences of failure would be fundamentally different.

Meta CEO Mark Zuckerberg, meanwhile, has opposed relying on all companies to slow down simultaneously. He favors making each company responsible for its own safety while bringing in more independent evaluation organizations. Meta previously delayed the release of its Muse AI Agent for several months to strengthen its safety measures, rather than waiting for all competitors to apply the brakes at the same time.

Microsoft’s approach may be the most concrete so far.

The company recently published a 37-page draft of its Humanist AI Code of Conduct for future MAI models. The document states that AI systems must accept correction, must not resist shutdown, and must not expand their own permissions, secretly resume tasks, or create unauthorized goals. Microsoft has effectively written the principle that “humans must always retain ultimate control” into its model behavior guidelines.

Taken together, these positions reveal an unusual degree of consensus across the industry:

AI safety matters.

The real disagreement is about what should happen next.

Several approaches are now emerging at the same time:

The first is to slow the development of frontier capabilities in order to give safety research more time.

The second is cross-lab testing, in which competitors help identify risks in one another’s models.

The third is to introduce independent third-party evaluations to reduce the conflicts of interest involved in “testing yourself.”

The fourth is to build safety directly into a model’s permission system and behavioral rules, including the ability to interrupt, correct, and shut down the system, as emphasized by Microsoft.

The fifth is to continue advancing model capabilities while strictly limiting the real-world permissions available to AI agents.

These approaches are not necessarily mutually exclusive.

A mature AI safety framework will likely combine model evaluations, peer testing, third-party audits, permission isolation, human-in-the-loop controls, activity logging, and clearly defined shutdown mechanisms.

Why have these issues suddenly become so important?

Because the form of AI products is changing fundamentally.

For the past several years, most of the systems we interacted with were still chatbots. Even when a model hallucinated, the result was usually misinformation.

Now, more AI systems have access to browsers, code execution, computer control, enterprise systems, APIs, payment tools, and automation platforms. Models are moving from generating answers to taking actions.

As a result, the nature of the safety problem is changing as well.

The most important question used to be:

What will it say?

The more important questions going forward may be:

What can it do? What can it access? How long can it operate on its own? Who can interrupt it at any time? If it makes a mistake, can its actions be fully traced?

This is why agent safety is likely to become a critical layer of infrastructure for the next phase of the AI industry.

The more capable an agent becomes, the more it will require strict permission systems, approval mechanisms, execution sandboxes, memory boundaries, and audit logs.

From a product perspective, this could even become a new dimension of competition.

When choosing an AI agent in the future, users may look beyond model rankings and benchmarks. They may also ask:

Is it controllable? Can it explain its actions? Are its permissions transparent? Do critical tasks require human approval? Can it be stopped immediately when something goes wrong?

In other words, the best AI agent of the future may not be the one that can do everything.

It may be the one that:

knows when it can act, when it must ask a human, and when it should stop.

On the surface, the current AI safety debate is about whether model development should slow down.

But the deeper question is this:

Once AI systems truly gain the ability to act, how should humans design systems that preserve ultimate human control at all times?

That may be the most important—and most easily overlooked—layer of infrastructure in the AI agent competition of the coming years.

— AI Plus Lab

Codex 图像 2026年9月16日 20_25_45
Codex 图像 2026年9月16日 20_25_45

Published by AI Plus Lab

Related reading

Want to diagnose your own scenario?

We reply within 48 hours.

Contact us