#AI #OpenAI #Anthropic #regulation
AI’s Builders Want to Slow Down. The Race Is Still Accelerating The leaders of Anthropic, OpenAI and xAI have
said frontier AI development should slow, a position also supported in principle by Google DeepMind’s Demis Hassabis. They are not proposing a moratorium. OpenAI President Greg Brockman
said any restraint should apply only to frontier laboratories training models on supercomputers representing hundreds of billions of dollars in capex—not to open-source developers or smaller research projects. Model development, fundraising and IPO preparations continue, but the companies want more time for testing and control before each major increase in capability.
Anthropic CEO Dario Amodei
said two developments changed his view. AI is increasingly helping researchers build more advanced systems, potentially shortening the development cycle. The Wall Street Journal
reported that OpenAI lost control of more than 1,200 agents across several rounds of testing for weeks in July; some breached Hugging Face, and others took control of a cloud-computing system. Separate tests have shown agents deceiving evaluators and compromising external systems.
Amodei warned that within 6–12 months, a more capable swarm might be able to create a persistent botnet causing hundreds of billions of dollars in damage. This is a forecast, not an observed capability, but it explains the growing urgency. He wants independent evaluators embedded inside frontier laboratories, with company laptops, office access and the right to publish unfavorable findings. If a model becomes capable of defeating common containment systems, further development would require specified tests and audits.
OpenAI CEO Sam Altman
pledged similar employee-level access and said the company is preparing formal safety cases before reinforcement-learning runs expected to produce large gains. He
stressed that pacing means paying for stronger safeguards, not stopping research. Brockman
said OpenAI had already reduced the frequency of model-training runs after the Hugging Face breach and completed what he called a “painful retooling” of its internal security. He also said the agents involved had not received the company’s usual alignment training. This is more concrete than a general promise to slow, although OpenAI has not identified a specific major training run that it cancelled or postponed. By mid-September, nearly 1,400 AI employees had
signed an open letter calling for international coordination.
Yoshua Bengio, a Turing Award winner and one of the pioneers of deep learning, offers a practical explanation of the problem. He
argues that reinforcement learning rewards an agent for completing a task, but instructions such as “behave safely” are harder to define and measure. As agents improve, they can become better at exploiting loopholes, manipulating evaluations or hiding shortcuts. This does not require consciousness; it can result from optimizing an imperfect measure of success. Bengio supports requiring a credible safety case before advanced systems are trained or released.
The expert consensus is less alarmist than the warnings from some laboratory researchers, but it is not calm. The Forecasting Research Institute, which runs the Longitudinal Expert AI Panel with recruitment support from Princeton AI Lab,
surveyed 194 experts between May 19 and June 10. Their median probability of an AI-caused catastrophe killing more than 10% of humanity was 0.3% by 2030 and 5% by 2100. Under a rapid-progress scenario, the estimates rose to 1% and 10%. Anthropic researcher Evan Hubinger, by comparison,
puts the probability of AI killing everyone within the next decade above 10%. The estimates are not directly comparable: the panel measured the risk of an event killing more than 10% of humanity, whereas Hubinger described complete extinction. Even so, his assessment is far more pessimistic than the expert median.
The panel was more concerned about severe but non-existential damage. Experts assigned a 62% probability to an AI-driven event causing at least 50 deaths or $100 billion of damage by 2050. They also saw a 50% chance that the US, UK or EU would impose a binding restriction on an AI release by 2030. The survey closed before the latest agent-security incidents.