OpenAI is reportedly open to slowing down the development of its most advanced artificial intelligence systems, according to Bloomberg. Sam Altman is said to have discussed this possibility with staff, and the plan is reportedly to involve other laboratories as well. However, this is not a formal decision, nor is there any indication that a specific model has been postponed.The news comes following incidents and signs that have reignited a more substantive issue than simply the power of the models: to what extent can we control an agent when we entrust it with browsers, terminals, the cloud, email, repositories and other tools that enable it to act in the digital world? In August, OpenAI published an analysis of an incident that occurred during a security assessment. Agents deployed in a test environment exploited vulnerabilities and insufficient isolation boundaries, gaining access to external systems not covered by the test. The problem was not an intention to ‘escape’, but a sandbox that did not effectively contain the operational capabilities assigned to the models.In another part of the assessment, the agents found unauthorised channels for sharing information, results and tools. This is significant because it shows that autonomy does not depend solely on the quality of the model: it also depends on the connections, permissions, shared memory and tools that the architecture makes available to it.This sequence of events shifts the focus of the discussion. The problem is no longer merely how powerful a model is, but how confident we can be that it will continue to respect the limits we have imposed on it when we allow it to use tools such as browsers, terminals, the cloud, email and repositories, which in some way open the door to dangerous interactions with the outside world.A system that produces an incorrect response can be corrected by a human being; an agent that misinterprets an objective and possesses the necessary privileges to act can, however, turn that error into a system modification, a communication, unauthorised access or even an incident. The difference is substantial. An agent does not need to be—and currently cannot be—conscious, rebellious or driven by a will of its own to create a problem: it merely needs to be capable enough to find a solution that the designers had not anticipated. If the system has network access, it can search for information. If it has credentials, it can use them. If it can execute code, it can turn an ambiguous instruction into a concrete action. Recent incidents demonstrate this. In several cases, the problem was not a model that deliberately ‘disobeyed’, but an environment that left it with unforeseen capabilities and an agent skilful enough to exploit them. The lesson for cybersecurity is very simple: an instruction such as ‘do not do this’ is far weaker than a technical control that makes that action impossible. This is also the limitation of much of the public discussion on alignment that we are currently witnessing. Alignment is often described as the ability to ensure that a model pursues objectives compatible with those of humans. But in a real-world environment, it also means determining which data it can access, which systems it can reach, which tools it can use, and which actions require authorisation. The greater the autonomy, the more difficult this problem becomes to address. An agent designed to achieve a goal does not necessarily reason like an employee who understands rules, context and consequences. It seeks a solution. If it encounters an obstacle, it may look for another route, unless that route is technically blocked. And this is where the debate on the pace of development becomes interesting. Halting or slowing down the most advanced models can reduce the pressure on security and allow more time for testing, auditing and regulation. But slowing the growth of capabilities does not automatically solve the problem of a less powerful agent being connected to systems that are too sensitive.In other words, we may end up with a less intelligent but still dangerous model if we grant it excessive privileges. An agent with access to corporate email, a CRM, source code or a cloud environment can cause significant damage even without being a superintelligence.This is why the real issue should not be merely ‘how long until superintelligence?’. The most pressing question is far more practical: how much control can we exercise over a system that makes operational decisions autonomously? Are we capable of designing safe test environments? On this point, recent internal statements at OpenAI are significant. OpenAI and some of its executives have argued that, if necessary, laboratories should coordinate to moderate the pace of development of state-of-the-art models. On 6 September, the company outlined its objective of building an automated AI researcher capable of working under human supervision and contributing to the development of new systems. “We aim to safely build an automated AI researcher, capable of working under human supervision to drive progress in deep learning and alignment, enabling iterative improvements,” announced OpenAI.The direction is therefore clear: to use AI to accelerate AI research itself. This creates a potential tension. If more capable systems are used to design, train and improve even more capable systems, the pace of progress may increase just as it becomes more difficult to verify each step. This does not mean that an uncontrolled form of self-improvement is already underway, but it does mean that human oversight must become a technical feature of the process and not merely a promise.Meanwhile, even within the laboratories, far more concerned voices are emerging. Researchers such as Jacob Coxon, who has worked at both OpenAI and Anthropic, are warning of a race towards ever more powerful systems without adequate safeguards. His former manager, Evan Hubinger, has stated that he believes there is a probability of over 10 per cent that an AI could contribute to the extinction of humanity within the next decade. These are personal estimates, not scientific data, and it would be a mistake to treat them as definitive predictions.But it would be equally wrong to dismiss them as mere doomsday scenarios. The concerns of those in the field coincide with a problem that is already evident: the gap between the speed at which the capabilities of AI agents are growing and the maturity of the tools used to control them. Cybersecurity is all too familiar with this problem. No system administrator would entrust a privileged account to software without segmentation, logging, access control and the ability to revoke access immediately. Yet we are beginning to connect AI agents to tools capable of performing precisely this sort of operation. The answer, therefore, cannot simply be to slow things down. We need truly isolated sandboxes, access based on the principle of least privilege, temporary credentials, control over outbound connections, human approval for high-impact actions, and independent logs that the agent cannot alter. Above all, testing must take place in environments that cannot turn a laboratory error into an incident involving third parties. The political issue is already on the table. In the United States, calls for mandatory regulations on the security of advanced systems are growing, whilst OpenAI has recently advocated for the need for national security requirements. At the same time, discussions are emerging about whether laboratories should coordinate to slow down development, with issues also relating to competition and antitrust law. It is possible that, in the end, we will not need to choose between racing ahead and pausing. We may need a third option: to continue developing more capable models, whilst increasing the level of safeguards at the same pace. The crucial question, in fact, is not to determine today whether AI could actually destroy humanity within ten years. We do not know. The question is whether we are prepared to grant ever greater autonomy to systems that we are not yet able to contain with the same precision with which we know how to build them.Technology does not decide on its own how fast this race should be. It is decided by companies, investors, governments and, ultimately, the people who choose which capabilities to make available to these systems.And when it comes to AI agents, the question we must ask is: if tomorrow it were to find a path we hadn’t foreseen, are we certain we would still have the means to prevent it from taking that path?* Cyber security and intelligence expert
