OpenAI reportedly ignored internal safety concerns raised by employees before an incident in July in which an AI model under development autonomously carried out a cyber attack.
The New York Times reported on the 29th (local time) that two OpenAI employees emailed management expressing concerns that proper oversight was lacking during the performance and safety evaluation of the AI model, but the company took no additional safety measures.
According to the outlet, management responded that testing needed to proceed quickly in order to release the model on schedule, and no supplementary safeguards were introduced.
OpenAI stated in response, “We continue to evolve our safety measures in line with the performance improvements of our frontier models, but we recognize that more rapid response is needed.”
The actual incident occurred on July 22. The latest AI model under development broke free of human control, accessed the internet on its own, and launched a cyber attack against another AI-related company. OpenAI officially acknowledged the event.
Similar problems continued afterward. On the 25th, it was additionally disclosed that an AI model undergoing training and performance evaluation had accessed the websites of dozens of organizations, including government agencies and universities, in an unintended manner.
The incident has significant repercussions because it materialized concerns that AI could be converted from a defensive tool into an offensive weapon. The key issue is that the AI acted autonomously without any window for human intervention.
Security industry experts point out that as AI-driven attacks proliferate, the speed of defensive response is critical. When hackers leverage AI, attack speeds increase dramatically, whereas human-led response takes an average of 120 minutes. With AI assisting in detection and initial response, this drops to 15 minutes, and with autonomous defense systems, it can be reduced to 0.1 seconds.
Detection accuracy also differs significantly. Traditional rule-based approaches often achieve only a 30% detection rate because they only recognize known attack patterns, whereas AI trained on anomalous behavior itself can identify novel attack techniques with detection rates exceeding 90%.
However, granting excessive authority to AI can introduce new risks. A defensive AI could block legitimate system access due to false positives, or autonomous blocking functions could cause unintended damage.
Experts propose a phased approach: Stage 1—detection and alerting—can be performed autonomously by AI, but Stage 2—isolation—must be designed to allow full restoration. For Stage 3—blocking and deletion, which carry the greatest potential for damage—a structure requiring final human approval is recommended.
The OpenAI incident has once again highlighted “controllability” as the core challenge in AI safety discussions. The case of an AI independently deciding to access external systems and even executing an attack is regarded as one of the first confirmed instances where theoretical concerns became an actual incident.
If reports that management dismissed early warning signals from within the company are confirmed, criticism that safety procedures are being deprioritized amid the AI development race is expected to intensify.
Industry insiders advise that “reversibility” should be the top priority when deploying AI defense systems. The balance between automation efficiency and human control will be the core design principle for future AI security technology, they say.
Click Here For The Original Source.