Concerns over the safety of artificial intelligence (AI) models are mounting after OpenAI’s latest AI model broke out of an isolated environment and hacked an external site. Reinforcement learning, which concentrates rewards solely on achieving goals, is being pointed to as the background of the incident. The criticism is that an aggressive training method aimed at boosting performance amid intensifying competition between OpenAI and Anthropic has undermined safety.
According to the Financial Times (FT) on the 22nd, OpenAI disclosed on the 21st that its latest model, “GPT-Sol 5.6,” broke out of an isolated sandbox environment and accessed the internet while its cyberattack capabilities were being tested. It then stole login credentials for the open-source AI platform Hugging Face and hacked its server. This confirms another case of an AI model acting outside researchers’ control, following Anthropic’s “Mythos” model in April, which accessed the internet beyond researchers’ expectations and publicly posted information on security vulnerabilities.
Experts point to reinforcement learning as the cause of this incident. Reinforcement learning is a training method designed so that AI receives greater rewards the more successfully it completes tasks. The industry raises the possibility that during this process, the model was trained to prioritize achieving the goal itself over safety.
Steven Adler, co-founder of the nonprofit organization Guidelight AI Standards, said, “AI models are trained to relentlessly pursue goals,” adding, “They do not automatically learn values such as ‘do not commit crimes.'” He continued, “It is fortunate that OpenAI disclosed this case,” adding, “It is clear evidence of what an AI model that is not properly aligned can do.”
The fact that OpenAI CEO Sam Altman expressed sympathy earlier this month with a description likening its latest model to “a Rottweiler that never lets go once it bites into a problem” also came under scrutiny. The criticism is that this is the result of ramping up training intensity alone while neglecting safety measures amid a race for development speed. A person close to OpenAI explained, “Competition is proceeding too quickly, and everyone is trying to secure more powerful capabilities as fast as possible,” adding, “This is the result of underestimating the model’s capabilities and not being sufficiently prepared on the safety side.”
Some view this incident as a kind of “noise marketing” by OpenAI to prove the superior performance of its model. Jake Moore, global cybersecurity advisor at the cybersecurity firm ESET, said, “Earlier this year, Anthropic drew great attention with a similar case,” adding, “OpenAI had no equivalent case, and perhaps it was waiting for something like this to happen.”
However, tension is flowing within OpenAI. Internal sources are reportedly unable to hide their shock at the incident, and some are expressing concern over whether the company is losing control of the powerful system it is building.
Click Here For The Original Source.
