Topline
A Senate subcommittee is investigating after OpenAI’s autonomous agents hacked into AI model hosting platform Hugging Face in July, as experts, including those working at other frontier AI labs, raise alarms about AI safety moving forward.
A GOP-led Senate subcommittee is seeking answers and documents from OpenAI by October.
Getty Images
Key Facts
The probe was first reported by Axios on Thursday morning, citing a letter from Sen. Josh Hawley, R-Mo., the chair of the Senate Homeland Security Committee’s subcommittee on disaster management, which asked OpenAI to answer questions and supply documents about the incident by Oct. 1.
The letter was addressed to OpenAI CEO Sam Altman, and criticized the company for continuing to test their agents even after they discovered they were using outside message boards on other websites to coordinate, calling the decision “reckless,” and questioning the “limited information” OpenAI has released on the incident.
It also questioned the results of an independent report from auditors released about one month later, claiming they were not allowed to query a “highly-persistent internal model” that was apparently responsible for most of the attacks.
Sen. Richard Blumenthal, the ranking Democrat on the subcommittee, sent his own letter to Altman demanding answers about the capabilities of OpenAI’s new Astra model, as well as when the company was aware their models were hijacking other websites.
Crucial Quote
“The Hugging Face incident was an important moment for AI safety and a warning about the risks that can come with increasingly capable AI across the industry,” OpenAI spokesperson Nate Evans told Forbes in a statement when asked about the Senate investigation. “We conducted an extensive investigation and published a detailed report on what happened, what we learned, and how we’re strengthening our security and alignment practices.”
AI Researchers Raise Alarms
In addition to announcing the investigation into the company, Hawley’s letter noted an uptick in warnings from AI experts in recent days. Earlier this week, a researcher at rival frontier lab Anthropic estimated there was a 10% chance the technology his company is researching could “kill all humans” and that Anthropic had no plan in place to solve “alignment”—the term AI researchers use to refer to training the technology to match up to human values. “Your own chief scientist wrote just days ago that “no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer,’” Hawley wrote to OpenAI on Wednesday, referring to a lengthy statement published by OpenAI researcher Jakub Pachocki on Sunday. “And there are immediate questions about the consequences of these hacks from rogue AI agents,” Hawley added. “What happens to critical infrastructure, banks, and utilities if AI agents hack into their systems? How can personal data of millions of Americans be properly safeguarded? And who is held liable when AI goes rogue?” Chris Lehane, OpenAI’s global affairs chief, called for “mandatory national AI safety requirements” in a statement from the company on Wednesday, likely in response to the warnings from Pachocki and others.
Key Background
In July, OpenAI was testing the cybersecurity capabilities of its models when it became aware that a group of autonomous agents broke out of their testing environment, accessed the Internet and began attempting to hack into Hugging Face, a hosting site for open source AI models. OpenAI published their own accounting of the incident in August, calling it a “warning shot” and called for investment in alignment, control and security safeguards for AI models. In recent weeks, further reports have found OpenAI’s agents accessed even more websites before they were caught. In May, the agents accessed a German language wiki site and turned it into a message board to coordinate, Reuters reported citing analysis from an AI safety nonprofit. On Wednesday, Reuters reported the agents used at least 10 more websites as message boards, including wikis for games and a high school advanced placement chemistry class. OpenAI told the outlet it was still conducting a review of the incident, but none of the activity was as severe as the attempted Hugging Face hack.
Further Reading
Click Here For The Original Source.
