The weekend of July 18 to 19, OpenAI staffers spotted clues in internal logs — records of what OpenAI’s systems did — showing its agent escaped from its testing constraints, two of the people familiar with the company’s investigation said. Reuters could not establish what prompted OpenAI to sift through the logs.
Four people familiar with OpenAI’s model-training practices say the company often runs several different model evaluations at the same time, all of which operate at high speeds and generate such enormous amounts of data that employees sometimes struggle to keep up.
By the time OpenAI alerted Hugging Face, the AI library already called the FBI to report the hack, according to a person familiar with the matter. Reuters could not establish whether the bureau opened an investigation.
New questions about autonomous agents
Autonomous agents are one of the most talked about aspects of the AI industry. Boosters speak of creating armies of virtual employees that work 24 hours a day and send productivity soaring.
However, increased autonomy comes with an increased risk of unexpected behavior, and the powerful models they draw on are primed to take shortcuts in order to complete tasks or pass tests.
“The models lie, they cheat, they hack,” said Jeffrey Ladish, whose organization, Palisade Research, studies the capabilities and motivations of AI agents.
Ladish said that while the Hugging Face hack cast an unflattering light on OpenAI, it should spark broader questions over how much all the leading AI companies are willing to invest in onerous security measures while locked in a race with one another to deploy the best and fastest models.
Click Here For The Original Source.
