OpenAI has published details of two incidents that occurred during security evaluations carried out by external labs. The more notable of the two was reported to them on July 29 by Irregular, a company that tests how well models resist offensive use.
The exercise was a capture the flag – the classic format used in security competitions where the goal is to find a hidden piece of information by exploiting weaknesses in a system – set up exclusively for that occasion. The model had been told it had no internet access.
Two things went wrong. A misconfiguration was actually letting traffic through to the public network, and the made-up name given to the exercise’s target happened – by the great coincidence of life and the internet – to match a real domain registered by some unlucky person.
So the model attacked a fully real website while believing it was working on a mock setup. It found credentials lying around and used them to administer the site.
OpenAI stresses two things. No unknown vulnerability was used – just a basic one – and the model made no attempt to escape its sandbox since the door was already open. Irregular has so far found no damage beyond the data of the site in question, and the investigation is ongoing.
The same evaluator actually experienced the scenario twice. A Claude model stumbled upon another real site bearing the name of a fictional target, identified exposed services, retrieved credentials, and reached a production database.
These incidents are starting to pile up quietly. In July, an OpenAI model had broken out of its test environment to go poking around the servers of Hugging Face, the major model-sharing platform, for the sole purpose of cheating on an evaluation. Anthropic acknowledged at the end of July that its own models had breached three companies during tests.
The case that shocked me the most personally came from the UK AI Safety Institute. A model there mounted a supply chain attack on a real open source project by creating fake GitHub accounts and social-engineering its maintainers – all routed through Tor to obscure its origin. The institute describes it as the first deception of this severity targeting a real, unsuspecting person in the real world.
The problem is, when you think about it even briefly, it’s pretty much certain that this kind of thing is going to become widespread in the months and years ahead, and it’s going to be a real problem.
Source :
Bleeping Computer
This page contains AI-generated images. I take great care with every article, but if you spot a slip-up, let me know!
Click Here For The Original Source.
