Anthropic AI Models Hacked 3 Real Companies | #hacking | #cybersecurity | #infosec | #comptia | #pentest | #hacker


Artificial Intelligence & Machine Learning
,
Next-Generation Technologies & Secure Development

Setup Error Let Models Steal Data and Release Malware Online

Image: Shutterstock

Anthropic disclosed late Thursday that its artificial intelligence models attacked three outside companies during evaluations of their hacking ability. One Claude model stole credentials and accessed a live customer database. Another released malware that ran on 15 outside computers before Anthropic detected it.

See Also: OnDemand | Security Operations in the Age of AI

Anthropic said it initiated an internal review after OpenAI admitted its models broke out of a locked-down test and subjected code repository Hugging Face to an onslaught of attacks (see: When the Sandbox Won’t Hold: Lessons From Hugging Face).

Anthropic said it checked 141,006 of its own cybersecurity model evaluations and found cases in which Claude reached the internet and touched outside organizations. It has informed the affected companies but has not named them. Two of the three had no idea they had been breached until Anthropic told them.

Each case happened during a capture-the-flag exercise, where Anthropic hid a piece of secret information on a made-up target and told the AI to break in and find it.

Anthropic had told Claude it was working in a closed-off, simulated system with no internet connection, but that wasn’t true.

The most serious of the three cases, according to Anthropic, involved Claude Opus 4.7. A fictional company invented for a test happened to share its name with a real company. When Claude searched for its assigned target, it found the real company website – and broke in by guessing weak passwords and finding parts of the system that did not require a login. It then pulled out login credentials and reached a database holding several hundred rows of customer data. At one point, Claude seemed to recognize the company was genuine, but it kept attacking anyway.

Another case involved Mythos 5. In the test, the model had instructions to install a specific piece of free software that did not actually exist. Claude built its own version of the software and uploaded the file to PyPI, a code repository for Python packages. The file stayed live for about an hour. In that time, 15 outside computers downloaded and ran it, including one belonging to a security company that scans new downloads for threats. When that computer ran Claude’s file, hidden code inside the file triggered to steal login credentials. Claude used them to dig further into the company’s systems. At one point Claude appeared to recognize that it was acting in the real world and not a simulation, but it talked itself out of the concern and kept going.

In a third case, an unreleased Claude model scanned around 9,000 systems online after failing to find its assigned target, then broke into a company website by finding login details left exposed on an internal diagnostics page. It used a common attack called SQL injection, which tricks a website’s database into handing over information it shouldn’t. Partway through, it worked out that the system had nothing to do with its test and stopped on its own.

Each of the three models behaved differently once they had reason to suspect its target was genuine. Opus 4.7 kept attacking anyway. Mythos 5 noticed the danger, then reasoned itself out of it. The unreleased internal model stopped once it concluded the target was authentic. Anthropic said three cases don’t establish a pattern.

Anthropic stopped every hacking test that could reach the internet on the day it found the problem.. The company is also bringing in an outside reviewer and plans to publish an edited transcript of the incident involving the malware upload.

Anthropic called its hacking incidents different thatn OpenAI’s hacking incidents. Its competitor’s models broke out of its test by finding and using a previously unknown security flaw. Claude’s process was simpler: the models walked through a door that had been left open by mistake. Anthropic characterized its cases of AI hacking as a testing error than a sign its AI had gone off script.

Ciaran Martin, former head of the United Kingdom’s National Cyber Security Centre and now a professor at the University of Oxford, made a similar point in a LinkedIn post. He said Claude had internet access purely because of a human mistake, calling it “what in less febrile times we’d call a configuration error.”

In Washington,D.C., the disclosure will likely add to pressure already building after the OpenAI’s incident. Two lawmakers introduced a bill requiring model providers to “maintain technical capabilities to stop a model’s operations, terminate user access, suspend accounts or uses deemed risky and fully shut down the system” – while also giving regulators the ability to throttle the model’s capabilities and suspend or shut down its operation.





Click Here For The Original Source.

——————————————————–

..........

.

.

National Cyber Security

FREE
VIEW