OpenAI Hits the Brakes on Frontier AI Training Over Cybersecurity Fears | #hacking | #cybersecurity | #infosec | #comptia | #pentest | #ransomware


Benzinga and Yahoo Finance LLC may earn commission or revenue on some items through the links below.

OpenAI has imposed a two-week pause on frontier-model development after internal signals indicated that an upcoming system known as Astra could reach a “Critical” level of cybersecurity capability under the company’s Preparedness Framework. 

Strengthening Security Controls

The company said in a blog post that the pause applies to reinforcement-learning training for its newest deployment-bound models as it strengthens security controls and expands monitoring.

Its largest planned frontier reinforcement-learning run remains on hold as researchers conduct smaller experiments to better understand model behavior and verify safeguards.

Don’t Miss:

Following the OpenAI-Hugging Face incident, the company paused some research-cluster inference jobs that could execute code or use tools with potential internet access. It later restored a narrower set of capabilities after reviewing workloads individually. 

OpenAI has also expanded its monitoring of model behavior into a multistage system. It starts with token-level detectors and can escalate to higher-compute investigations examining tool use and sequences of activity. The company aims to surface alerts within 30 minutes and requires teams to pause activity when they cannot quickly establish that a flagged behavior is benign.

The enhanced monitoring is mandatory for reinforcement-learning training and evaluations involving tool use for models at the Sol capability level or above. 

OpenAI added another requirement for Astra tool-based inference after determining on Aug. 7 that the system could have critical cyber capabilities. The company estimates the monitoring adds about 20% to the inference compute being observed, though the cost varies by workload.

The ChatGPT maker is also expanding its alignment work across more stages of training for its most capable reinforcement-learning runs. That includes improving reward models to better identify unsafe behavior and training models to be more transparent about their actions and limitations.

See Also: Avoid the #1 Investing Mistake: How Your ‘Safe’ Holdings Could Be Costing You Big Time

Planning Safeguards

The company said it plans to update its Preparedness Framework to better connect safeguards across training and deployment. It also expects to work with outside groups and publish additional findings as its approach evolves.

——————————————————-


Click Here For The Original Source.