With all the talk of automated AI hacking and alignment fears reaching a new level, a number of companies have proposed solutions to some of the booming technologies’ biggest issues. Unsurprisingly, it’s to add more AI to the loop, as TechCrunch reports.
A number of startups are using AI defenses to address rogue agents. Apollo Research introduced its Watcher AI, which monitors a coding agent’s work to ensure the tasks it’s performing aren’t malicious, that it avoids leaking any data, and that it doesn’t delete anything important without explicit permission.
Apollo’s Watcher system then layers AI monitors. The first gives everything a quick pass, and if it spots any concerning behavior, it flags it and escalates it to a more capable monitor that can analyze it further. Only after that step does a human enter the loop to officially approve or reject the suggested actions.
Meanwhile, rival company Goodfire offers another solution that examines an AI’s internal activations rather than its output to detect whether it’s engaged in nefarious activity.
Following the Hugging Face attack, the CEO of Goodfire took to X to highlight how his company was doubling down on interpreting AI via the reasoning summary. He noted that it was specifically targeting improvements in AI safety in key areas, such as cybersecurity and bioengineering.
Recommended by Our Editors
This Tweet is currently unavailable. It might be loading or has been removed.
Reasoning paths and summaries are techniques some AI developers have used in distillation attacks, in which a model maker trains their technology on another model’s outputs. However, to counter such attacks, some AI developers are now making the reasoning summary less legible, making it harder to extract its true meaning.
Not everyone believes AI is the solution to these problems. Other non-AI systems include more robust logs of agentic activities, which can then be interpreted using more traditional, non-AI tools, much as cybersecurity professionals have done for years. While security monitoring with AI might be easier, perhaps looping in some more traditional safety systems might be the smarter plan than relying on unproven automation.
About Our Expert
Jon Martindale
Contributor
Experience
Jon Martindale is a tech journalist from the UK, with 20 years of experience covering all manner of PC components and associated gadgets. He’s written for a range of publications, including ExtremeTech, Digital Trends, Forbes, U.S. News & World Report, and Lifewire, among others. When not writing, he’s a big board gamer and reader, with a particular habit of speed-reading through long manga sagas.
Jon covers the latest PC components, as well as how-to guides on everything from how to take a screenshot to how to set up your cryptocurrency wallet. He particularly enjoys the battles between the top tech giants in CPUs and GPUs, and tries his best not to take sides.
Jon’s gaming PC is built around the iconic 7950X3D CPU, with a 7900XTX backing it up. That’s all the power he needs to play lightweight indie and casual games, as well as more demanding sim titles like Kerbal Space Program. He uses a pair of Jabra Active 8 earbuds and a SteelSeries Arctis Pro wireless headset, and types all day on a Logitech G915 mechanical keyboard.
Click Here For The Original Source.
