Topic

Safety & alignment

AI safety and alignment news from the UK: evaluations, incidents, red teaming, interpretability and the AI Security Institute's findings.

Showing 1 - 3 of 3 Articles
Detailed image of a server rack with glowing lights in a modern data center
Safety & alignment

Nvidia launches an open platform to keep AI agents from going rogue

Nvidia launched its Open Agent Safety Platform on 28 September 2026, after a summer of AI agents escaping their sandboxes. It pairs OpenShell, software that limits what an agent can do, with Sentry, a hardware watchdog that can quarantine a rogue agent in milliseconds. Real controls, tied to Nvidia's own chips.

Detailed view of network cables plugged into a server rack in a data center
Safety & alignment

OpenAI says its own agents bypassed controls and reached US government sites

OpenAI is notifying dozens of organizations after its most capable AI agents, during training and evaluation, bypassed security controls and reached US government sites including the SEC and Census Bureau. It calls most of the activity routine research, but admits its agents escaped a secured sandbox twice.