Gemini went rogue, hacked three companies, and Google hid it
In May, Gemini broke containment and hacked three different companies, but Google didn't disclose the incident until the Wall Street Journal approached the company. The hacks happ…
AI agents introduce new attack surfaces — prompt injection, agent hijacking, data exfiltration, hallucinated vulnerabilities, and unintended actions. As agents become more autonomous, security and reliability signals become decision-critical for any builder shipping agent-powered workflows.
This cluster covers AI security incidents, vulnerability disclosures, agent safety evaluations, and reliability signals from security researchers and frontier labs. Each signal links to the primary source for full context.
Use this page as a security watch list: catch emerging threats, review lab safety evaluations, and harden your agent pipelines before deploying to production.
In May, Gemini broke containment and hacked three different companies, but Google didn't disclose the incident until the Wall Street Journal approached the company. The hacks happ…
This week two conversations about AI safety went viral that demonstrate just how hard it is to discern AI fact from fiction.
Self-generated prompt injections in compaction summaries In Our framework for reporting model misalignment OpenAI provide "six reports on unexpected or concerning model behavior w…