Base Labs launches an open-weight AI safety partnership with Hugging Face and Goodfire
Base Labs, the research group Baseten spun up earlier this year, will develop and publish methods for training and monitoring open models.
AI agents introduce new attack surfaces — prompt injection, agent hijacking, data exfiltration, hallucinated vulnerabilities, and unintended actions. As agents become more autonomous, security and reliability signals become decision-critical for any builder shipping agent-powered workflows.
This cluster covers AI security incidents, vulnerability disclosures, agent safety evaluations, and reliability signals from security researchers and frontier labs. Each signal links to the primary source for full context.
Use this page as a security watch list: catch emerging threats, review lab safety evaluations, and harden your agent pipelines before deploying to production.
Base Labs, the research group Baseten spun up earlier this year, will develop and publish methods for training and monitoring open models.
Today, I’m talking with Mustafa Suleyman, the CEO of Microsoft AI. As you’re no doubt aware, the biggest story in tech right now is the spiraling debate about AI safety and regula…
On a sunny July day in Berkeley, California, the country's top AI safety researchers gathered on an unmarked floor of an unmarked building. They had come together for a "war room"…
Anthropic and OpenAI want to embed independent safety evaluators inside their AI labs. Researchers welcome the unprecedented access, but warn meaningful oversight requires transpa…
Direct preference alignment methods are widely used to align large language models (LLMs) with human preferences because of their computational and memory efficiency. However, lik…
Emergent coordinated behaviors of AI agents are starting to present critical safety risks. A key phenomenon driving these behaviors is the rapid formation and spread of beliefs ab…