Tagged:
case-study2 posts
-
Physics Is All You Need? What One Physicist's AI Supervision Log Reveals About Trustworthy Agent Output
A new ICML 2026 paper presents a rare quantified case study: 57 agent sessions, 15 bugs, and three failures that oracle tests could not catch. The central finding — that supervision design, not model capability, determined whether the agent's output was trustworthy — has direct implications for AI security.
-
423 Security Fixes in One Month: Inside Mozilla's AI-Powered Vulnerability Pipeline
Mozilla shipped 423 Firefox security fixes in April 2026 — nearly 20x the monthly average — by combining Anthropic's Claude Mythos Preview with a custom agentic harness. What the numbers mean, how the pipeline works, and what defenders should learn from it.