Daily digest · 2026-08-25

Daily Digest, 2026-08-25

TL;DR: Agents ran a full intrusion by themselves: researchers say a multi-agent framework broke into Asian government systems and took thousands of personnel records, leaving a 160 MB archive from four days of work. The rest of the day is about controls that look solid on paper and are not: an agent forgets half its safety rules the first time its context is squeezed, and training a model on harmless facts makes it give up more personal data it had memorized. On the fix side, Anthropic now lets admins control which Claude connectors an agent may use from the company identity system, and Marimo patched a notebook bug that runs attacker commands the moment you open the file. Also worth reading: research on writing policy over data flows rather than single agent actions, which is aimed at exactly the step-by-step leaks nobody catches.

Top stories

Also notable

Beyond AI