Daily digest · 2026-09-14

Daily Digest, 2026-09-14

TL;DR: Today's batch is mostly about proving what an agent actually did, and about the gap between what benchmarks claim and what they measure. A new record format argues agent logs are unusable as evidence because they can be edited, can't be replayed, and don't say whether secrets were stripped; two separate papers show benchmark validity decays across pipeline stages and that a real audit of agent eval files found metrics measuring the wrong thing. On the data side, researchers show that ordinary public posts can be aggregated into sensitive attributes with links back to the source, and that one unredacted prompt seeds personal data into memory, tool arguments and logs across an entire pipeline. Brussels also sent a US-EU biometric border-sharing framework to the Council ahead of a year-end Visa Waiver deadline.

Top stories

Also notable

Beyond AI