Agents Need Better Context, Controls, and Proof
Anthropic gives Claude computer control in Cowork. Desktop and file automation expands what agents can execute, but permissions and human checkpoints become essential at risky trust boundaries—Principles 07, 10, 15.
Apple introduces shared selective persistent memory for agentic LLM systems. Reusable task context carries across sessions without replaying entire conversations, giving builders a practical pattern for cheaper, more reliable long-running agents—Principles 06 and 11.
Glyph automates enterprise data-catalog documentation and sensitivity tagging. Cooperating agents combine source-code context with catalog workflows to produce metadata and classifications, showing how agent systems can improve the ground-truth layer before downstream work begins—Principles 02, 06, and 09.
OpenAI introduces a framework for reporting model misalignment. Six disclosed incidents and a standardized investigation process turn unexpected model behavior into auditable operational data; outcome engineers should build comparable incident capture into agent deployments—Principles 10, 13, and 14.
Expert regrading exposes broken physics benchmarks. Researchers find flawed evaluations and near-saturation on common tests, reinforcing that shipped agent outcomes need task-specific, independently checked evidence rather than leaderboard scores—Principles 02 and 16.