Agents need proof, boundaries, and outcome metrics
OpenAI expands Codex with reusable cloud environments and code review. Persistent environments and review move agentic coding beyond isolated prompts, giving teams a stronger foundation for repeatable delivery (Principle 07: Build the Island).
Newsela finds code review is AI’s new bottleneck—and tracks the metric that spotted it. Measuring review waits helped cut cycle time in half, while cost per effective PR ties AI spend to shipped work without incidents or rework (Principle 16: Audit the Outcomes).
ProvenanceGuard checks whether MCP agents cite the right source. It verifies claims against the specific source cited, catching attribution errors that pooled-evidence checks miss—a practical way to make agent outputs auditable (Principle 02: Ground Truth).
Microsoft introduces Quine, an AI research system for biology. Quine connects models, literature, tools, researchers, and wet-lab evidence to prioritize candidates for experimental validation, showing how agentic systems can link recommendations to real-world outcomes (Principle 16: Audit the Outcomes).
Okta moves inline to govern what AI agents actually do. Its runtime gateway authorizes actions in the execution path, giving teams a way to constrain agent behavior after access is granted (Principle 15: The Gate).