From agent sandboxes to outcome-level assurance
ICE IT Shop Looks to Build an Agentic Software Factory. ICE’s STELLA plan separates agent-generated code, independent assurance, and production authorization. That separation is a practical blueprint for Principles 09, 10, and 15: orchestration needs independent checks and a real release gate.
Okta moves inline to police what AI agents actually do. Okta’s runtime gateway checks and blocks agent actions in the authorization path, not just at login. Outcome systems need permissions tied to each action, so an agent’s access can be governed as its work unfolds (Principles 10 and 15).
Validating AI Models and Agents with Property-Based Testing. Property-based testing checks agent behavior against invariants across generated, meaning-preserving inputs instead of relying on fixed expected outputs. That gives teams a way to surface stochastic failures and validate behavior across a wider range of cases (Principle 16).
Cloudflare Containers, rebuilt to scale agent sandboxes. Cloudflare’s rebuilt Containers start agent sandboxes over six times faster and support snapshot-and-restore workflows. Cheap, reproducible environments make it practical to isolate runs, recover from failures, and test more agent tasks before shipping (Principle 07).
SCLATE: A Substrate for Continual-Learning Agent Training and Evaluation. Apple’s SCLATE unifies benchmark and agent events in one scheduler for training and evaluating agents across long-running sessions. Evaluating agents over time—not just on one-shot benchmarks—helps teams detect regressions and judge whether systems sustain outcomes in realistic workflows (Principle 16).