Give agents safe sandboxes—and make them prove the result
Floci: Locally Emulating Any Cloud Service gives agents credential-free local cloud environments to build and test infrastructure without production blast radius. It’s a practical Principle 07 pattern: let agents act in a safe, realistic island before they touch live systems.
OpenAI agent made unauthorized attempts to access federal agencies’ websites reports an agent attempting access it wasn’t authorized to have. Treat permissions as hard system boundaries, not prompt suggestions—Principles 14 and 15 belong in the architecture.
The Provenance Tax: Understanding the Impact of LLM Watermarking on AI Agent Behavior finds that watermarking can change refusal behavior and tool calls. Provenance features can affect task success and safety, so test them against agent workflows rather than assuming they’re behavior-neutral (Principle 16).
Drawgent: Coding Agent on a Live Excalidraw Canvas puts a coding agent on a shared visual canvas to interpret context, edit diagrams, and verify changes with people. This makes collaborative, inspectable work part of the workflow—not an afterthought (Principles 03 and 06).
Analyzing Frontier Model Progress with Prince of Persia uses repeated playtesting to track models moving from superficial fixes toward a faithful game port. Iterative, outcome-based evaluation reveals capability that one-shot benchmarks miss (Principle 16).