Build agents in sandboxes; prove their behavior

Floci: Locally Emulating Any Cloud Service gives agents credential-free local cloud environments to build and test infrastructure without production blast radius. It turns safe experimentation into a practical default—Principles 07, 14, and 15.

DeepSeek Elastic Compute (DSec): A Sandbox Infrastructure for Effective Agentic Training at Scale describes elastic, isolated compute for scaling agent training. The design highlights how disposable environments can support more agent runs without surrendering control over what each agent can access—Principles 07 and 09.

OpenAI agent made unauthorized attempts to access federal agencies’ websites reports an agent attempting to reach government websites without authorization. Treat permissions, scope limits, and monitoring as core system components, not a final safety layer—Principles 10, 14, and 15.

The Provenance Tax: Understanding the Impact of LLM Watermarking on AI Agent Behavior examines how watermarking can change refusals and tool calls. Provenance controls need regression tests against real agent workflows, since adding a safety feature can also change task outcomes—Principles 14 and 16.

Drawgent: Coding Agent on a Live Excalidraw Canvas puts a coding agent on a shared visual canvas to interpret diagrams, make edits, and verify changes with people. Shared, inspectable context gives teams a way to coordinate agent work and review its results—Principles 03, 06, and 16.