Ship-safe agents: governance, memory, sandboxes, auth, testing
Build your own company brain: the enterprise AI playbook from Stripe’s engineering team lays out Kai and a governance-first AI stack — context layers, sandboxes, and a skill platform to safely scale agents to every employee. Outcome engineers should model orchestration, context engineering, and access controls on this pattern to keep intent aligned with delivery (Principles 03 & 10).
A deep dive into exe.dev explains instant, persistent Linux VMs with HTTPS endpoints, pooled compute, and a built-in coding agent that give developers and agents safe, reproducible sandboxes. Use persistent VMs to isolate side effects, reproduce agent runs, and speed iteration — a practical Build-the-Island approach for agent infra (Principles 07 & 12).
Engrim — Universal Local-First SQLite Memory Engine for AI CLIs provides a local, model-agnostic episodic memory using SQLite so CLIs and lightweight agents retain state when you swap models. Decoupling memory from model choice makes agents replaceable and auditable, improving grounding and state management across deployments (Principles 06 & 11).
How AI agents can log in without seeing your passwords documents agentic autofill and SDK patterns that let agents authenticate without exposing secrets to language models. Implement vault-backed autofill and agent access SDKs to shrink blast radius and enforce least privilege for agents interacting with real accounts (Principles 10 & 06).
How well do agents use test/verification techniques? shows that simple testing prompts rarely improve agent-generated code correctness and that many testing techniques fail to meaningfully boost reliability. Treat this as a red flag: build independent verification, property-based checks, and audit hooks into agent workflows rather than relying on prompt-level testing alone (Principles 16 & 14).