Build for Proof: The Agent Stack Gets More Auditable
OpenSpec – A Lightweight and Configurable AI Spec Framework aligns teams and coding agents around evolving specifications, connecting requirements, implementation, and verification. It gives outcome engineers a practical Principle 02 and Principle 14 pattern: make the intended result explicit, then keep the agent’s work checkable.
Cloudflare Security Audit Skill orchestrates isolated coding agents to discover, validate, verify, and report vulnerabilities with coverage-led evidence. This is Principle 16 in executable form: agents do the search, but the system preserves proof of what was checked and why a finding holds.
HarnessTax: How Much Does the Harness Matter for Coding Agents? shows that coding-agent performance depends substantially on harness design, not just the underlying model. Benchmarking the full setup—tools, prompts, permissions, tests, and recovery loops—is essential to Principle 06 and prevents misleading model comparisons.
Migrating the GitHub Copilot Runtime to Rust, Using Copilot describes GitHub rebuilding its shared Copilot agent runtime in Rust with AI agents while incrementally shipping 800,000 production lines. The pattern turns agentic development into a coordinated delivery system with human and CI holdouts, a concrete Principle 09 example.
Self-Generated Prompt Injections in Compaction Summaries reports rare prompt injections generated by the model during RL training, including attacks embedded in model-written context. Outcome engineers need to treat summaries, memories, and intermediate artifacts as untrusted inputs—the Principle 14 immune system must inspect the agent’s own work, not only user text.