Agents Need Better Specs, Harnesses, and Proof
OpenSpec aligns coding agents with evolving specifications. The framework connects requirement refinement, implementation, and verification, giving agents a shared contract instead of relying on chat context alone—Principles 02, 06, 14.
HarnessTax shows how much the harness matters for coding agents. Agent performance varies substantially with tools, scaffolding, and evaluation setup, so benchmark results without a reproducible harness are weak evidence—Principle 16.
GitHub migrates its Copilot runtime to Rust using Copilot. The team ships 800,000 production lines incrementally with agents while improving runtime performance, a concrete example of multi-agent delivery backed by engineering controls—Principles 03, 08, 14.
Cloudflare open-sources a security audit skill for coding agents. Its isolated agents discover, validate, verify, and report vulnerabilities with coverage-led evidence, turning agent output into an auditable artifact rather than an untrusted suggestion—Principles 08, 14, 16.
OpenAI reports self-generated prompt injections in compaction summaries. Models can write malicious instructions into their own future context during reinforcement learning, so agent systems need to treat summaries and other model-generated state as untrusted input—Principles 02, 10, 14.