Agents as Engineers: Testing, Sandboxes, Terminals, and Trust
Cognition helps Devin test its own work with GPT‑6 Astra. Devin uses GPT‑6 Astra to autonomously test and prove its software, reducing engineers’ review burden and accelerating shipping. Embedding agent self‑tests and verifiable proofs is a practical way to scale delivery without losing accountability — Principle 14.
Perplexity trusts GPT-6 Astra with end-to-end systems. Perplexity delegates autonomous edits, production management, and communications to GPT‑6 Astra with far fewer human check‑ins. Treat this as a wake‑up call for orchestration and operational risk controls if you plan to let agents drive outcomes — Principle 09.
Quoting Boris Cherny. Anthropic enforces strict automated guardrails—lint, tests, fuzzers, and reviews—for Claude‑written production code. Mirror that pipeline: if agents produce code in your system, you need CI, fuzzing, and human review gates to keep outputs reliable and auditable — Principles 10 & 14.
gpty: Godot + Rust PTY multiplexer for terminal panes and AI control. gpty exposes a tiling PTY grid and JSON‑RPC/MCP control so agents can spawn panes, inject input, and observe terminal output without brittle TUI scraping. Direct terminal integration improves observability and deterministic tool integration for agents acting in developer environments — Principles 06 & 11.
Sprites, disposable Linux VMs for YOLO AI coding. Sprites delivers instant, disposable Linux VMs with preinstalled AI tooling, enabling safe, ephemeral AI coding sandboxes and web‑accessible checkpoints. Use disposable VMs to contain agent experiments, prevent persistent side effects, and make rollbacks and audits straightforward — Principle 07.