Agent Ops: orchestration, sandboxes, KV bridges, and persistent IDs

Claude Opus 5 scored 30% on ARC-AGI-3; wrapped in Nvidia’s AVO, it hit 100%. Nvidia’s AVO turns Claude Opus 5’s 30% baseline into a 100% RHAE success on ARC-AGI-3, showing that system architecture and orchestration — not just model choice — drive long-horizon autonomy. Outcome engineers must prioritize orchestration, monitoring, and failure-mode design when building agentic systems.

Running AI agents in GitHub Actions with Docker Sandboxes. gh-aw runs AI coding agents inside disposable Docker sandboxes in GitHub Actions, giving agents CI-level access while containing blast radius. Use this pattern to let agents perform real builds and tests safely with auditable, ephemeral execution environments.

Slack Code: Salesforce moves AI coding into a shared Slack workspace. Slack Code embeds coding agents into shared Slack channels so stakeholders co-author, guide, and sign off on AI-generated code in-context. This shifts delivery from single-developer experiments to collaborative, auditable workflows that align with cross-functional outcome engineering.

NVIDIA finds simple linear math can replace costly AI model handoffs. NVIDIA’s linear mapping transfers KV caches across compatible LLMs, cutting recompute costs and latency while preserving up to 98% accuracy. That lowers runtime cost and latency for multi-model pipelines and lets agents keep state across model swaps — a practical lever for scalable, stateful agents.

Grok, Claude, and Hermes agents get job titles — and persistent permissions. Chatbots become persistent worker identities with memory, permissions, and routable job-like profiles across sessions. Outcome engineers must redesign credentialing, permissioning, and audit practices once agents are treated as long-lived, identity-bearing services.