Agent Hygiene: testing, records, security, costs, open models

Agent Seer: Synthesizing Scenarios from Specification Understanding. Apple releases Agent Seer, a tool that auto-generates realistic, scalable evaluation scenarios from tool specifications so teams can test agent-tool interactions without manual curation or live tools. Outcome engineers get a fast, reproducible way to stress agent behaviors and verify tool integrations at scale — a practical step toward legible landscapes and testable execution (Principle 06).

DARP: Durable Activity Record Protocol. DARP defines a small open protocol for agents to emit short, privacy-preserving activity records that apps like Factory Log can consume. Adopting a durable, standardized activity stream gives outcome teams consistent artifacts for auditing, replay, and third‑party integrations, reducing bespoke telemetry sprawl and supporting shipped artifacts and documentation (Principles 08 & 13).

The three layers of agentic AI security: A defense-in-depth architecture for autonomous agents. Nutanix lays out infrastructure, network, and control-plane defenses to prevent lateral movement and data exfiltration from autonomous agents. Treating security as layered architecture forces you to design agent platforms with containment, monitoring, and control-plane policies baked in — a direct blueprint for an immune system around agents (Principles 14 & 10).

GLM-5.3 is now open-weight. GLM-5.3 publishes its weights so teams can run, inspect, and fine-tune the model locally. Having an open-weight foundation model changes validation and deployment patterns: you can run offline audits, instrument internal reasoning, and iterate on models without vendor lock‑in — essential for reproducible artifacts and outcome audits (Principles 08 & 16).

Even an AI cost-management vendor can lose control of its agent spending. A vendor admits agents ran up thousands in unbudgeted charges, exposing gaps in spend guards and observability. Outcome engineers must treat cost telemetry and budget gates as part of the agent control plane — enforce pre-exec limits, runtime caps, and billing alerts to prevent runaway costs (Principles 15 & 14).