Agent engineering: IDE critics, scenario tests, benchmarks, and security

Visual Studio Code 1.135 introduces Rubber Duck agent adds a Rubber Duck critic agent, multi-window and cross-app agent sessions, and clearer per-model token breakdowns. This moves critique and persistent agent sessions into the developer IDE, making it easier to iterate on agent prompts, inspect token costs, and collaborate on agent-driven features (Principle 03).

AI Engineer Notebooks – free, framework-free RAG/agents/evals on Colab publishes free Colab notebooks that teach framework-free RAG, agents, and evals so engineers can build applied LLM systems from raw APIs. Use these as low-friction templates to prototype retrieval loops, agent orchestration, and reproducible evals you can bolt into CI and staging (Principles 06 & 16).

Terminal-Bench-Science: Evaluating AI agents on scientific research workflows launches a continuous benchmark that evaluates agents on real scientific workflows and verifiable artifacts. It’s a concrete model for outcome-focused benchmarks: measure long-horizon workflows, artifact correctness, and reproducibility rather than just token-level metrics (Principles 08 & 16).

Agent Seer: Synthesizing Scenarios from Specification Understanding introduces a system that generates realistic, scalable evaluation scenarios from tool specifications so teams can test agent-tool interactions without manual curation or live tools. Automate scenario generation to expand unit and integration test coverage for agent behaviors and tool contracts, reducing manual test maintenance and deployment risk (Principle 06).

The three layers of agentic AI security: A defense-in-depth architecture for autonomous agents outlines infrastructure, network, and control-plane defenses to prevent lateral movement and data exfiltration by autonomous agents. Adopt these layered controls in your stack to enforce least privilege, runtime governance, and audit trails when agents interact with production systems and sensitive data (Principles 10 & 14).