Agent Ops: routers, observability, budgeted agents, cheap search, pricing

Nvidia moves into hot market for model routers. Nvidia launches NeMo Switchyard to route prompts across models, lowering inference costs and enabling system-of-models for efficient agent workloads. Outcome engineers can use model routing to compose specialist models, optimize cost/latency tradeoffs, and build predictable multi-model pipelines (Principles 09 & 12).

Dynatrace agrees to acquire Arize for $915M to accelerate AI observability. Dynatrace buys Arize to embed AI observability and lifecycle telemetry across its platform for monitoring, drift detection, and root-cause analysis. Practitioners should treat observability as infrastructure — instrument outcomes, lineage, and alerts now to make agentic systems maintainable and auditable (Principle 14).

Mole — Deep research agent for your terminal. Mole enforces per-run model budgets, verifies source quotations, and analyzes local data in a privacy-preserving terminal agent. Use this pattern as a reference implementation for budget enforcement, provenance checks, and local-data isolation when you build developer-facing agents (Principles 02 & 07).

Introducing Toast 1. Toast 1 ships a low-cost, high-throughput search subagent that promises frontier-quality retrieval up to 10× cheaper and 12× faster. Slot specialized retrieval agents like Toast into your orchestration layer to cut inference spend and speed up agent loops without reworking your core models (Principles 06 & 09).

DeepSeek-V4-Pro GA Release and Peak/Off-Peak Pricing Update. DeepSeek releases V4-Pro with agent-focused upgrades, flexible reasoning modes, OpenAI Responses support, and peak/off-peak API pricing that discounts off-peak runs by 50%. Model-level features plus time-based pricing change scheduling, SLAs, and cost models — design your outcome pipelines to exploit off-peak windows and reasoning modes for predictable unit economics (Principles 12 & 06).