Agent Ops: routing, subagents, open weights, and budgeted research
Nvidia moves into hot market for model routers. NVIDIA launches NeMo Switchyard to route prompts across models, lowering inference costs and enabling system-of-models for agent workloads. That gives outcome engineers a production-ready router to implement model selection, fallbacks, and cost-aware routing across specialized agents — Principle 09/12.
DeepSeek-V4-Pro GA Release and Peak/Off-Peak Pricing Update. DeepSeek ships DeepSeek‑V4‑Pro with agent-focused upgrades, flexible reasoning modes, OpenAI Responses support, and new peak/off‑peak API pricing (off‑peak 50%). This changes agent economics and integration patterns: schedule heavy reasoning off‑peak, pick cheaper reasoning modes for background tasks, and design agents around price-aware inference — Principle 12.
Introducing Toast 1. Toast 1 surfaces frontier-quality search up to 10× cheaper and 12× faster as a specialized retrieval subagent for agent stacks. Use it to split retrieval from reasoning, lower token and latency costs, and compose retrieval subagents that scale independently of your reasoning models — Principle 09/06.
Alibaba’s new model promises Opus 4.6-level performance on your laptop. Alibaba releases open-weight Qwen3.8 variants, including a dense 27B with very long contexts that runs on consumer hardware. Open weights plus long context let outcome engineers deploy local agents with full control over data, latency, and auditing — Principle 07/16.
Mole — Deep research agent for your terminal. Mole provides a terminal research agent that enforces per-run model budgets, verifies quoted sources, and analyzes local data privately. It’s a compact pattern for building accountable, cost-aware research agents with verifiable outputs you can adapt into larger agent workflows — Principle 02/07/10.