Agent Architecture, Memory, and Latency: Tools & Case Studies

Asana cleared 5 years of engineering work in 2 weeks with Codex. Asana uses Codex to replace five years of testing work in two weeks for about $12K, slashing backlog and cost. This is a concrete outcome-engineering playbook for replacing manual QA backlog with agentic automation—operationalizing Principle 04 (The Backlog Is Dead).

Agent, skill, or MCP? Which to use and when to use them | AWS’ Clare Liguori. Clare Liguori lays out an MCP-first decision framework so teams can choose between skills, agents, or stateless MCP servers to scale enterprise AI without heavy scaffolding. Outcome engineers get actionable architecture guidance for layering state, context, and orchestration—useful for semantic-layer and context-engineering decisions (Principle 06/11).

Open Sourcing Comfy MCP on Local. Comfy open-sources a local MCP that runs hardware-aware ComfyUI workflows across local and cloud environments. Local, hardware-aware MCPs give you runtime control over latency, privacy, and reproducibility—an explicit step toward building isolated, auditable agent islands (Principle 07/06).

Agentic AI has a latency problem that more compute won’t solve. The piece shows agentic stacks suffer end-to-end latency from multi-hop, CPU-bound work and that GPUs alone won’t meet real-world SLOs. Outcome engineers must benchmark multi-stage hops, optimize orchestration, and redesign pipelines for tail latency—an operational constraint on agent orchestration and validation (Principle 09/16).

How Much Memory Does Your Agent Actually Need?. IBM Research demonstrates that calibrated memory dosing and selective retrieval often outperform full guideline injection, saving tokens while improving task completion. Use this as a practical guide: tune retrieval frequency, compress memories, and prefer on-demand context to reduce cost and improve reliability (Principle 06/12).