Agents, Data & Compute: Legal, infra, and tools for outcome engineers

vLLM v0.28.0 massively speeds multi-GPU inference with Kimi-K3 optimizations, speculative decoding, memory sharding, and ROCm support. Outcome engineers get cheaper, lower-latency agent loops and more predictable multi-GPU scaling — use this to redesign orchestration and capacity planning (Principle 09, 12).

Domain-Driven Agents argues for making messy legacy codebases agent-ready by building shared language and context and letting humans decide while agents execute tactical work. This gives a practical pattern for context engineering and human-in-the-loop gates so agents operate inside legible domains (Principle 01, 03, 06).

Debian votes to allow “Responsible Use of Generative AI” formally permits responsible use of generative AI while holding contributors legally and technically accountable for reviewed, tested, and compliant submissions. For outcome engineers this signals stronger upstream governance models — embed compliance checks, provenance, and testable reviews into contribution and CI pipelines (Principle 10, 15).

Sony Music and Warner Chappell sue Anthropic, Dario Amodei, and Benjamin Mann, alleging tens of thousands of copyrighted songs were used to train Claude’s LLMs alleges large-scale use of copyrighted music in training and opens multi-billion-dollar litigation. Outcome engineers must treat training and fine-tuning pipelines as legal risk vectors — invest in provenance, opt-out tooling, and immutable audit trails to keep Gate and Law aligned (Principle 10).

Samsung’s Processing-in-Memory (PIM) at Hot Chips 2026 embeds MAC arrays into LPDDR5X banks, enabling multi-bank in-memory compute with 614 GB/s internal bandwidth and standard DDR command access. This shifts where compute lives; outcome engineers should reconsider model partitioning, data layout, and latency budgets to exploit PIM for tighter agent feedback loops and lower inference cost (Principle 06, 11, 12).