Agent Ops — Observability, Sandboxes, Plugins, Stateless MCP, CI

FriskAI launches with $3.6M to show enterprises what their AI agents are doing. FriskAI launches a runtime recorder and auditor to capture what enterprise AI agents do in production, adding visibility and accountability across agent activity. Outcome engineers get a practical path to build auditable agent pipelines and satisfy recordkeeping and validation requirements (Principles 02, 13).

Electric joins Databricks to bring WASM Postgres to AI agent sandboxes. Databricks integrates Electric’s PGlite WASM Postgres and real-time sync so agents can run a local Postgres in-sandbox that syncs back to Lakebase. That makes stateful SQL workflows reproducible inside sandboxes, letting teams test data-connected skills and ship artifacts from isolated islands (Principles 07, 09).

Agent Plugins package your skills, tools, and more. Google publishes Agent Plugins 1.0.0 as a portable plugin spec to package agent skills, tools, and MCP-compatible servers for cross-vendor reuse. Portable skills reduce integration friction, enable registries and CI validation of capabilities, and make agent behavior more modular and auditable (Principles 06, 11).

Scaling AI Agent Infrastructure with the MCP Stateless Updates. Google updates the Model Context Protocol with a stateless spec and SDKs that replace stateful agent servers, enabling cloud-native horizontal scaling and long-running task patterns. Moving state out of the agent host simplifies autoscaling, deployment, and CI-driven rollouts—important if you treat agents as production infrastructure (Principles 06, 11).

Agent and Model Evaluations in Gemini Enterprise Agent Platform are now GA. Google launches a GA evaluation service with unified metrics, adaptive rubrics, and CI integration for testing agent and model quality. Built-in evaluation pipelines give outcome engineers the gates needed to audit, validate, and continuously measure outcomes before agents reach users (Principles 14, 16).