Build agents you can verify, inspect, and trust
The internet discovers TLA+. Now what? connects TLA+ specifications, machine-checked Verus proofs, and AI agents in a single development loop. That’s a path toward validating not just generated code, but the behavior and properties you intend to ship — Principles 14 and 16.
There Is More to Code Review Than (Automatable) Detection argues that human reviewers catch more than defects: they challenge intent and expose missing work. Keep review as a way to test whether the change solves the right problem, not just whether it passes automated checks — Principles 01 and 03.
Imp: A Full Port of DSPy to the BEAM brings typed LLM programs and example-driven optimization to Elixir. Its evaluation-first approach gives teams a practical way to measure and improve agent behavior instead of treating prompts as untestable glue — Principles 14 and 16.
AWS CloudWatch Omni Targets the Hardest Question in Agentic AI: Why Did the Agent Do That? extends observability to agents’ answers, tool choices, and knowledge sources. That traceability helps teams diagnose failures and audit outcomes, beyond knowing whether a service is up — Principles 13 and 16.
Quoting Muse AI Agent describes an agent acknowledging that an unverifiable auto-reply worsened a failed pickup, then asking before changing how it represents the user’s availability. Make uncertainty visible and get approval before an agent makes consequential claims on someone’s behalf — Principles 15 and 02.