THE Tech Island THE Immune System
Is Sandboxing Sufficient to Contain Rogue Agents?
simonwillison.net
It was never about the code
Outcome Development
Next steps for engineering and product development in an agentic world
Most recent 0d ago Software Doesn't Need to Be Readable Anymore. It Needs to Be Explainable.
THE Tech Island THE Immune System
simonwillison.net
THE Law THE Gate
fortune.com
THE Law THE Immune System
nytimes.com
THE Law THE Immune System
infoworld.com
THE Immune System THE Gate
aisi.gov.uk
THE Law THE Gate
washingtonpost.com
THE Law THE Immune System
thehill.com
THE Tech Island THE Law
openai.com
THE Map THE Graph
github.com
THE Law THE Gate
cnbc.com
The central risk is no longer whether frontier AI can do consequential work, but whether its safeguards hold when capability meets real access. The New York Times report on OpenAI’s internal safety-testing warnings raises a hard operational question: if teams raise concerns about inadequate testing and those warnings do not change release decisions, what does a safety process actually control? The report describes allegations, not a final finding, but the gap it spotlights is familiar to anyone shipping agent systems: documented safeguards matter only if they can block deployment or constrain behavior.
Anthropic’s analysis of GLM-5.3’s cyber capabilities makes that gap more concrete. The researchers report that the model can construct end-to-end exploits in simulated tests and that simple techniques bypass its limited safeguards. Read alongside the OpenAI report, this shifts attention from broad safety assurances to evidence about specific capabilities, attack conditions, and failure modes. Evaluation must test what a system can do under pressure—not just what its policy says it should do. That is the force of Audit the Outcomes: a safety claim needs a reproducible test and a decision attached to the result.
The day’s security releases point toward a practical response: narrow authority at the moment of action. Okta’s agent runtime gateway places authorization in the execution path, where organizations can govern or block an action after an agent has authenticated. Equals Money’s MCP server draws an even simpler boundary: customer AI tools may read financial data but cannot initiate payments. These are complementary controls. Identity answers who or what is acting; scoped permissions define which actions are available. Both make the Gate and the Immune System enforceable properties of a workflow, not entries in a policy document.
Assurance also has to persist after release. Livenerf proposes a preregistered, reproducible baseline for detecting whether a frontier model’s performance drifts after launch. That matters because teams often build evaluations around a model snapshot, then depend on a changing service in production. A stable baseline turns “the model seems worse” into a question that can be investigated—and gives operators a reason to retest integrations when behavior shifts.
Together, these stories describe a move from trust-by-assertion to trust earned through controls and continuing evidence. Runtime authorization limits what an agent can do; adversarial tests expose where safeguards fail; repeatable baselines reveal when the system changes. None substitutes for accountable release decisions, but each makes those decisions harder to obscure. Build agent safety as a loop: test capability, constrain authority, and keep checking the behavior you actually depend on.
Who's instigating and driving conversations
Reach
First mover
Coverage
Reach
First mover
Coverage
Reach
First mover
Coverage
Share of trailing 7-day coverage per frontier lab
Anthropic OpenAI Google Meta DeepSeek Mistral xAI
Per-article sentiment with 7-day net approval
7-day net approval
Trailing 7-day balance of creation vs oversight principles
Building (warm) Governing (cool)
Stories per principle, last 7 days