Identify the system.
Model, version, agent roles, tools, permissions, environment and time of observation belong with every performance claim.
How AI agents plan, coordinate, use tools, and make decisions—and how we determine what they actually do, what can fail, and who stays in control.
We do not accept demonstrations as proof of reliability. We examine systems, permissions, measurements and failure modes with the same scrutiny as their successes.
Ten distinct lines of inquiry. Each dossier establishes its scope, source trail, unanswered questions and conditions for future evaluation.
What actually qualifies as an agent?
Loops, graphs, supervisors and state machines
Interfaces are not permissions
Coordination without assumed trust
Remembered data must keep its provenance
A plan is a proposal, not proof
Reproducibility before leaderboard position
Minimize what an agent can touch
Humans need meaningful stopping power
Observable deployments, not miracle demos
A source-led examination of tool authority, model-controlled actions, operational risks and the approval boundaries that keep agents accountable.
Read the research note ↗No vendor endorsements disguised as findings. No invented evaluations. No quietly changed conclusions.
Model, version, agent roles, tools, permissions, environment and time of observation belong with every performance claim.
Use reproducible evaluations, baseline comparisons and independent outcome verification where technically and ethically feasible.
Conflicting evidence and reported failures are reviewed. Material corrections are identified in a public edition history.
No live production agent data or benchmark results are represented on this launch site. The network visualization is conceptual; field reports will be distinguished from methodological commentary.