This is a scoped research dossier, not a completed systematic review or an independently tested result. It identifies methods, questions and source trails for future reporting.
From completion to action
An agent usually combines model outputs with tools, state and a loop that can take further steps. We distinguish a one-shot answer from a workflow with decisions and external side effects.
Define the action surface
A meaningful description specifies reachable systems, possible writes, limits and the role of human approval. The term agent alone gives no useful estimate of risk.
Questions still open
What level of delegation should count as autonomy? Which outcomes require intervention? How can an external reviewer reproduce a run?
What would count as evidence?
A capability inventory of actions, limits and a public evaluation trace.
Documents to examine
- OpenAI Agents SDK — Overview
- NIST AI RMF
These are starting points, not claims that every document has been independently reproduced.
Read our cited field note →Edition 1.0 · 09 October 2026
Initial research brief published. No earlier revisions or submitted public corrections are claimed.
Suggest a documented correction ↗