This is a scoped research dossier, not a completed systematic review or an independently tested result. It identifies methods, questions and source trails for future reporting.
Long-horizon tasks
Task splitting may help progress but also compounds errors. Each step needs an observable success criterion and recovery path.
Checkpoints and budgets
Tool-call ceilings, timeouts and approval gates keep a plan bounded. A model should not be able to expand its own authority mid-run.
Evidence of progress
Claims like “completed” should be reconciled with artifacts, tests and external state, not accepted as a self-report.
What would count as evidence?
Archive the stated plan, actual tool events and independent outcome verification.
Documents to examine
- OpenAI Agents SDK — Running Agents
- NIST AI RMF
These are starting points, not claims that every document has been independently reproduced.
Read our cited field note →Edition 1.0 · 09 October 2026
Initial research brief published. No earlier revisions or submitted public corrections are claimed.
Suggest a documented correction ↗