futureofagents.org
RESEARCH DOSSIER / 2026.10INITIAL EDITION · OPEN FOR REVIEW
← Back to research index

Reproducibility before leaderboard position

Evaluation & Benchmarks.

Benchmarks are measurements under assumptions, not universal capability certificates.

EVIDENCE STATUS / LAUNCH

This is a scoped research dossier, not a completed systematic review or an independently tested result. It identifies methods, questions and source trails for future reporting.

01

Build a reliable task set

Report datasets, versions, scoring rules, contamination risks and representative negative cases. Hidden rubrics should not hide measurement design.

02

Measure interventions

Useful metrics include success, time, expense, policy violations, false completions and how often a human had to rescue the system.

03

Adversarial scenarios

Include tool failures, prompt injection attempts, stale context and cross-agent confusion. Separate detection accuracy from ability to recover.

SUGGESTED VERIFICATION METHOD

What would count as evidence?

Create a versioned evaluation card with tests, failures, budgets and confidence limits.

STARTING SOURCE TRAIL

Documents to examine

  • NIST AI RMF
  • OWASP — Excessive Agency

These are starting points, not claims that every document has been independently reproduced.

Read our cited field note →
EDITORIAL / VERSION RECORD

Edition 1.0 · 09 October 2026

Initial research brief published. No earlier revisions or submitted public corrections are claimed.

Suggest a documented correction ↗