futureofagents.org
EDITORIAL PROTOCOL / 02

Trust is a
measurable claim.

Our standard for evaluating agent systems, multi-agent architectures and operational security.

STANDARD 01 / PROVENANCE

Name every system.

Record publication date, model and tool versions, orchestration framework, tool manifest, data sources, operating environment and who conducted the review.

STANDARD 02 / ACCESS CONTROL

Show the permission boundary.

Enumerate read, write and destructive actions; roles, scopes, secrets handling, approval gates, command boundaries and isolation conditions.

STANDARD 03 / EXPERIMENTS

Publish a reproducible method.

Define task set, comparison baseline, scoring code, intervention policy, repetition count, costs, failure categories, and known contamination or selection effects.

STANDARD 04 / ADVERSARIAL CASES

Test the failure surface.

Include indirect prompt injection, compromised tool responses, deceptive peer messages, timeouts, tool unavailability, unintended state changes and partial completion.

STANDARD 05 / LIMITATIONS

Keep uncertainty legible.

Separate observed properties from hypothetical architectures. State missing tests, bias, sample limits, excluded environments, and the difference between voluntary guidance and binding requirements.

STANDARD 06 / REVISION

Publish corrections and incident notes.

Material changes should be dated, linked to evidence, and understandable without seeing the previous edition. Operational secrets and personal data must be redacted.

CURRENT STATUS / INITIAL EDITION

Our research index is an agenda, not a results database.

At launch we provide topic briefs, standards and a cited methodology article. We do not imply that independent agent tests, incident investigations or comparative benchmarks have already been completed.