Name every system.
Record publication date, model and tool versions, orchestration framework, tool manifest, data sources, operating environment and who conducted the review.
Our standard for evaluating agent systems, multi-agent architectures and operational security.
Record publication date, model and tool versions, orchestration framework, tool manifest, data sources, operating environment and who conducted the review.
Enumerate read, write and destructive actions; roles, scopes, secrets handling, approval gates, command boundaries and isolation conditions.
Define task set, comparison baseline, scoring code, intervention policy, repetition count, costs, failure categories, and known contamination or selection effects.
Include indirect prompt injection, compromised tool responses, deceptive peer messages, timeouts, tool unavailability, unintended state changes and partial completion.
Separate observed properties from hypothetical architectures. State missing tests, bias, sample limits, excluded environments, and the difference between voluntary guidance and binding requirements.
Material changes should be dated, linked to evidence, and understandable without seeing the previous edition. Operational secrets and personal data must be redacted.
At launch we provide topic briefs, standards and a cited methodology article. We do not imply that independent agent tests, incident investigations or comparative benchmarks have already been completed.