A minimum reporting standard for AI evaluations
A short proposal for the least an evaluation result has to disclose before it can be compared with another result or repeated by anyone else.
Topic
How to measure what an AI system can do, and how to report the result so it can be checked and compared.
3 entries
A short proposal for the least an evaluation result has to disclose before it can be compared with another result or repeated by anyone else.
Measuring how language-model agents fail over long tasks. Per-step accuracy barely separates the systems; end-to-end success separates them a lot.
The code I use to run and record my own evaluations, released so the results here can be reproduced.