Failure compounds faster than per-step accuracy suggests
Measuring how language-model agents fail over long tasks. Per-step accuracy barely separates the systems; end-to-end success separates them a lot.
Topic
How systems fail over long tasks, how errors compound, and what recovers from them.
1 entry
Measuring how language-model agents fail over long tasks. Per-step accuracy barely separates the systems; end-to-end success separates them a lot.