Research
Research
The four areas this practice works in, and what is published under each.
The work falls into four areas. They overlap, and the boundaries matter less than the questions.
Areas
Evaluation
How to measure what an AI system can do, and how to report the result so it can be checked and compared.
3 entriesReliability
How systems fail over long tasks, how errors compound, and what recovers from them.
1 entryInterpretability
What can be established about what is happening inside a model, and how much weight those findings carry.
1 entryOpen science
Code, datasets and reporting practice released so that others can repeat the work.
3 entriesHow questions get picked
A question is worth taking on here if it can be answered properly by one person with modest compute, and if the answer would change what someone does. Questions that need a large team or a frontier training run are not ones this practice can honestly attempt.