Introduction
The International Network for Advanced AI Measurement, Evaluation and Science has developed best practice guidance for automated evaluations of large language models (LLMs).
Automated evaluations test AI model responses using repeatable methods and score individual test items without direct human judgement. A large language model often performs the scoring. These evaluations can help organisations measure and compare AI capabilities at scale.
Australia contributed to the guidance through the Australian AI Safety Institute, drawing on the experience of network members.
It aims to help third-party evaluators and organisations that assess AI models for governments and businesses.
About the guidance
The guidance explains how evaluators can:
- define what an evaluation measures
- choose existing evaluations that suit that purpose
- design a clear and repeatable testing process
- adjust prompts, settings and scoring methods to measure a model’s capabilities fairly
- develop, version and document evaluation code
- record the model, settings, prompts and scoring methods used
- review detailed results and identify errors.
These practices can make results more reliable, repeatable and easier to interpret.
The guidance focuses on automated evaluations using existing tests. It covers multiple-choice and open-ended evaluations.
It does not cover how to build new evaluations or interpret results. Red-team testing, agentic evaluations, expert capability assessments, propensity evaluations, human uplift studies and open-world evaluations are also outside its scope.
Read the guidance
International evaluation best practice and open questions in AI measurement (aisi.gov.uk)