Humanity's Last Exam: A Benchmark of Expert-Level Academic Questions to Assess AI Capabilities
Center for AI Safety, Scale AI, HLE Contributors Consortium
Humanity's Last Exam is a multimodal benchmark of 2,500 expert-level closed-ended academic questions designed to measure frontier AI capabilities beyond saturated benchmarks.
Benchmarks are important tools for tracking the rapid advancements in large language model capabilities. However, benchmarks are not keeping pace in difficulty: LLMs now achieve more than 90% accuracy on popular benchmarks such as MMLU, limiting informed measurement of state-of-the-art LLM capabilities. In response, Humanity’s Last Exam (HLE) introduces a multimodal benchmark at the frontier of human knowledge, designed as an expert-level closed-ended academic benchmark with broad subject coverage. HLE consists of 2,500 questions across dozens of subjects, including mathematics, humanities, and the natural sciences. It was developed globally by subject-matter experts and contains multiple-choice and short-answer questions suitable for automated grading. Each question has a known solution that is unambiguous and easily verifiable but cannot be quickly answered by internet retrieval. State-of-the-art LLMs demonstrate low accuracy and calibration on HLE, highlighting a marked gap between current model capabilities and the expert human frontier on closed-ended academic questions.
@article{centerforaisafety2026benchmark,
author = {{Center for AI Safety} and
{Scale AI} and
{HLE Contributors Consortium}},
title = {A Benchmark of Expert-Level Academic Questions to Assess {AI}
Capabilities},
journal = {Nature},
volume = {649},
number = {8099},
pages = {1139-1146},
year = {2026},
doi = {10.1038/s41586-025-09962-4}
}
