Skip to main content
A serious professional discussion among diverse experts in a modern boardroom, debating the necessity and usefulness of human

Editorial illustration for 60% of Experts Say Humanity's Last Exam Is Necessary and Useful

Experts Back Humanity's Last Exam as AI Benchmark

60% of Experts Say Humanity's Last Exam Is Necessary and Useful

Updated: 4 min read

Most AI benchmarks are useless now. The good ones are too easy. The latest large language models score over 90% on tests that were considered hard just a few years ago, rendering them incapable of showing which system is actually better. So a group of researchers built something meaner.

They call it Humanity’s Last Exam. Published in *Nature* in early 2026 by the Center for AI Safety and Scale AI, it is a 2,500-question gauntlet across more than a hundred expert-level fields. It is not a memory test.

It demands reasoning and deep knowledge. Even the best AI models currently fail more than half the time. The goal is to create a contemporary, brutal successor to the Turing test.

The question is whether this new standard is a vital tool or an academic sideshow.

HLE is Truly Useful and Necessary About 60% of the opinions lean toward this collective opinion, according to which there is a technical reason why HLE is paramount at present: previous benchmarks and testing frameworks for AI systems, including not-so-old language model benchmarks like Massive Multitask Language Understanding (MMLU), became saturated or obsolete, with nearly every modern AI scoring over 90% on them. This made it impossible to truly compare the latest models against each other to determine which one is best. One salient reason why HLE is praised by many experts is that it measures whether the AI is willing to say "I don't know" instead of hallucinating about complex problems or questions it can't address. HLE is a Distraction From Real AI This skeptical viewpoint is adopted by about 30% of the opinions.

The sixty percent majority has a point. We need a test that doesn't break immediately. But the drama of the name matters.

Calling something the "last" exam frames progress as a finite race with a finish line, which is a marketing gimmick. Intelligence is not a mountain you summit. It is a set of capabilities we barely know how to define, let alone measure with a single score.

HLE is a harder test, which is useful. It is not the final word. The real work happens far from any leaderboard, in the slow engineering of systems that can actually think and, just as critically, know when they cannot.

Further Reading

Common Questions Answered

Why do 60% of experts consider Humanity's Last Exam necessary for AI evaluation?

Experts believe HLE is necessary because previous AI benchmarks like MMLU have become saturated, with nearly every modern AI model scoring over 90% on them. This saturation makes it impossible to meaningfully differentiate between the latest AI systems, so a more challenging evaluation framework is required to accurately assess their true capabilities.

How does Humanity's Last Exam differ from traditional AI benchmarks?

Unlike traditional benchmarks, HLE is designed as an extreme evaluation framework where even the most advanced AI models fail more than half the time. It represents a radical departure from conventional testing by focusing on pushing AI reasoning and deep knowledge capabilities to their absolute limits, rather than simply measuring performance on standard tasks.

What is the relationship between Humanity's Last Exam and the Turing test?

HLE is conceived as a contemporary evolution of the Turing test, maintaining the original test's goal of assessing machine intelligence but using modern evaluation methodologies. Both tests aim to measure whether AI can demonstrate human-level reasoning and understanding, though HLE employs a more rigorous and comprehensive framework suited to today's advanced AI systems.

What concern does the article raise about using HLE as the primary measure of AI intelligence?

The article warns that a hard test should not be mistaken for a meaningful one, especially when success depends heavily on academic esoterica rather than real-world applicability. While HLE offers a tougher challenge than outdated predecessors, its branding and difficulty level may overshadow its actual utility in measuring intelligence that matters for practical applications.

LIVE00:31DeepSeek's V4 Flash Agent Tasks Falter Amid Price Restructuring