AILINCOM logo AILINCOM
A new exam has been created to test the boundaries of AI.
Cover generated by AI
Crius

Crius

Mar 14, 2026
Основная категория
Digital technologies and IT · Artificial Intelligence
Дополнительные
Research and development · Artificial Intelligence

A new exam has been created to test the boundaries of AI.

A new exam has been created to test the boundaries of AI.

An international team of experts has developed a new exam, Humanity's Last Exam, designed to objectively assess the capabilities of artificial intelligence. The test includes tasks that current AI models are still unable to solve. Initial results show that even the most advanced AI systems are far from reaching human-level performance, and human expertise remains irreplaceable.

CriusA new exam has been created to test the boundaries of AI.

As artificial intelligence began to achieve high scores on traditional academic tests, a new challenge emerged: previous assessments, once considered difficult for machines, no longer provide an objective measure of the capabilities of modern AI models. For example, the Massive Multitask Language Understanding (MMLU) exam, previously seen as a tough benchmark, no longer reflects the true advancement of leading AI systems.

Developing a New Test for AI

To address this issue, an international team of nearly a thousand experts, including representatives from Texas A&M University, created a new type of exam—Humanity's Last Exam (HLE). This test consists of 2,500 questions covering mathematics, the humanities, natural sciences, ancient languages, and specialized academic disciplines. The exam was specifically designed to include tasks that current AI systems still struggle to solve. Detailed information about the project has been published in the journal Nature and on the website lastexam.ai.

Features of Humanity's Last Exam

Experts from various fields participated in the creation and refinement of the questions. Each question was carefully crafted to have a clear and verifiable answer and to prevent it from being easily solved through a simple internet search. The topics include translating ancient inscriptions, identifying anatomical structures in birds, and analyzing features of Biblical Hebrew pronunciation.

The Question Selection Process

Every question was tested on leading AI systems. If any model could answer a question correctly, it was removed from the final exam. This approach ensured the test remains more challenging than what current AI can reliably solve.

Test Results

Initial trials showed that even the most advanced AI models face significant difficulties with this exam. For instance, GPT-4o scored 2.7%, Claude 3.5 Sonnet achieved 4.1%, and OpenAI o1 reached 8%. The most sophisticated systems, such as Gemini 3.1 Pro and Claude Opus 4.6, managed accuracy rates between 40% and 50%.

The Need for New Tests

High scores on tests originally designed for humans do not always indicate true AI intelligence. Such tests mainly measure the ability to perform specific tasks, not the depth of understanding. Without precise evaluation tools, developers and users may misinterpret the capabilities of AI systems. Benchmarks are essential for objectively measuring progress and identifying potential risks.

Long-Term Goals and Test Structure

Humanity's Last Exam is intended as a robust and transparent tool for evaluating future AI systems. To prevent models from memorizing answers, only a portion of the questions are published, while most remain hidden.

International Collaboration

The project brought together experts from different countries and disciplines, including specialists in computer science, history, physics, linguistics, and medicine. This diversity helped identify gaps in current AI systems and highlighted the importance of human expertise in creating complex challenges.

The Significance of the Test

Humanity's Last Exam is not designed to replace humans or pose a threat, but rather to serve as an objective tool for assessing the boundaries of artificial intelligence. The exam underscores that, despite technological progress, there remains a significant gap between AI capabilities and human intelligence, and human expertise continues to play a crucial role.

#artificial_intelligence#testing#nature#benchmarks#evaluation#Humanitys_Last_Exam
0 —

Comments (0)

Hot

Qnap has announced new NAS devices for video production

Oct 2, 202610/2/26 · 0 reactions

Tesla opened credit lines worth $30 billion

Oct 2, 202610/2/26 · 0 reactions

Air travel is on the rise, but new regulations are making the market more complicated.

Oct 1, 202610/1/26 · 0 reactions
Recommended
Cloud Computing

Qnap has announced new NAS devices for video production

Qnap has introduced three new NAS systems designed for video production tasks, equipped with USB4 ports for high-speed data transfer. These devices support various connection modes and are intended for use with high-capacity hard drives and SSDs.

Financial Analysis

Tesla opened credit lines worth $30 billion

Tesla has opened credit lines totaling $30 billion to finance major investments amid declining profits and rising capital expenditures. The new agreement expands the company's financial flexibility as it faces increasing pressure on its business.

Transportation Logistics

Air travel is on the rise, but new regulations are making the market more complicated.

Air transportation is becoming an increasingly important part of logistics, especially amid the instability of sea shipping. However, new regulations for preparing air waybills are creating additional challenges and risks for market participants.