llmboard.ai
Benchmarks
CompareRankings
llmboard.ai
Benchmarks
CompareRankings
HomeBenchmarksmathHumanity's Last Exam

math benchmark

Humanity's Last Exam

Humanity's Last Exam (HLE) is a multi-modal academic benchmark with 2,500 questions across mathematics, humanities, and natural sciences, designed to test LLM capabilities at the frontier of human knowledge with unambiguous, verifiable solutions

Updated Aug 7, 2026

Published models92
Registry coverage92
MetricScore
EvidenceB

On this page

  • Ranking
  • Distribution
  • Highlights
  • About
  • FAQ

Humanity's Last Exam leaderboard

Sorted by the source-provided rank. Higher score is better according to the registry.

92 rows
Columns

Show columns

1ANClaude Mythos PreviewAnthropic64.7%100.0%92CAug 7, 2026
2ANClaude Opus 5Anthropic64.7%98.9%92CAug 7, 2026
3ANClaude Fable 5Anthropic64.5%97.8%92CAug 7, 2026
4MEMuse Spark 1.1Meta62.1%96.7%92CAug 7, 2026
5MEMuse SparkMeta58.4%95.6%92CAug 7, 2026
6ANClaude Opus 4.8Anthropic57.9%94.5%92CAug 7, 2026
7ANClaude Sonnet 5Anthropic57.4%93.4%92CAug 7, 2026
8OPGPT-5.5 ProOpenAI57.2%92.3%92CAug 7, 2026
9MAKimi K3Moonshot AI56.0%91.2%92CAug 7, 2026
10BYSeed 2.1 ProByteDance55.7%90.1%92CAug 7, 2026
11ANClaude Opus 4.7Anthropic54.7%89.0%92CAug 7, 2026
12ZAGLM-5.2Zhipu AI54.7%87.9%92CAug 7, 2026
13BYSeed 2.1 TurboByteDance54.6%86.8%92CAug 7, 2026
14ANClaude Opus 4.6Anthropic53.1%85.7%92CAug 7, 2026
15ZAGLM-5.1Zhipu AI52.3%84.6%92CAug 7, 2026
16OPGPT-5.5OpenAI52.2%83.5%92CAug 7, 2026
17GOGemini 3.1 ProGoogle51.4%82.4%92CAug 7, 2026
18MAKimi K2-Thinking-0905Moonshot AI51.0%81.3%92CAug 7, 2026
19XAGrok-4 HeavyxAI50.7%80.2%92CAug 7, 2026
20MAKimi K2.5Moonshot AI50.2%79.1%92CAug 7, 2026
21ANClaude Sonnet 4.6Anthropic49.0%78.0%92CAug 7, 2026
22ACQwen3.5-27BAlibaba Cloud / Qwen Team48.5%76.9%92CAug 7, 2026
23DEDeepSeek-V4-Pro-MaxDeepSeek48.2%75.8%92CAug 7, 2026
24ACQwen3.5-122B-A10BAlibaba Cloud / Qwen Team47.5%74.7%92CAug 7, 2026
25ACQwen3.5-35B-A3BAlibaba Cloud / Qwen Team47.4%73.6%92CAug 7, 2026
26GOGemini 3 ProGoogle45.8%72.5%92CAug 7, 2026
27DEDeepSeek-V4-Flash-MaxDeepSeek45.1%71.4%92CAug 7, 2026
28ACQwen3.8 MaxAlibaba Cloud / Qwen Team43.6%70.3%92CAug 7, 2026
29GOGemini 3 FlashGoogle43.5%69.2%92CAug 7, 2026
30ZAGLM-4.7Zhipu AI42.8%68.1%92CAug 7, 2026
31ACQwen3.7 MaxAlibaba Cloud / Qwen Team41.4%67.0%92CAug 7, 2026
32DEDeepSeek-V3.2DeepSeek40.8%65.9%92CAug 7, 2026
33GOGemini 3.5 FlashGoogle40.2%64.8%92CAug 7, 2026
34XAGrok-4xAI40.0%63.7%92CAug 7, 2026
35OPGPT-5.4OpenAI39.8%62.6%92CAug 7, 2026
36BAERNIE 5.0Baidu39.0%61.5%92CAug 7, 2026
37NVNemotron 3 Ultra (550B A55B)NVIDIA37.4%60.4%92CAug 7, 2026
38OPGPT-5.2 ProOpenAI36.6%59.3%92CAug 7, 2026
39MAKimi K2.6Moonshot AI36.4%58.2%92CAug 7, 2026
40ACQwen3.7-PlusAlibaba Cloud / Qwen Team34.7%57.1%92CAug 7, 2026
41OPGPT-5.2OpenAI34.5%56.0%92CAug 7, 2026
42XIMiMo-V2.5-ProXiaomi34.0%55.0%92CAug 7, 2026
43DEDeepSeek-V3.2-SpecialeDeepSeek30.6%53.9%92CAug 7, 2026
44ACQwen3.6 PlusAlibaba Cloud / Qwen Team28.8%52.8%92CAug 7, 2026
45ACQwen3.5-397B-A17BAlibaba Cloud / Qwen Team28.7%51.6%92CAug 7, 2026
46OPGPT-5.4 miniOpenAI28.2%50.5%92CAug 7, 2026
47GOGemma 4 31BGoogle26.5%49.5%92CAug 7, 2026
48MELongCat-Flash-Thinking-2601Meituan25.2%48.4%92CAug 7, 2026
49DEDeepSeek-V3.2 (Thinking)DeepSeek25.1%47.3%92CAug 7, 2026
50OPGPT-5OpenAI24.8%46.1%92CAug 7, 2026
51OPGPT-5.4 nanoOpenAI24.3%45.0%92CAug 7, 2026
52ACQwen3.6-27BAlibaba Cloud / Qwen Team24.0%44.0%92CAug 7, 2026
53NVNemotron 3 Super (120B A12B)NVIDIA22.8%42.9%92CAug 7, 2026
54XIMiMo-V2-FlashXiaomi22.1%41.8%92CAug 7, 2026
55MIMiniMax M2.1MiniMax22.0%40.7%92CAug 7, 2026
56GOGemini 2.5 Pro Preview 06-05Google21.6%39.6%92CAug 7, 2026
57ACQwen3.6-35B-A3BAlibaba Cloud / Qwen Team21.4%38.5%92CAug 7, 2026
58XAGrok 4 FastxAI20.0%37.4%92CAug 7, 2026
59DEDeepSeek-V3.2-ExpDeepSeek19.8%36.3%92CAug 7, 2026
60ACQwen3-235B-A22B-Thinking-2507Alibaba Cloud / Qwen Team18.2%35.2%92CAug 7, 2026
61MIMAI-Code-1-FlashMicrosoft18.0%34.1%92CAug 7, 2026
62GOGemini 2.5 ProGoogle17.8%33.0%92CAug 7, 2026
63DEDeepSeek-R1-0528DeepSeek17.7%31.9%92CAug 7, 2026
64GOGemma 4 26B-A4BGoogle17.2%30.8%92CAug 7, 2026
65ZAGLM-4.6Zhipu AI17.2%29.7%92CAug 7, 2026
66OPGPT-5 miniOpenAI16.7%28.6%92CAug 7, 2026
67GOGemini 3.1 Flash-LiteGoogle16.0%27.5%92CAug 7, 2026
68DEDeepSeek-V3.1DeepSeek15.9%26.4%92CAug 7, 2026
69NVNemotron 3 Nano (30B A3B)NVIDIA15.5%25.3%92CAug 7, 2026
70OPGPT OSS 120BOpenAI14.9%24.2%92CAug 7, 2026
71OPo3OpenAI14.7%23.1%92CAug 7, 2026
72OPo4-miniOpenAI14.7%22.0%92CAug 7, 2026
73ZAGLM-4.5Zhipu AI14.4%20.9%92CAug 7, 2026
74ZAGLM-4.7-FlashZhipu AI14.4%19.8%92CAug 7, 2026
75ACQwen3 VL 235B A22B ThinkingAlibaba Cloud / Qwen Team13.6%18.7%92CAug 7, 2026
76MIMiniMax M2MiniMax12.5%17.6%92CAug 7, 2026
77GODiffusionGemma 26B-A4BGoogle11.9%16.5%92CAug 7, 2026
78SASarvam-105BSarvam AI11.2%15.4%92CAug 7, 2026
79GOGemini 2.5 FlashGoogle11.0%14.3%92CAug 7, 2026
80OPGPT OSS 20BOpenAI10.9%13.2%92CAug 7, 2026
81ZAGLM-4.5-AirZhipu AI10.6%12.1%92CAug 7, 2026
82MAMagistral MediumMistral AI9.0%11.0%92CAug 7, 2026
83OPGPT-5 nanoOpenAI8.7%9.9%92CAug 7, 2026
84MIMiniMax M1 80KMiniMax8.4%8.8%92CAug 7, 2026
85MIMiniMax M1 40KMiniMax7.2%7.7%92CAug 7, 2026
86OPGPT-4.1OpenAI5.4%6.6%92CAug 7, 2026
87OPGPT-4oOpenAI5.3%5.5%92CAug 7, 2026
88GOGemma 4 12BGoogle5.2%4.4%92CAug 7, 2026
89GOGemini 2.5 Flash-LiteGoogle5.1%3.3%92CAug 7, 2026
90MAKimi K2 InstructMoonshot AI4.7%2.2%92CAug 7, 2026
91MAKimi K2-Instruct-0905Moonshot AI4.7%1.1%92CAug 7, 2026
92OPGPT-4.1 miniOpenAI3.7%0.0%92CAug 7, 2026

Score distribution

Top published rows on the benchmark's original scale.

Humanity's Last Exam

Humanity's Last Exam highlights

The top published results on this benchmark's own scale.

Rank #1Claude Mythos Preview64.7%Rank #2Claude Opus 564.7%Rank #3Claude Fable 564.5%Rank #4Muse Spark 1.162.1%

What is Humanity's Last Exam?

Definition and scoring fields from the benchmark registry.

Humanity's Last Exam (HLE) is a multi-modal academic benchmark with 2,500 questions across mathematics, humanities, and natural sciences, designed to test LLM capabilities at the frontier of human knowledge with unambiguous, verifiable solutions

Scores are shown in ratio. The current registry marks this benchmark as not independently verified with evidence level B.

Family
Humanity's Last Exam
Modality
multimodal
Primary category
math
Score direction
higher
LLMBoard eligible
No
Evaluation key
humanity's-last-exam|llm-stats-current

Source-native results are preserved. Eligibility for the overall LLMBoard score is a separate policy decision.

FAQ

Common questions about Humanity's Last Exam.

Which model scores highest on Humanity's Last Exam?

Claude Mythos Preview is currently ranked first with 64.7%.

What does Humanity's Last Exam measure?

Humanity's Last Exam (HLE) is a multi-modal academic benchmark with 2,500 questions across mathematics, humanities, and natural sciences, designed to test LLM capabilities at the frontier of human knowledge with unambiguous, verifiable solutions

Is a higher score better?

Yes. Higher values rank better for this benchmark.

How many models are compared?

92 unique published model results are currently shown.

Does this benchmark affect the overall score?

This benchmark is preserved as source-native evidence but is not eligible for the current overall score.

Rankings

OverallCodingText ArenaPricing

Modalities

Image GenerationVideo GenerationSpeech-to-TextEmbeddings

Benchmarks

All BenchmarksReasoningMathCoding

Vendors

All VendorsOpenAIAnthropicGoogle
llmboard.aiCopyright 2026 llmboard.ai