llmboard.ai
Benchmarks
CompareRankings
llmboard.ai
Benchmarks
CompareRankings
HomeBenchmarksreasoningHellaSwag

reasoning benchmark

HellaSwag

A challenging commonsense natural language inference dataset that uses Adversarial Filtering to create questions trivial for humans (>95% accuracy) but difficult for state-of-the-art models, requiring completion of sentence endings based on physical situations and everyday activities

Updated Aug 7, 2026

Published models27
Registry coverage27
MetricScore
EvidenceB

On this page

  • Ranking
  • Distribution
  • Highlights
  • About
  • FAQ

HellaSwag leaderboard

Sorted by the source-provided rank. Higher score is better according to the registry.

27 rows
Columns

Show columns

1ANClaude 3 OpusAnthropic95.4%100.0%27CAug 7, 2026
2OPGPT-4OpenAI95.3%96.2%27CAug 7, 2026
3GOGemini 1.5 ProGoogle93.3%92.3%27CAug 7, 2026
4XIMiMo-V2.5-ProXiaomi89.8%88.5%27CAug 7, 2026
5ANClaude 3 SonnetAnthropic89.0%84.6%27CAug 7, 2026
6COCommand R+Cohere88.6%80.8%27CAug 7, 2026
7NRHermes 3 70BNous Research88.2%76.9%27CAug 7, 2026
8ACQwen2 72B InstructAlibaba Cloud / Qwen Team87.6%73.1%27CAug 7, 2026
9GOGemini 1.5 FlashGoogle86.5%69.2%27CAug 7, 2026
10GOGemma 2 27BGoogle86.4%65.4%27CAug 7, 2026
11ANClaude 3 HaikuAnthropic85.9%61.5%27CAug 7, 2026
12NVLlama 3.1 Nemotron 70B InstructNVIDIA85.6%57.7%27CAug 7, 2026
13ACQwen2.5 32B InstructAlibaba Cloud / Qwen Team85.2%53.9%27CAug 7, 2026
14MIPhi-3.5-MoE-instructMicrosoft83.8%50.0%27CAug 7, 2026
15MAMistral NeMo InstructMistral AI83.5%46.1%27CAug 7, 2026
16ACQwen2.5-Coder 32B InstructAlibaba Cloud / Qwen Team83.0%42.3%27CAug 7, 2026
17GOGemma 2 9BGoogle81.9%38.5%27CAug 7, 2026
18IBGranite 3.3 8B BaseIBM80.1%34.6%27CAug 7, 2026
19GOGemma 3n E4BGoogle78.6%30.8%27CAug 7, 2026
20GOGemma 3n E4B Instructed LiteRT PreviewGoogle78.6%26.9%27CAug 7, 2026
21ACQwen2.5-Coder 7B InstructAlibaba Cloud / Qwen Team76.8%23.1%27CAug 7, 2026
22GOGemma 3n E2BGoogle72.2%19.2%27CAug 7, 2026
23GOGemma 3n E2B Instructed LiteRT (Preview)Google72.2%15.4%27CAug 7, 2026
24MELlama 3.2 3B InstructMeta69.8%11.5%27CAug 7, 2026
25MIPhi-3.5-mini-instructMicrosoft69.4%7.7%27CAug 7, 2026
26MIPhi 4 MiniMicrosoft69.1%3.9%27CAug 7, 2026
27BAERNIE 4.5Baidu33.0%0.0%27CAug 7, 2026

Score distribution

Top published rows on the benchmark's original scale.

HellaSwag

HellaSwag highlights

The top published results on this benchmark's own scale.

Rank #1Claude 3 Opus95.4%Rank #2GPT-495.3%Rank #3Gemini 1.5 Pro93.3%Rank #4MiMo-V2.5-Pro89.8%

What is HellaSwag?

Definition and scoring fields from the benchmark registry.

A challenging commonsense natural language inference dataset that uses Adversarial Filtering to create questions trivial for humans (>95% accuracy) but difficult for state-of-the-art models, requiring completion of sentence endings based on physical situations and everyday activities

Scores are shown in ratio. The current registry marks this benchmark as not independently verified with evidence level B.

Family
HellaSwag
Modality
text
Primary category
reasoning
Score direction
higher
LLMBoard eligible
Yes
Evaluation key
hellaswag|llm-stats-current

Source-native results are preserved. Eligibility for the overall LLMBoard score is a separate policy decision.

FAQ

Common questions about HellaSwag.

Which model scores highest on HellaSwag?

Claude 3 Opus is currently ranked first with 95.4%.

What does HellaSwag measure?

A challenging commonsense natural language inference dataset that uses Adversarial Filtering to create questions trivial for humans (>95% accuracy) but difficult for state-of-the-art models, requiring completion of sentence endings based on physical situations and everyday activities

Is a higher score better?

Yes. Higher values rank better for this benchmark.

How many models are compared?

27 unique published model results are currently shown.

Does this benchmark affect the overall score?

This benchmark is marked as eligible for the current LLMBoard capability methodology.

Rankings

OverallCodingText ArenaPricing

Modalities

Image GenerationVideo GenerationSpeech-to-TextEmbeddings

Benchmarks

All BenchmarksReasoningMathCoding

Vendors

All VendorsOpenAIAnthropicGoogle
llmboard.aiCopyright 2026 llmboard.ai