llmboard.ai
Benchmarks
CompareRankings
llmboard.ai
Benchmarks
CompareRankings
HomeBenchmarksreasoningTriviaQA

reasoning benchmark

TriviaQA

A large-scale reading comprehension dataset containing over 650K question-answer-evidence triples. TriviaQA includes 95K question-answer pairs authored by trivia enthusiasts and independently gathered evidence documents (six per question on average) that provide high quality distant supervision for answering the questions. The dataset features relatively complex, compositional questions with considerable syntactic and lexical variability, requiring cross-sentence reasoning to find answers.

Updated Aug 7, 2026

Published models18
Registry coverage18
MetricScore
EvidenceB

On this page

  • Ranking
  • Distribution
  • Highlights
  • About
  • FAQ

TriviaQA leaderboard

Sorted by the source-provided rank. Higher score is better according to the registry.

18 rows
Columns

Show columns

1MAKimi K2 BaseMoonshot AI85.1%100.0%18CAug 7, 2026
2GOGemma 2 27BGoogle83.7%94.1%18CAug 7, 2026
3XIMiMo-V2.5-ProXiaomi81.3%88.2%18CAug 7, 2026
4MAMistral Small 3.1 24B BaseMistral AI80.5%82.3%18CAug 7, 2026
5MAMistral Small 3.1 24B InstructMistral AI80.5%76.5%18CAug 7, 2026
6MAMistral Small 3 24B BaseMistral AI80.3%70.6%18CAug 7, 2026
7IBGranite 3.3 8B BaseIBM78.2%64.7%18CAug 7, 2026
8GOGemma 2 9BGoogle76.6%58.8%18CAug 7, 2026
9MAMinistral 3 (14B Base 2512)Mistral AI74.9%52.9%18CAug 7, 2026
10MAMistral Large 3Mistral AI74.9%47.1%18CAug 7, 2026
11MAMistral NeMo InstructMistral AI73.8%41.2%18CAug 7, 2026
12GOGemma 3n E4BGoogle70.2%35.3%18CAug 7, 2026
13GOGemma 3n E4B Instructed LiteRT PreviewGoogle70.2%29.4%18CAug 7, 2026
14MAMinistral 3 (8B Base 2512)Mistral AI68.1%23.5%18CAug 7, 2026
15MAMinistral 8B InstructMistral AI65.5%17.6%18CAug 7, 2026
16GOGemma 3n E2BGoogle60.8%11.8%18CAug 7, 2026
17GOGemma 3n E2B Instructed LiteRT (Preview)Google60.8%5.9%18CAug 7, 2026
18MAMinistral 3 (3B Base 2512)Mistral AI59.2%0.0%18CAug 7, 2026

Score distribution

Top published rows on the benchmark's original scale.

TriviaQA

TriviaQA highlights

The top published results on this benchmark's own scale.

Rank #1Kimi K2 Base85.1%Rank #2Gemma 2 27B83.7%Rank #3MiMo-V2.5-Pro81.3%Rank #4Mistral Small 3.1 24B Base80.5%

What is TriviaQA?

Definition and scoring fields from the benchmark registry.

A large-scale reading comprehension dataset containing over 650K question-answer-evidence triples. TriviaQA includes 95K question-answer pairs authored by trivia enthusiasts and independently gathered evidence documents (six per question on average) that provide high quality distant supervision for answering the questions. The dataset features relatively complex, compositional questions with considerable syntactic and lexical variability, requiring cross-sentence reasoning to find answers.

Scores are shown in ratio. The current registry marks this benchmark as not independently verified with evidence level B.

Family
TriviaQA
Modality
text
Primary category
reasoning
Score direction
higher
LLMBoard eligible
Yes
Evaluation key
triviaqa|llm-stats-current

Source-native results are preserved. Eligibility for the overall LLMBoard score is a separate policy decision.

FAQ

Common questions about TriviaQA.

Which model scores highest on TriviaQA?

Kimi K2 Base is currently ranked first with 85.1%.

What does TriviaQA measure?

A large-scale reading comprehension dataset containing over 650K question-answer-evidence triples. TriviaQA includes 95K question-answer pairs authored by trivia enthusiasts and independently gathered evidence documents (six per question on average) that provide high quality distant supervision for answering the questions. The dataset features relatively complex, compositional questions with considerable syntactic and lexical variability, requiring cross-sentence reasoning to find answers.

Is a higher score better?

Yes. Higher values rank better for this benchmark.

How many models are compared?

18 unique published model results are currently shown.

Does this benchmark affect the overall score?

This benchmark is marked as eligible for the current LLMBoard capability methodology.

Rankings

OverallCodingText ArenaPricing

Modalities

Image GenerationVideo GenerationSpeech-to-TextEmbeddings

Benchmarks

All BenchmarksReasoningMathCoding

Vendors

All VendorsOpenAIAnthropicGoogle
llmboard.aiCopyright 2026 llmboard.ai