llmboard.ai
Benchmarks
CompareRankings
llmboard.ai
Benchmarks
CompareRankings
HomeBenchmarksreasoningQasper

reasoning benchmark

Qasper

QASPER is a dataset of 5,049 information-seeking questions and answers anchored in 1,585 NLP research papers. Questions are written by NLP practitioners who read only titles and abstracts, while answers require understanding the full paper text and provide supporting evidence. The dataset challenges models with complex reasoning across document sections for academic document question answering. Each question seeks information present in the full text and is answered by a separate set of NLP practitioners who also provide supporting evidence to answers.

Updated Aug 7, 2026

Published models2
Registry coverage2
MetricScore
EvidenceB

On this page

  • Ranking
  • Distribution
  • Highlights
  • About
  • FAQ

Qasper leaderboard

Sorted by the source-provided rank. Higher score is better according to the registry.

2 rows
Columns

Show columns

1MIPhi-3.5-mini-instructMicrosoft41.9%100.0%2CAug 7, 2026
2MIPhi-3.5-MoE-instructMicrosoft40.0%0.0%2CAug 7, 2026

Score distribution

Top published rows on the benchmark's original scale.

Qasper

Qasper highlights

The top published results on this benchmark's own scale.

Rank #1Phi-3.5-mini-instruct41.9%Rank #2Phi-3.5-MoE-instruct40.0%

What is Qasper?

Definition and scoring fields from the benchmark registry.

QASPER is a dataset of 5,049 information-seeking questions and answers anchored in 1,585 NLP research papers. Questions are written by NLP practitioners who read only titles and abstracts, while answers require understanding the full paper text and provide supporting evidence. The dataset challenges models with complex reasoning across document sections for academic document question answering. Each question seeks information present in the full text and is answered by a separate set of NLP practitioners who also provide supporting evidence to answers.

Scores are shown in ratio. The current registry marks this benchmark as not independently verified with evidence level B.

Family
Qasper
Modality
text
Primary category
reasoning
Score direction
higher
LLMBoard eligible
No
Evaluation key
qasper|llm-stats-current

Source-native results are preserved. Eligibility for the overall LLMBoard score is a separate policy decision.

FAQ

Common questions about Qasper.

Which model scores highest on Qasper?

Phi-3.5-mini-instruct is currently ranked first with 41.9%.

What does Qasper measure?

QASPER is a dataset of 5,049 information-seeking questions and answers anchored in 1,585 NLP research papers. Questions are written by NLP practitioners who read only titles and abstracts, while answers require understanding the full paper text and provide supporting evidence. The dataset challenges models with complex reasoning across document sections for academic document question answering. Each question seeks information present in the full text and is answered by a separate set of NLP practitioners who also provide supporting evidence to answers.

Is a higher score better?

Yes. Higher values rank better for this benchmark.

How many models are compared?

2 unique published model results are currently shown.

Does this benchmark affect the overall score?

This benchmark is preserved as source-native evidence but is not eligible for the current overall score.

Rankings

OverallCodingText ArenaPricing

Modalities

Image GenerationVideo GenerationSpeech-to-TextEmbeddings

Benchmarks

All BenchmarksReasoningMathCoding

Vendors

All VendorsOpenAIAnthropicGoogle
llmboard.aiCopyright 2026 llmboard.ai