llmboard.ai
Benchmarks
CompareRankings
llmboard.ai
Benchmarks
CompareRankings
HomeBenchmarksmathBIG-Bench

math benchmark

BIG-Bench

Beyond the Imitation Game Benchmark (BIG-bench) is a collaborative benchmark consisting of 204+ tasks designed to probe large language models and extrapolate their future capabilities. It covers diverse domains including linguistics, mathematics, common-sense reasoning, biology, physics, social bias, software development, and more. The benchmark focuses on tasks believed to be beyond current language model capabilities and includes both English and non-English tasks across multiple languages.

Updated Aug 7, 2026

Published models3
Registry coverage3
MetricScore
EvidenceB

On this page

  • Ranking
  • Distribution
  • Highlights
  • About
  • FAQ

BIG-Bench leaderboard

Sorted by the source-provided rank. Higher score is better according to the registry.

3 rows
Columns

Show columns

1GOGemini 1.0 ProGoogle75.0%100.0%3BAug 7, 2026
2GOGemma 2 27BGoogle74.9%50.0%3CAug 7, 2026
3GOGemma 2 9BGoogle68.2%0.0%3CAug 7, 2026

Score distribution

Top published rows on the benchmark's original scale.

BIG-Bench

BIG-Bench highlights

The top published results on this benchmark's own scale.

Rank #1Gemini 1.0 Pro75.0%Rank #2Gemma 2 27B74.9%Rank #3Gemma 2 9B68.2%

What is BIG-Bench?

Definition and scoring fields from the benchmark registry.

Beyond the Imitation Game Benchmark (BIG-bench) is a collaborative benchmark consisting of 204+ tasks designed to probe large language models and extrapolate their future capabilities. It covers diverse domains including linguistics, mathematics, common-sense reasoning, biology, physics, social bias, software development, and more. The benchmark focuses on tasks believed to be beyond current language model capabilities and includes both English and non-English tasks across multiple languages.

Scores are shown in ratio. The current registry marks this benchmark as not independently verified with evidence level B.

Family
BIG-Bench
Modality
text
Primary category
math
Score direction
higher
LLMBoard eligible
No
Evaluation key
big-bench|llm-stats-current

Source-native results are preserved. Eligibility for the overall LLMBoard score is a separate policy decision.

FAQ

Common questions about BIG-Bench.

Which model scores highest on BIG-Bench?

Gemini 1.0 Pro is currently ranked first with 75.0%.

What does BIG-Bench measure?

Beyond the Imitation Game Benchmark (BIG-bench) is a collaborative benchmark consisting of 204+ tasks designed to probe large language models and extrapolate their future capabilities. It covers diverse domains including linguistics, mathematics, common-sense reasoning, biology, physics, social bias, software development, and more. The benchmark focuses on tasks believed to be beyond current language model capabilities and includes both English and non-English tasks across multiple languages.

Is a higher score better?

Yes. Higher values rank better for this benchmark.

How many models are compared?

3 unique published model results are currently shown.

Does this benchmark affect the overall score?

This benchmark is preserved as source-native evidence but is not eligible for the current overall score.

Rankings

OverallCodingText ArenaPricing

Modalities

Image GenerationVideo GenerationSpeech-to-TextEmbeddings

Benchmarks

All BenchmarksReasoningMathCoding

Vendors

All VendorsOpenAIAnthropicGoogle
llmboard.aiCopyright 2026 llmboard.ai