llmboard.ai
Benchmarks
CompareRankings
llmboard.ai
Benchmarks
CompareRankings
HomeBenchmarksreasoningARC-AGI

reasoning benchmark

ARC-AGI

The Abstraction and Reasoning Corpus for Artificial General Intelligence (ARC-AGI) is a benchmark designed to test general intelligence and abstract reasoning capabilities through visual grid-based transformation tasks. Each task consists of 2-5 demonstration pairs showing input grids transformed into output grids according to underlying rules, with test-takers required to infer these rules and apply them to novel test inputs. The benchmark uses colored grids (up to 30x30) with 10 discrete colors/symbols, designed to measure human-like general fluid intelligence and skill-acquisition efficiency with minimal prior knowledge.

Updated Aug 7, 2026

Published models7
Registry coverage7
MetricScore
EvidenceB

On this page

  • Ranking
  • Distribution
  • Highlights
  • About
  • FAQ

ARC-AGI leaderboard

Sorted by the source-provided rank. Higher score is better according to the registry.

7 rows
Columns

Show columns

1OPGPT-5.5OpenAI95.0%100.0%7CAug 7, 2026
2OPGPT-5.4OpenAI93.7%83.3%7CAug 7, 2026
3OPGPT-5.2 ProOpenAI90.5%66.7%7CAug 7, 2026
4OPo3OpenAI88.0%50.0%7CAug 7, 2026
5OPGPT-5.2OpenAI86.2%33.3%7CAug 7, 2026
6MELongCat-Flash-ThinkingMeituan50.3%16.7%7CAug 7, 2026
7ACQwen3-235B-A22B-Instruct-2507Alibaba Cloud / Qwen Team41.8%0.0%7CAug 7, 2026

Score distribution

Top published rows on the benchmark's original scale.

ARC-AGI

ARC-AGI highlights

The top published results on this benchmark's own scale.

Rank #1GPT-5.595.0%Rank #2GPT-5.493.7%Rank #3GPT-5.2 Pro90.5%Rank #4o388.0%

What is ARC-AGI?

Definition and scoring fields from the benchmark registry.

The Abstraction and Reasoning Corpus for Artificial General Intelligence (ARC-AGI) is a benchmark designed to test general intelligence and abstract reasoning capabilities through visual grid-based transformation tasks. Each task consists of 2-5 demonstration pairs showing input grids transformed into output grids according to underlying rules, with test-takers required to infer these rules and apply them to novel test inputs. The benchmark uses colored grids (up to 30x30) with 10 discrete colors/symbols, designed to measure human-like general fluid intelligence and skill-acquisition efficiency with minimal prior knowledge.

Scores are shown in ratio. The current registry marks this benchmark as not independently verified with evidence level B.

Family
ARC-AGI
Modality
image
Primary category
reasoning
Score direction
higher
LLMBoard eligible
No
Evaluation key
arc-agi|llm-stats-current

Source-native results are preserved. Eligibility for the overall LLMBoard score is a separate policy decision.

FAQ

Common questions about ARC-AGI.

Which model scores highest on ARC-AGI?

GPT-5.5 is currently ranked first with 95.0%.

What does ARC-AGI measure?

The Abstraction and Reasoning Corpus for Artificial General Intelligence (ARC-AGI) is a benchmark designed to test general intelligence and abstract reasoning capabilities through visual grid-based transformation tasks. Each task consists of 2-5 demonstration pairs showing input grids transformed into output grids according to underlying rules, with test-takers required to infer these rules and apply them to novel test inputs. The benchmark uses colored grids (up to 30x30) with 10 discrete colors/symbols, designed to measure human-like general fluid intelligence and skill-acquisition efficiency with minimal prior knowledge.

Is a higher score better?

Yes. Higher values rank better for this benchmark.

How many models are compared?

7 unique published model results are currently shown.

Does this benchmark affect the overall score?

This benchmark is preserved as source-native evidence but is not eligible for the current overall score.

Rankings

OverallCodingText ArenaPricing

Modalities

Image GenerationVideo GenerationSpeech-to-TextEmbeddings

Benchmarks

All BenchmarksReasoningMathCoding

Vendors

All VendorsOpenAIAnthropicGoogle
llmboard.aiCopyright 2026 llmboard.ai