llmboard.ai
Benchmarks
CompareRankings
llmboard.ai
Benchmarks
CompareRankings
HomeBenchmarksreasoningARC-C

reasoning benchmark

ARC-C

The AI2 Reasoning Challenge (ARC) Challenge Set is a multiple-choice question-answering benchmark containing grade-school level science questions that require advanced reasoning capabilities. ARC-C specifically contains questions that were answered incorrectly by both retrieval-based and word co-occurrence algorithms, making it a particularly challenging subset designed to test commonsense reasoning abilities in AI systems.

Updated Aug 11, 2026

Models34
Model coverage34
MetricScore
EvidenceB

On this page

  • Ranking
  • Distribution
  • Highlights
  • About
  • FAQ

ARC-C Ranking

Higher score ranks better on this benchmark.

34 rows
Columns

Show columns

1XIMiMo-V2.5-ProXiaomi97.2%100.0%34CAug 11, 2026
2MELlama 3.1 405B InstructMeta96.9%97.0%34CAug 11, 2026
3ANClaude 3 OpusAnthropic96.4%93.9%34CAug 11, 2026
4MELlama 3.1 70B InstructMeta94.8%90.9%34CAug 11, 2026
5AMNova ProAmazon94.8%87.9%34CAug 11, 2026
6ANClaude 3 SonnetAnthropic93.2%84.8%34CAug 11, 2026
7ALJamba 1.5 LargeAI21 Labs93.0%81.8%34CAug 11, 2026
8AMNova LiteAmazon92.4%78.8%34CAug 11, 2026
9MAMistral Small 3 24B BaseMistral AI91.3%75.8%34CAug 11, 2026
10MIPhi-3.5-MoE-instructMicrosoft91.0%72.7%34CAug 11, 2026
11AMNova MicroAmazon90.2%69.7%34CAug 11, 2026
12ANClaude 3 HaikuAnthropic89.2%66.7%34CAug 11, 2026
13ALJamba 1.5 MiniAI21 Labs85.7%63.6%34CAug 11, 2026
14MIPhi-3.5-mini-instructMicrosoft84.6%60.6%34CAug 11, 2026
15MIPhi 4 MiniMicrosoft83.7%57.6%34CAug 11, 2026
16MELlama 3.1 8B InstructMeta83.4%54.5%34CAug 11, 2026
17MELlama 3.2 3B InstructMeta78.6%51.5%34CAug 11, 2026
18MAMinistral 8B InstructMistral AI71.9%48.5%34CAug 11, 2026
19GOGemma 2 27BGoogle71.4%45.5%34CAug 11, 2026
20COCommand R+Cohere71.0%42.4%34CAug 11, 2026
21ACQwen2.5-Coder 32B InstructAlibaba Cloud / Qwen Team70.5%39.4%34CAug 11, 2026
22ACQwen2.5 32B InstructAlibaba Cloud / Qwen Team70.4%36.4%34CAug 11, 2026
23NVLlama 3.1 Nemotron 70B InstructNVIDIA69.2%33.3%34CAug 11, 2026
24ACQwen2 72B InstructAlibaba Cloud / Qwen Team68.9%30.3%34CAug 11, 2026
25GOGemma 2 9BGoogle68.4%27.3%34CAug 11, 2026
26ACQwen2.5 14B InstructAlibaba Cloud / Qwen Team67.3%24.2%34CAug 11, 2026
27NRHermes 3 70BNous Research65.5%21.2%34CAug 11, 2026
28GOGemma 3n E4BGoogle61.6%18.2%34CAug 11, 2026
29GOGemma 3n E4B Instructed LiteRT PreviewGoogle61.6%15.2%34CAug 11, 2026
30ACQwen2.5-Coder 7B InstructAlibaba Cloud / Qwen Team60.9%12.1%34CAug 11, 2026
31GOGemma 3n E2BGoogle51.7%9.1%34CAug 11, 2026
32GOGemma 3n E2B Instructed LiteRT (Preview)Google51.7%6.1%34CAug 11, 2026
33IBGranite 3.3 8B BaseIBM50.8%3.0%34CAug 11, 2026
34BAERNIE 4.5Baidu40.6%0.0%34CAug 11, 2026

ARC-C Score Distribution

A closer view of the leading scores on this benchmark.

ARC-C

ARC-C Highlights

The leading models and scores on this benchmark.

Rank #1MiMo-V2.5-Pro97.2%Rank #2Llama 3.1 405B Instruct96.9%Rank #3Claude 3 Opus96.4%Rank #4Llama 3.1 70B Instruct94.8%

What is ARC-C?

What ARC-C measures and how its scores work.

The AI2 Reasoning Challenge (ARC) Challenge Set is a multiple-choice question-answering benchmark containing grade-school level science questions that require advanced reasoning capabilities. ARC-C specifically contains questions that were answered incorrectly by both retrieval-based and word co-occurrence algorithms, making it a particularly challenging subset designed to test commonsense reasoning abilities in AI systems.

Scores are shown in ratio. This benchmark is not independently verified and has an evidence level of B.

Family
ARC-C
Modality
text
Primary category
reasoning
Score direction
higher
LLMBoard eligible
Yes
Evaluation key
arc-c|llm-stats-current

Benchmark scores retain their original unit. Overall score eligibility is shown separately.

FAQ

Common questions about ARC-C.

Which model scores highest on ARC-C?

MiMo-V2.5-Pro is currently ranked first with 97.2%.

What does ARC-C measure?

The AI2 Reasoning Challenge (ARC) Challenge Set is a multiple-choice question-answering benchmark containing grade-school level science questions that require advanced reasoning capabilities. ARC-C specifically contains questions that were answered incorrectly by both retrieval-based and word co-occurrence algorithms, making it a particularly challenging subset designed to test commonsense reasoning abilities in AI systems.

Is a higher score better?

Yes. Higher values rank better for this benchmark.

How many models are compared?

34 model results are currently shown.

Does this benchmark affect the overall score?

Yes. This benchmark can contribute to the current LLMBoard capability score.

Rankings

OverallCodingText ArenaPricing

Modalities

All ModelsImage GenerationImage EditingVideo GenerationImage-to-VideoVideo EditingText-to-SpeechSpeech-to-TextEmbeddings

Benchmarks

All BenchmarksReasoningMathCoding

Vendors

All VendorsOpenAIAnthropicGoogle
llmboard.aiCopyright 2026 llmboard.ai