llmboard.ai
Benchmarks
CompareRankings
llmboard.ai
Benchmarks
CompareRankings
HomeBenchmarksmultimodalCharXiv-R

multimodal benchmark

CharXiv-R

CharXiv-R is the reasoning component of the CharXiv benchmark, focusing on complex reasoning questions that require synthesizing information across visual chart elements. It evaluates multimodal large language models on their ability to understand and reason about scientific charts from arXiv papers through various reasoning tasks.

Updated Aug 7, 2026

Published models47
Registry coverage47
MetricScore
EvidenceB

On this page

  • Ranking
  • Distribution
  • Highlights
  • About
  • FAQ

CharXiv-R leaderboard

Sorted by the source-provided rank. Higher score is better according to the registry.

47 rows
Columns

Show columns

1ANClaude Mythos PreviewAnthropic93.2%100.0%47CAug 7, 2026
2MAKimi K3Moonshot AI91.3%97.8%47CAug 7, 2026
3ANClaude Opus 4.7Anthropic91.0%95.7%47CAug 7, 2026
4ANClaude Opus 4.8Anthropic89.9%93.5%47CAug 7, 2026
5GOGemini 3.6 FlashGoogle89.4%91.3%47CAug 7, 2026
6MEMuse Spark 1.1Meta88.4%89.1%47CAug 7, 2026
7ANClaude Sonnet 5Anthropic88.3%87.0%47CAug 7, 2026
8MAKimi K2.6Moonshot AI86.7%84.8%47CAug 7, 2026
9MEMuse SparkMeta86.4%82.6%47CAug 7, 2026
10BYSeed 2.1 ProByteDance86.4%80.4%47CAug 7, 2026
11ACQwen3.7-PlusAlibaba Cloud / Qwen Team85.9%78.3%47CAug 7, 2026
12GOGemini 3.5 FlashGoogle84.2%76.1%47CAug 7, 2026
13BYSeed 2.1 TurboByteDance83.6%73.9%47CAug 7, 2026
14OPGPT-5.2OpenAI82.1%71.7%47CAug 7, 2026
15OPGPT-5.5 InstantOpenAI81.6%69.6%47CAug 7, 2026
16ACQwen3.6 PlusAlibaba Cloud / Qwen Team81.5%67.4%47CAug 7, 2026
17GOGemini 3 ProGoogle81.4%65.2%47CAug 7, 2026
18OPGPT-5OpenAI81.1%63.0%47CAug 7, 2026
19XIMiMo-V2.5Xiaomi81.0%60.9%47CAug 7, 2026
20GOGemini 3 FlashGoogle80.3%58.7%47CAug 7, 2026
21ACQwen3.5-27BAlibaba Cloud / Qwen Team79.5%56.5%47CAug 7, 2026
22OPo3OpenAI78.6%54.4%47CAug 7, 2026
23ACQwen3.6-27BAlibaba Cloud / Qwen Team78.4%52.2%47CAug 7, 2026
24ACQwen3.6-35B-A3BAlibaba Cloud / Qwen Team78.0%50.0%47CAug 7, 2026
25MAKimi K2.5Moonshot AI77.5%47.8%47CAug 7, 2026
26ACQwen3.5-35B-A3BAlibaba Cloud / Qwen Team77.5%45.6%47CAug 7, 2026
27ANClaude Opus 4.6Anthropic77.4%43.5%47CAug 7, 2026
28ACQwen3.5-122B-A10BAlibaba Cloud / Qwen Team77.2%41.3%47CAug 7, 2026
29GOGemini 3.5 Flash-LiteGoogle76.5%39.1%47CAug 7, 2026
30GOGemini 3.1 Flash-LiteGoogle73.2%37.0%47CAug 7, 2026
31OPo4-miniOpenAI72.0%34.8%47CAug 7, 2026
32ACQwen3 VL 235B A22B ThinkingAlibaba Cloud / Qwen Team66.1%32.6%47CAug 7, 2026
33ACQwen3 VL 32B ThinkingAlibaba Cloud / Qwen Team65.2%30.4%47CAug 7, 2026
34ACQwen3 VL 32B InstructAlibaba Cloud / Qwen Team62.8%28.3%47CAug 7, 2026
35ACQwen3 VL 235B A22B InstructAlibaba Cloud / Qwen Team62.1%26.1%47CAug 7, 2026
36OPGPT-4oOpenAI58.8%23.9%47CAug 7, 2026
37OPGPT-4.1 miniOpenAI56.8%21.7%47CAug 7, 2026
38OPGPT-4.1OpenAI56.7%19.6%47CAug 7, 2026
39ACQwen3 VL 30B A3B ThinkingAlibaba Cloud / Qwen Team56.6%17.4%47CAug 7, 2026
40OPGPT-4.5OpenAI55.4%15.2%47CAug 7, 2026
41ACQwen3 VL 8B ThinkingAlibaba Cloud / Qwen Team53.0%13.0%47CAug 7, 2026
42COCommand A+Cohere52.7%10.9%47CAug 7, 2026
43ACQwen3 VL 4B ThinkingAlibaba Cloud / Qwen Team50.3%8.7%47CAug 7, 2026
44ACQwen3 VL 30B A3B InstructAlibaba Cloud / Qwen Team48.9%6.5%47CAug 7, 2026
45ACQwen3 VL 8B InstructAlibaba Cloud / Qwen Team46.4%4.3%47CAug 7, 2026
46OPGPT-4.1 nanoOpenAI40.5%2.2%47CAug 7, 2026
47ACQwen3 VL 4B InstructAlibaba Cloud / Qwen Team39.7%0.0%47CAug 7, 2026

Score distribution

Top published rows on the benchmark's original scale.

CharXiv-R

CharXiv-R highlights

The top published results on this benchmark's own scale.

Rank #1Claude Mythos Preview93.2%Rank #2Kimi K391.3%Rank #3Claude Opus 4.791.0%Rank #4Claude Opus 4.889.9%

What is CharXiv-R?

Definition and scoring fields from the benchmark registry.

CharXiv-R is the reasoning component of the CharXiv benchmark, focusing on complex reasoning questions that require synthesizing information across visual chart elements. It evaluates multimodal large language models on their ability to understand and reason about scientific charts from arXiv papers through various reasoning tasks.

Scores are shown in ratio. The current registry marks this benchmark as not independently verified with evidence level B.

Family
CharXiv-R
Modality
multimodal
Primary category
multimodal
Score direction
higher
LLMBoard eligible
No
Evaluation key
charxiv-r|llm-stats-current

Source-native results are preserved. Eligibility for the overall LLMBoard score is a separate policy decision.

FAQ

Common questions about CharXiv-R.

Which model scores highest on CharXiv-R?

Claude Mythos Preview is currently ranked first with 93.2%.

What does CharXiv-R measure?

CharXiv-R is the reasoning component of the CharXiv benchmark, focusing on complex reasoning questions that require synthesizing information across visual chart elements. It evaluates multimodal large language models on their ability to understand and reason about scientific charts from arXiv papers through various reasoning tasks.

Is a higher score better?

Yes. Higher values rank better for this benchmark.

How many models are compared?

47 unique published model results are currently shown.

Does this benchmark affect the overall score?

This benchmark is preserved as source-native evidence but is not eligible for the current overall score.

Rankings

OverallCodingText ArenaPricing

Modalities

Image GenerationVideo GenerationSpeech-to-TextEmbeddings

Benchmarks

All BenchmarksReasoningMathCoding

Vendors

All VendorsOpenAIAnthropicGoogle
llmboard.aiCopyright 2026 llmboard.ai