llmboard.ai
Benchmarks
CompareRankings
llmboard.ai
Benchmarks
CompareRankings
HomeBenchmarksmathSciCode

math benchmark

SciCode

SciCode is a research coding benchmark curated by scientists that challenges language models to code solutions for scientific problems. It contains 338 subproblems decomposed from 80 challenging main problems across 16 natural science sub-fields including mathematics, physics, chemistry, biology, and materials science. Problems require knowledge recall, reasoning, and code synthesis skills.

Updated Aug 7, 2026

Published models18
Registry coverage18
MetricScore
EvidenceB

On this page

  • Ranking
  • Distribution
  • Highlights
  • About
  • FAQ

SciCode leaderboard

Sorted by the source-provided rank. Higher score is better according to the registry.

18 rows
Columns

Show columns

1BYSeed 2.1 ProByteDance59.8%100.0%18CAug 7, 2026
2GOGemini 3.1 ProGoogle59.0%94.1%18CAug 7, 2026
3BYSeed 2.1 TurboByteDance57.8%88.2%18CAug 7, 2026
4ACQwen3.7 MaxAlibaba Cloud / Qwen Team53.5%82.3%18CAug 7, 2026
5MAKimi K2.6Moonshot AI52.2%76.5%18CAug 7, 2026
6ACQwen3.7-PlusAlibaba Cloud / Qwen Team51.3%70.6%18CAug 7, 2026
7MAKimi K2.5Moonshot AI48.7%64.7%18CAug 7, 2026
8MAKimi K2-Thinking-0905Moonshot AI44.8%58.8%18CAug 7, 2026
9NVNemotron 3 Ultra (550B A55B)NVIDIA44.6%52.9%18CAug 7, 2026
10NVNemotron 3 Super (120B A12B)NVIDIA42.0%47.1%18CAug 7, 2026
11ZAGLM-4.5Zhipu AI41.7%41.2%18CAug 7, 2026
12MIMiniMax M2.1MiniMax39.0%35.3%18CAug 7, 2026
13CONorth Mini Code 1.0Cohere38.2%29.4%18CAug 7, 2026
14COCommand A+Cohere38.0%23.5%18CAug 7, 2026
15INMercury 2Inception38.0%17.6%18CAug 7, 2026
16ZAGLM-4.5-AirZhipu AI37.3%11.8%18CAug 7, 2026
17MIMiniMax M2MiniMax36.0%5.9%18CAug 7, 2026
18NVNemotron 3 Nano (30B A3B)NVIDIA33.3%0.0%18CAug 7, 2026

Score distribution

Top published rows on the benchmark's original scale.

SciCode

SciCode highlights

The top published results on this benchmark's own scale.

Rank #1Seed 2.1 Pro59.8%Rank #2Gemini 3.1 Pro59.0%Rank #3Seed 2.1 Turbo57.8%Rank #4Qwen3.7 Max53.5%

What is SciCode?

Definition and scoring fields from the benchmark registry.

SciCode is a research coding benchmark curated by scientists that challenges language models to code solutions for scientific problems. It contains 338 subproblems decomposed from 80 challenging main problems across 16 natural science sub-fields including mathematics, physics, chemistry, biology, and materials science. Problems require knowledge recall, reasoning, and code synthesis skills.

Scores are shown in ratio. The current registry marks this benchmark as not independently verified with evidence level B.

Family
SciCode
Modality
text
Primary category
math
Score direction
higher
LLMBoard eligible
Yes
Evaluation key
scicode|llm-stats-current

Source-native results are preserved. Eligibility for the overall LLMBoard score is a separate policy decision.

FAQ

Common questions about SciCode.

Which model scores highest on SciCode?

Seed 2.1 Pro is currently ranked first with 59.8%.

What does SciCode measure?

SciCode is a research coding benchmark curated by scientists that challenges language models to code solutions for scientific problems. It contains 338 subproblems decomposed from 80 challenging main problems across 16 natural science sub-fields including mathematics, physics, chemistry, biology, and materials science. Problems require knowledge recall, reasoning, and code synthesis skills.

Is a higher score better?

Yes. Higher values rank better for this benchmark.

How many models are compared?

18 unique published model results are currently shown.

Does this benchmark affect the overall score?

This benchmark is marked as eligible for the current LLMBoard capability methodology.

Rankings

OverallCodingText ArenaPricing

Modalities

Image GenerationVideo GenerationSpeech-to-TextEmbeddings

Benchmarks

All BenchmarksReasoningMathCoding

Vendors

All VendorsOpenAIAnthropicGoogle
llmboard.aiCopyright 2026 llmboard.ai