llmboard.ai
Benchmarks
CompareRankings
llmboard.ai
Benchmarks
CompareRankings
HomeBenchmarksagentsProgram Bench

agents benchmark

Program Bench

Program Bench evaluates code-generation agents by asking them to recreate a program's behavior from only a compiled binary and documentation. It spans 200 tasks from small CLI tools to large systems such as FFmpeg and SQLite, with submissions judged against more than 248,000 fuzz-generated behavioral tests.

Updated Aug 11, 2026

Models5
Model coverage5
MetricScore
EvidenceB

On this page

  • Ranking
  • Distribution
  • Highlights
  • About
  • FAQ

Program Bench Ranking

Higher score ranks better on this benchmark.

5 rows
Columns

Show columns

1MAKimi K3Moonshot AI77.8%100.0%5CAug 11, 2026
2ZAGLM-5.2Zhipu AI63.7%75.0%5CAug 11, 2026
3MAKimi K2.7 CodeMoonshot AI53.6%50.0%5CAug 11, 2026
4BYSeed 2.1 ProByteDance50.3%25.0%5CAug 11, 2026
5BYSeed 2.1 TurboByteDance49.4%0.0%5CAug 11, 2026

Program Bench Score Distribution

A closer view of the leading scores on this benchmark.

Program Bench

Program Bench Highlights

The leading models and scores on this benchmark.

Rank #1Kimi K377.8%Rank #2GLM-5.263.7%Rank #3Kimi K2.7 Code53.6%Rank #4Seed 2.1 Pro50.3%

What is Program Bench?

What Program Bench measures and how its scores work.

Program Bench evaluates code-generation agents by asking them to recreate a program's behavior from only a compiled binary and documentation. It spans 200 tasks from small CLI tools to large systems such as FFmpeg and SQLite, with submissions judged against more than 248,000 fuzz-generated behavioral tests.

Scores are shown in ratio. This benchmark is not independently verified and has an evidence level of B.

Family
Program Bench
Modality
text
Primary category
agents
Score direction
higher
LLMBoard eligible
Yes
Evaluation key
program-bench|llm-stats-current

Benchmark scores retain their original unit. Overall score eligibility is shown separately.

FAQ

Common questions about Program Bench.

Which model scores highest on Program Bench?

Kimi K3 is currently ranked first with 77.8%.

What does Program Bench measure?

Program Bench evaluates code-generation agents by asking them to recreate a program's behavior from only a compiled binary and documentation. It spans 200 tasks from small CLI tools to large systems such as FFmpeg and SQLite, with submissions judged against more than 248,000 fuzz-generated behavioral tests.

Is a higher score better?

Yes. Higher values rank better for this benchmark.

How many models are compared?

5 model results are currently shown.

Does this benchmark affect the overall score?

Yes. This benchmark can contribute to the current LLMBoard capability score.

Rankings

OverallCodingText ArenaPricing

Modalities

All ModelsImage GenerationImage EditingVideo GenerationImage-to-VideoVideo EditingText-to-SpeechSpeech-to-TextEmbeddings

Benchmarks

All BenchmarksReasoningMathCoding

Vendors

All VendorsOpenAIAnthropicGoogle
llmboard.aiCopyright 2026 llmboard.ai