llmboard.ai
Benchmarks
CompareRankings
llmboard.ai
Benchmarks
CompareRankings
HomeModelsKimi K2 Thinking

Moonshot AI model product

Kimi K2 Thinking

Kimi K2 Thinking is the latest, most capable version of open-source thinking model. Starting with Kimi K2, it is built as a thinking agent that reasons step-by-step while dynamically invoking tools. It sets a new state-of-the-art on Humanity's Last Exam (HLE), BrowseComp, and other benchmarks by dramatically scaling multi-step reasoning depth and maintaining stable tool-use across 200–300 sequential calls. At the same time, K2 Thinking is a native INT4 quantization model with 256k context window, achieving lossless reductions in inference latency and GPU memory usage. Key features include deep thinking & tool orchestration with end-to-end training to interleave chain-of-thought reasoning with function calls, native INT4 quantization via Quantization-Aware Training (QAT) achieving lossless 2x speed-up, and stable long-horizon agency maintaining coherent goal-directed behavior across up to 200–300 consecutive tool invocations.

Updated Aug 10, 2026. Default version: Kimi K2-Thinking-0905

Compare
LLMBoard score59.4Kimi K2-Thinking-0905
Coverage100%10 benchmark families
Context window262.1KTokens
Official input priceN/AOfficial price unavailable

On this page

  • Specification
  • Capability
  • Benchmarks
  • Arena
  • Pricing
  • Versions
  • About
  • Compare
  • Similar models
  • FAQ

Model specification

Structured fields from the published default version.

Version
Kimi K2-Thinking-0905
Released
Sep 5, 2025
Knowledge cutoff
Unknown
Parameters
1T
Context window
262.1K
Max output
262.1K
Inputs
text
Outputs
text
Open weights
No
License
MIT

Capability profile

This profile uses the latest version under this unique model that has a calculated LLMBoard score. Arena and price are excluded.

Kimi K2-Thinking-0905 category scores

Benchmark results

Published benchmark records for the scored version Kimi K2-Thinking-0905.

21 rows
Columns

Show columns

FinSearchComp-T30.511100.0%CAug 7, 2026
FRAMES0.912100.0%CAug 7, 2026
HealthBench0.62987.5%CAug 7, 2026
OJBench0.52987.5%CAug 7, 2026
Seal-00.62680.0%CAug 7, 2026
Terminal-Bench0.532591.7%CAug 7, 2026
HMMT 20251.043390.6%CAug 7, 2026
Multi-SWE-Bench0.44640.0%CAug 7, 2026
AIME 20251.0511496.5%CAug 7, 2026
MMLU-Redux0.954891.5%CAug 7, 2026
BrowseComp-zh0.681341.7%CAug 7, 2026
SciCode0.481858.8%CAug 7, 2026
LiveCodeBench v60.8135376.9%CAug 7, 2026
WritingBench0.715150.0%CAug 7, 2026
IMO-AnswerBench0.8171911.1%CAug 7, 2026
Humanity's Last Exam0.5189281.3%CAug 7, 2026
MMLU-Pro0.82412982.0%CAug 7, 2026
SWE-bench Multilingual0.6253427.3%CAug 7, 2026
BrowseComp0.6365838.6%CAug 7, 2026
SWE-Bench Verified0.75410448.5%CAug 7, 2026
GPQA0.85823375.4%CAug 7, 2026

Arena results

Preference and agent-evaluation signals from published Arena datasets.

No published Arena match

The default version has no published Arena rows, or its source alias has not been resolved.

Pricing

Official vendor API PAYG pricing is summarized first. The table then lists individual provider offerings without treating their minimum as the official price.

Official API
N/A
Official provider
N/A
Lowest third-party
N/A
Tracked offerings
0
No published price snapshot

The default version has no provider offering with current input or output token prices.

Official prices use only the vendor's configured official Provider and positive standard USD PAYG rates. Third-party offers remain explicitly labeled.

Kimi K2 Thinking versions

All published versions linked to this unique model. The score columns identify the version used by the current overall ranking.

1 rows
Columns

Show columns

Kimi K2-Thinking-0905Sep 5, 202559.41T262.1K262.1KNoMIT

What is Kimi K2 Thinking?

A concise description based on the published model registry.

Kimi K2 Thinking is the latest, most capable version of open-source thinking model. Starting with Kimi K2, it is built as a thinking agent that reasons step-by-step while dynamically invoking tools. It sets a new state-of-the-art on Humanity's Last Exam (HLE), BrowseComp, and other benchmarks by dramatically scaling multi-step reasoning depth and maintaining stable tool-use across 200–300 sequential calls. At the same time, K2 Thinking is a native INT4 quantization model with 256k context window, achieving lossless reductions in inference latency and GPU memory usage. Key features include deep thinking & tool orchestration with end-to-end training to interleave chain-of-thought reasoning with function calls, native INT4 quantization via Quantization-Aware Training (QAT) achieving lossless 2x speed-up, and stable long-horizon agency maintaining coherent goal-directed behavior across up to 200–300 consecutive tool invocations.

Use the benchmark, Arena and pricing sections above as separate evidence. A missing field means the current data snapshot does not support that claim.

Data snapshot: 2026-08-07. Editorial model content is not available in the backend.

Kimi K2 Thinking vs nearby models

Open a comparison with the three ranked models immediately above and below this model.

Kimi K2 ThinkingvsQwen3.5 122B A10BKimi K2 ThinkingvsLongCat Flash ThinkingKimi K2 ThinkingvsNemotron Nano 9BKimi K2 ThinkingvsNemotron 3 SuperKimi K2 ThinkingvsClaude Opus 3Kimi K2 ThinkingvsGPT-4.5

Models similar to Kimi K2 Thinking

Recommendations prioritize the same model type and family, then the closest published LLMBoard score.

#80-8.0
MA

Kimi K2

Moonshot AI

51.4 LLMBoard

DetailsCompare
#29+12.2
MA

Kimi K2.5

Moonshot AI

71.6 LLMBoard

DetailsCompare
#7+23.4
MA

Kimi K2.6

Moonshot AI

82.8 LLMBoard

DetailsCompare
#54-0.3
NV

Nemotron 3 Super

NVIDIA

59.1 LLMBoard

DetailsCompare
#55-0.6
AN

Claude Opus 3

Anthropic

58.8 LLMBoard

DetailsCompare
#56-0.6
OP

GPT-4.5

OpenAI

58.8 LLMBoard

DetailsCompare

FAQ

Common questions about Kimi K2 Thinking.

When was Kimi K2 Thinking released?

Kimi K2 Thinking's default version was released on Sep 5, 2025.

How much does Kimi K2 Thinking cost?

No official standard PAYG price is currently available for Kimi K2 Thinking.

Who created Kimi K2 Thinking?

Kimi K2 Thinking is published under Moonshot AI in the model registry.

What is the context window for Kimi K2 Thinking?

The default version has a 262.1K token context window.

Is Kimi K2 Thinking open weight?

No. The default version is not marked as having publicly available weights.

How many API providers offer Kimi K2 Thinking?

No published provider offering is currently linked to the default version.

What models should I compare Kimi K2 Thinking with?

Nearby ranked alternatives include Qwen3.5 122B A10B, LongCat Flash Thinking, Nemotron Nano 9B.

Rankings

OverallCodingText ArenaPricing

Modalities

Image GenerationVideo GenerationSpeech-to-TextEmbeddings

Benchmarks

All BenchmarksReasoningMathCoding

Vendors

All VendorsOpenAIAnthropicGoogle
llmboard.aiCopyright 2026 llmboard.ai