llmboard.ai
Benchmarks
CompareRankings
llmboard.ai
Benchmarks
CompareRankings
HomeModelsGrok 4.1 Thinking

xAI model product

Grok 4.1 Thinking

Grok 4.1 Thinking (code name: quasarflux) is the reasoning variant of Grok 4.1, bringing significant improvements to the real-world usability of Grok. The model is exceptionally capable in creative, emotional, and collaborative interactions. It is more perceptive to nuanced intent, compelling to speak with, and coherent in personality, while fully retaining the razor-sharp intelligence and reliability of its predecessors. Grok 4.1 Thinking uses thinking tokens for deeper reasoning, achieving state-of-the-art performance on blind human preference evaluations. The model features reduced hallucinations compared to previous versions, with significant improvements in factual accuracy for information-seeking prompts. To achieve this, xAI used large scale reinforcement learning infrastructure to optimize style, personality, helpfulness, and alignment, developing new methods that use frontier agentic reasoning models as reward models to autonomously evaluate and iterate on responses at scale. Grok 4.1 Thinking includes comprehensive safety mitigations including refusal training, input filters for restricted knowledge, and adversarial robustness measures. The model demonstrates strong refusal rates on harmful queries (7% answer rate on chat refusals, 14% on agentic refusals) and improved honesty training to reduce deception.

Updated Aug 10, 2026. Default version: Grok-4.1 Thinking

Compare
LLMBoard scoreN/ANot scored
CoverageN/ANo published score
Context window256KTokens
Official input priceN/AOfficial price unavailable

On this page

  • Specification
  • Capability
  • Benchmarks
  • Arena
  • Pricing
  • Versions
  • About
  • Similar models
  • FAQ

Model specification

Structured fields from the published default version.

Version
Grok-4.1 Thinking
Released
Nov 17, 2025
Knowledge cutoff
Unknown
Parameters
N/A
Context window
256K
Max output
8K
Inputs
image, text
Outputs
text
Open weights
No
License
Proprietary

Capability profile

This profile uses the latest version under this unique model that has a calculated LLMBoard score. Arena and price are excluded.

No capability profile

No published version under this unique model currently has enough benchmark coverage to calculate a score.

Benchmark results

Published benchmark records for the scored version currently unavailable.

No benchmark rows

No published benchmark result is linked to the scored version.

Arena results

Preference and agent-evaluation signals from published Arena datasets.

89 rows
Columns

Show columns

text style controlindustry medicine and healthcare91502.24,288N/AAug 6, 2026
text style controlpolish141488.21,853N/AAug 6, 2026
text factualitykorean231396.2495N/AAug 6, 2026
text factualitypolish251428.41,189N/AAug 6, 2026
textpolish281457.01,853N/AAug 6, 2026
text style controlfrench301487.11,584N/AAug 6, 2026
text style controlkorean301429.91,072N/AAug 6, 2026
text factualitygerman311398.7536N/AAug 6, 2026
text style controlspanish321465.01,708N/AAug 6, 2026
text style controlindustry legal and government331477.14,714N/AAug 6, 2026
text factualityspanish381438.91,119N/AAug 6, 2026
text style controlexclude ties381473.547,995N/AAug 6, 2026
text style controlnon english391455.434,869N/AAug 6, 2026
text style controloverall401466.065,044N/AAug 6, 2026
text style controlenglish421471.730,172N/AAug 6, 2026
text style controlgerman441460.81,094N/AAug 6, 2026
text style controlindustry business and management and financial operations441460.612,471N/AAug 6, 2026
text style controlindustry entertainment and sports and media461434.612,297N/AAug 6, 2026
text style controlindustry life and physical and social science471479.310,694N/AAug 6, 2026
text style controlindustry software and it services471496.823,485N/AAug 6, 2026
text style controlrussian481459.36,638N/AAug 6, 2026
textindustry medicine and healthcare501453.94,288N/AAug 6, 2026
text factualityfrench501453.41,072N/AAug 6, 2026
text style controlhard prompts521477.637,141N/AAug 6, 2026
textkorean531396.81,072N/AAug 6, 2026
text factualitycreative writing541424.26,222N/AAug 6, 2026
text factualityindustry entertainment and sports and media571419.28,284N/AAug 6, 2026
textgerman581436.61,094N/AAug 6, 2026
text style controlcreative writing581428.59,725N/AAug 6, 2026
textindustry entertainment and sports and media591408.612,297N/AAug 6, 2026
text style controlhard prompts english591478.718,067N/AAug 6, 2026
text style controlchinese601490.33,443N/AAug 6, 2026
text style controlcoding601498.614,988N/AAug 6, 2026
textnon english621427.834,869N/AAug 6, 2026
textrussian621431.86,638N/AAug 6, 2026
textexclude ties631432.247,995N/AAug 6, 2026
text style controlindustry writing and literature and language631428.914,667N/AAug 6, 2026
text style controljapanese631407.8555N/AAug 6, 2026
text style controlmulti turn631459.712,052N/AAug 6, 2026
textfrench641451.11,584N/AAug 6, 2026
textspanish641438.51,708N/AAug 6, 2026
textoverall651437.665,044N/AAug 6, 2026
text style controlmath651442.93,795N/AAug 6, 2026
text factualitymath671423.72,313N/AAug 6, 2026
textcreative writing681406.89,725N/AAug 6, 2026
text factualityrussian681438.73,983N/AAug 6, 2026
textenglish701442.430,172N/AAug 6, 2026
textindustry legal and government701440.04,714N/AAug 6, 2026
text factualityindustry medicine and healthcare701462.62,701N/AAug 6, 2026
text factualityindustry mathematical721416.42,031N/AAug 6, 2026
text style controlexpert731467.54,512N/AAug 6, 2026
text style controlinstruction following751430.518,355N/AAug 6, 2026
text factualityindustry writing and literature and language761418.99,974N/AAug 6, 2026
text style controllonger query761447.319,791N/AAug 6, 2026
textindustry software and it services781451.923,485N/AAug 6, 2026
text factualitymulti turn781440.38,194N/AAug 6, 2026
text factualitynon english781421.834,114N/AAug 6, 2026
text factualityindustry legal and government801439.42,989N/AAug 6, 2026
textmulti turn811432.412,052N/AAug 6, 2026
text factualitychinese811451.02,293N/AAug 6, 2026
text factualityexpert811445.82,907N/AAug 6, 2026
texthard prompts821434.537,141N/AAug 6, 2026
textjapanese841374.5555N/AAug 6, 2026
text factualityexclude ties841433.347,145N/AAug 6, 2026
textcoding871444.714,988N/AAug 6, 2026
text factualityoverall871434.064,137N/AAug 6, 2026
text factualityindustry business and management and financial operations871430.98,378N/AAug 6, 2026
text style controlindustry mathematical871437.73,012N/AAug 6, 2026
textindustry writing and literature and language881406.314,667N/AAug 6, 2026
textindustry business and management and financial operations891417.712,471N/AAug 6, 2026
textindustry life and physical and social science891438.310,694N/AAug 6, 2026
text factualitycoding901474.910,619N/AAug 6, 2026
text factualityenglish911443.329,270N/AAug 6, 2026
texthard prompts english941435.418,067N/AAug 6, 2026
textmath951422.83,795N/AAug 6, 2026
text factualityindustry software and it services961463.319,703N/AAug 6, 2026
text factualityhard prompts english981453.213,259N/AAug 6, 2026
text factualityinstruction following991413.613,645N/AAug 6, 2026
text factualityhard prompts1011446.036,381N/AAug 6, 2026
text factualityindustry life and physical and social science1011440.17,100N/AAug 6, 2026
textchinese1021454.13,443N/AAug 6, 2026
webdevwebdev-html1031220.9944N/AAug 6, 2026
webdevoverall1041210.1944N/AAug 6, 2026
webdevwebdev1041210.1944N/AAug 6, 2026
textexpert1081419.84,512N/AAug 6, 2026
text factualitylonger query1091426.615,251N/AAug 6, 2026
textinstruction following1111396.418,355N/AAug 6, 2026
textlonger query1111408.619,791N/AAug 6, 2026
textindustry mathematical1181412.83,012N/AAug 6, 2026

Pricing

Official vendor API PAYG pricing is summarized first. The table then lists individual provider offerings without treating their minimum as the official price.

Official API
N/A
Official provider
N/A
Lowest third-party
N/A
Tracked offerings
0
No published price snapshot

The default version has no provider offering with current input or output token prices.

Official prices use only the vendor's configured official Provider and positive standard USD PAYG rates. Third-party offers remain explicitly labeled.

Grok 4.1 Thinking versions

All published versions linked to this unique model. The score columns identify the version used by the current overall ranking.

1 rows
Columns

Show columns

Grok-4.1 ThinkingNov 17, 2025N/AN/A256K8KNoProprietary

What is Grok 4.1 Thinking?

A concise description based on the published model registry.

Grok 4.1 Thinking (code name: quasarflux) is the reasoning variant of Grok 4.1, bringing significant improvements to the real-world usability of Grok. The model is exceptionally capable in creative, emotional, and collaborative interactions. It is more perceptive to nuanced intent, compelling to speak with, and coherent in personality, while fully retaining the razor-sharp intelligence and reliability of its predecessors. Grok 4.1 Thinking uses thinking tokens for deeper reasoning, achieving state-of-the-art performance on blind human preference evaluations. The model features reduced hallucinations compared to previous versions, with significant improvements in factual accuracy for information-seeking prompts. To achieve this, xAI used large scale reinforcement learning infrastructure to optimize style, personality, helpfulness, and alignment, developing new methods that use frontier agentic reasoning models as reward models to autonomously evaluate and iterate on responses at scale. Grok 4.1 Thinking includes comprehensive safety mitigations including refusal training, input filters for restricted knowledge, and adversarial robustness measures. The model demonstrates strong refusal rates on harmful queries (7% answer rate on chat refusals, 14% on agentic refusals) and improved honesty training to reduce deception.

Use the benchmark, Arena and pricing sections above as separate evidence. A missing field means the current data snapshot does not support that claim.

Data snapshot: 2026-08-07. Editorial model content is not available in the backend.

Models similar to Grok 4.1 Thinking

Recommendations prioritize the same model type and family, then the closest published LLMBoard score.

#156
XA

Grok 1.5

xAI

19.3 LLMBoard

DetailsCompare
#37
XA

Grok 4 Fast

xAI

66.2 LLMBoard

DetailsCompare
#170
AC

Qwen3.5 0.8B

Alibaba Cloud / Qwen Team

0.3 LLMBoard

DetailsCompare
#169
GO

Gemma 3 1B

Google

2.7 LLMBoard

DetailsCompare
#168
GO

Gemma 3n E2B

Google

6.9 LLMBoard

DetailsCompare
#167
AC

Qwen3 VL 4B

Alibaba Cloud / Qwen Team

7.7 LLMBoard

DetailsCompare

FAQ

Common questions about Grok 4.1 Thinking.

When was Grok 4.1 Thinking released?

Grok 4.1 Thinking's default version was released on Nov 17, 2025.

How much does Grok 4.1 Thinking cost?

No official standard PAYG price is currently available for Grok 4.1 Thinking.

Who created Grok 4.1 Thinking?

Grok 4.1 Thinking is published under xAI in the model registry.

What is the context window for Grok 4.1 Thinking?

The default version has a 256K token context window.

Is Grok 4.1 Thinking open weight?

No. The default version is not marked as having publicly available weights.

How many API providers offer Grok 4.1 Thinking?

No published provider offering is currently linked to the default version.

What models should I compare Grok 4.1 Thinking with?

No nearby ranked alternatives are currently available.

Rankings

OverallCodingText ArenaPricing

Modalities

Image GenerationVideo GenerationSpeech-to-TextEmbeddings

Benchmarks

All BenchmarksReasoningMathCoding

Vendors

All VendorsOpenAIAnthropicGoogle
llmboard.aiCopyright 2026 llmboard.ai