reasoning benchmark
WinoGrande: An Adversarial Winograd Schema Challenge at Scale. A large-scale dataset of 44,000 pronoun resolution problems designed to test machine commonsense reasoning. Uses adversarial filtering to reduce spurious biases and provides a more robust evaluation of whether AI systems truly understand commonsense or exploit statistical shortcuts. Current best AI methods achieve 59.4-79.1% accuracy, significantly below human performance of 94.0%.
Updated Aug 7, 2026
Sorted by the source-provided rank. Higher score is better according to the registry.
| 1 | OP | 87.5% | 100.0% | 22 | C | |
| 2 | XI | 85.6% | 95.2% | 22 | C | |
| 3 | CO | 85.4% | 90.5% | 22 | C | |
| 4 | AC | 85.1% | 85.7% | 22 | C | |
| 5 | NV | 84.5% | 81.0% | 22 | C | |
| 6 | GO | 83.7% | 76.2% | 22 | C | |
| 7 | NR | 83.2% | 71.4% | 22 | C | |
| 8 | AC | 82.0% | 66.7% | 22 | C | |
| 9 | MI | 81.3% | 61.9% | 22 | C | |
| 10 | AC | 80.8% | 57.1% | 22 | C | |
| 11 | GO | 80.6% | 52.4% | 22 | C | |
| 12 | MA | 76.8% | 47.6% | 22 | C | |
| 13 | MA | 75.3% | 42.9% | 22 | C | |
| 14 | IB | 74.4% | 38.1% | 22 | C | |
| 15 | AC | 72.9% | 33.3% | 22 | C | |
| 16 | GO | 71.7% | 28.6% | 22 | C | |
| 17 | GO | 71.7% | 23.8% | 22 | C | |
| 18 | MI | 68.5% | 19.1% | 22 | C | |
| 19 | MI | 67.0% | 14.3% | 22 | C | |
| 20 | GO | 66.8% | 9.5% | 22 | C | |
| 21 | GO | 66.8% | 4.8% | 22 | C | |
| 22 | BA | 51.3% | 0.0% | 22 | C |
Top published rows on the benchmark's original scale.
The top published results on this benchmark's own scale.
Definition and scoring fields from the benchmark registry.
WinoGrande: An Adversarial Winograd Schema Challenge at Scale. A large-scale dataset of 44,000 pronoun resolution problems designed to test machine commonsense reasoning. Uses adversarial filtering to reduce spurious biases and provides a more robust evaluation of whether AI systems truly understand commonsense or exploit statistical shortcuts. Current best AI methods achieve 59.4-79.1% accuracy, significantly below human performance of 94.0%.
Scores are shown in ratio. The current registry marks this benchmark as not independently verified with evidence level B.
Source-native results are preserved. Eligibility for the overall LLMBoard score is a separate policy decision.
Common questions about Winogrande.
GPT-4 is currently ranked first with 87.5%.
WinoGrande: An Adversarial Winograd Schema Challenge at Scale. A large-scale dataset of 44,000 pronoun resolution problems designed to test machine commonsense reasoning. Uses adversarial filtering to reduce spurious biases and provides a more robust evaluation of whether AI systems truly understand commonsense or exploit statistical shortcuts. Current best AI methods achieve 59.4-79.1% accuracy, significantly below human performance of 94.0%.
Yes. Higher values rank better for this benchmark.
22 unique published model results are currently shown.
This benchmark is marked as eligible for the current LLMBoard capability methodology.