reasoning benchmark
A graph reasoning benchmark that evaluates language models' ability to find parent nodes in graphs with context length over 128k tokens, testing long-context reasoning and graph structure understanding.
Updated Aug 7, 2026
Sorted by the source-provided rank. Higher score is better according to the registry.
| 1 | AN | 95.4% | 100.0% | 18 | C | |
| 4 | AN | 83.3% | 82.3% | 18 | C | |
| 9 | OP | 58.5% | 52.9% | 18 | C | |
| 14 | OP | 32.4% | 23.5% | 18 | C | |
| 15 | OP | 25.0% | 17.6% | 18 | C | |
| 16 | OP | 11.0% | 11.8% | 18 | C | |
| 18 | OP | 5.6% | 0.0% | 18 | C |
Top published rows on the benchmark's original scale.
The top published results on this benchmark's own scale.
Definition and scoring fields from the benchmark registry.
A graph reasoning benchmark that evaluates language models' ability to find parent nodes in graphs with context length over 128k tokens, testing long-context reasoning and graph structure understanding.
Scores are shown in ratio. The current registry marks this benchmark as not independently verified with evidence level B.
Source-native results are preserved. Eligibility for the overall LLMBoard score is a separate policy decision.
Common questions about Graphwalks parents >128k.
Claude Opus 4.6 is currently ranked first with 95.4%.
A graph reasoning benchmark that evaluates language models' ability to find parent nodes in graphs with context length over 128k tokens, testing long-context reasoning and graph structure understanding.
Yes. Higher values rank better for this benchmark.
7 unique published model results are currently shown.
This benchmark is preserved as source-native evidence but is not eligible for the current overall score.