reasoning benchmark
The Internal Research Debugging Evaluation measures whether models can debug 41 real bugs from internal OpenAI research experiments (plus alignment-auditing tasks), where the original solutions took experienced researchers hours to days. Passing corresponds to providing assistance that would unblock the user, including partial root-cause explanations or fixes.
Updated Aug 7, 2026
Sorted by the source-provided rank. Higher score is better according to the registry.
| 1 | OP | 68.3% | 100.0% | 3 | C | |
| 2 | OP | 67.8% | 50.0% | 3 | C | |
| 3 | OP | 50.8% | 0.0% | 3 | C |
Top published rows on the benchmark's original scale.
The top published results on this benchmark's own scale.
Definition and scoring fields from the benchmark registry.
The Internal Research Debugging Evaluation measures whether models can debug 41 real bugs from internal OpenAI research experiments (plus alignment-auditing tasks), where the original solutions took experienced researchers hours to days. Passing corresponds to providing assistance that would unblock the user, including partial root-cause explanations or fixes.
Scores are shown in ratio. The current registry marks this benchmark as not independently verified with evidence level B.
Source-native results are preserved. Eligibility for the overall LLMBoard score is a separate policy decision.
Common questions about Internal Research Debugging Evaluation.
GPT-5.6 Sol is currently ranked first with 68.3%.
The Internal Research Debugging Evaluation measures whether models can debug 41 real bugs from internal OpenAI research experiments (plus alignment-auditing tasks), where the original solutions took experienced researchers hours to days. Passing corresponds to providing assistance that would unblock the user, including partial root-cause explanations or fixes.
Yes. Higher values rank better for this benchmark.
3 unique published model results are currently shown.
This benchmark is preserved as source-native evidence but is not eligible for the current overall score.