llmboard.ai
Benchmarks
CompareRankings
llmboard.ai
Benchmarks
CompareRankings
HomeBenchmarksreasoningBFCL_v3_MultiTurn

reasoning benchmark

BFCL_v3_MultiTurn

Berkeley Function Calling Leaderboard (BFCL) V3 MultiTurn benchmark that evaluates large language models' ability to handle multi-turn and multi-step function calling scenarios. The benchmark introduces complex interactions requiring models to manage sequential function calls, handle conversational context across multiple turns, and make dynamic decisions about when and how to use available functions. BFCL V3 uses state-based evaluation by verifying the actual state of API systems after function execution, providing more realistic assessment of function calling capabilities in agentic applications.

Updated Aug 7, 2026

Published models2
Registry coverage2
MetricScore
EvidenceB

On this page

  • Ranking
  • Distribution
  • Highlights
  • About
  • FAQ

BFCL_v3_MultiTurn leaderboard

Sorted by the source-provided rank. Higher score is better according to the registry.

2 rows
Columns

Show columns

1MIMiniMax M2.5MiniMax76.8%100.0%2CAug 7, 2026
2NVNemotron Nano 9B v2NVIDIA66.9%0.0%2CAug 7, 2026

Score distribution

Top published rows on the benchmark's original scale.

BFCL_v3_MultiTurn

BFCL_v3_MultiTurn highlights

The top published results on this benchmark's own scale.

Rank #1MiniMax M2.576.8%Rank #2Nemotron Nano 9B v266.9%

What is BFCL_v3_MultiTurn?

Definition and scoring fields from the benchmark registry.

Berkeley Function Calling Leaderboard (BFCL) V3 MultiTurn benchmark that evaluates large language models' ability to handle multi-turn and multi-step function calling scenarios. The benchmark introduces complex interactions requiring models to manage sequential function calls, handle conversational context across multiple turns, and make dynamic decisions about when and how to use available functions. BFCL V3 uses state-based evaluation by verifying the actual state of API systems after function execution, providing more realistic assessment of function calling capabilities in agentic applications.

Scores are shown in ratio. The current registry marks this benchmark as not independently verified with evidence level B.

Family
BFCL_v3_MultiTurn
Modality
text
Primary category
reasoning
Score direction
higher
LLMBoard eligible
No
Evaluation key
bfcl-v3-multiturn|llm-stats-current

Source-native results are preserved. Eligibility for the overall LLMBoard score is a separate policy decision.

FAQ

Common questions about BFCL_v3_MultiTurn.

Which model scores highest on BFCL_v3_MultiTurn?

MiniMax M2.5 is currently ranked first with 76.8%.

What does BFCL_v3_MultiTurn measure?

Berkeley Function Calling Leaderboard (BFCL) V3 MultiTurn benchmark that evaluates large language models' ability to handle multi-turn and multi-step function calling scenarios. The benchmark introduces complex interactions requiring models to manage sequential function calls, handle conversational context across multiple turns, and make dynamic decisions about when and how to use available functions. BFCL V3 uses state-based evaluation by verifying the actual state of API systems after function execution, providing more realistic assessment of function calling capabilities in agentic applications.

Is a higher score better?

Yes. Higher values rank better for this benchmark.

How many models are compared?

2 unique published model results are currently shown.

Does this benchmark affect the overall score?

This benchmark is preserved as source-native evidence but is not eligible for the current overall score.

Rankings

OverallCodingText ArenaPricing

Modalities

Image GenerationVideo GenerationSpeech-to-TextEmbeddings

Benchmarks

All BenchmarksReasoningMathCoding

Vendors

All VendorsOpenAIAnthropicGoogle
llmboard.aiCopyright 2026 llmboard.ai