llmboard.ai
Benchmarks
CompareRankings
llmboard.ai
Benchmarks
CompareRankings
HomeBenchmarksreasoningTAU-bench Airline

reasoning benchmark

TAU-bench Airline

Part of τ-bench (TAU-bench), a benchmark for Tool-Agent-User interaction in real-world domains. The airline domain evaluates language agents' ability to interact with users through dynamic conversations while following domain-specific rules and using API tools. Agents must handle airline-related tasks and policies reliably.

Updated Aug 7, 2026

Published models23
Registry coverage23
MetricScore
EvidenceB

On this page

  • Ranking
  • Distribution
  • Highlights
  • About
  • FAQ

TAU-bench Airline leaderboard

Sorted by the source-provided rank. Higher score is better according to the registry.

23 rows
Columns

Show columns

1ANClaude Sonnet 4.5Anthropic70.0%100.0%23CAug 7, 2026
2MIMiniMax M1 80KMiniMax62.0%95.5%23CAug 7, 2026
3ZAGLM-4.5-AirZhipu AI60.8%90.9%23CAug 7, 2026
4ZAGLM-4.5Zhipu AI60.4%86.4%23CAug 7, 2026
5ANClaude Sonnet 4Anthropic60.0%81.8%23CAug 7, 2026
6MIMiniMax M1 40KMiniMax60.0%77.3%23CAug 7, 2026
7ACQwen3-Coder 480B A35B InstructAlibaba Cloud / Qwen Team60.0%72.7%23CAug 7, 2026
8ANClaude Opus 4Anthropic59.6%68.2%23CAug 7, 2026
9ANClaude 3.7 SonnetAnthropic58.4%63.6%23CAug 7, 2026
10ANClaude Opus 4.1Anthropic56.0%59.1%23CAug 7, 2026
11OPGPT-4.5OpenAI50.0%54.5%23CAug 7, 2026
12OPo1OpenAI50.0%50.0%23CAug 7, 2026
13OPGPT-4.1OpenAI49.4%45.5%23CAug 7, 2026
14OPo4-miniOpenAI49.2%40.9%23CAug 7, 2026
15ACQwen3-Next-80B-A3B-ThinkingAlibaba Cloud / Qwen Team49.0%36.4%23CAug 7, 2026
16ANClaude 3.5 SonnetAnthropic46.0%31.8%23CAug 7, 2026
17ACQwen3-235B-A22B-Thinking-2507Alibaba Cloud / Qwen Team46.0%27.3%23CAug 7, 2026
18ACQwen3-Next-80B-A3B-InstructAlibaba Cloud / Qwen Team44.0%22.7%23CAug 7, 2026
19OPGPT-4oOpenAI42.8%18.2%23CAug 7, 2026
20OPGPT-4.1 miniOpenAI36.0%13.6%23CAug 7, 2026
21OPo3-miniOpenAI32.4%9.1%23CAug 7, 2026
22ANClaude 3.5 HaikuAnthropic22.8%4.5%23CAug 7, 2026
23OPGPT-4.1 nanoOpenAI14.0%0.0%23CAug 7, 2026

Score distribution

Top published rows on the benchmark's original scale.

TAU-bench Airline

TAU-bench Airline highlights

The top published results on this benchmark's own scale.

Rank #1Claude Sonnet 4.570.0%Rank #2MiniMax M1 80K62.0%Rank #3GLM-4.5-Air60.8%Rank #4GLM-4.560.4%

What is TAU-bench Airline?

Definition and scoring fields from the benchmark registry.

Part of τ-bench (TAU-bench), a benchmark for Tool-Agent-User interaction in real-world domains. The airline domain evaluates language agents' ability to interact with users through dynamic conversations while following domain-specific rules and using API tools. Agents must handle airline-related tasks and policies reliably.

Scores are shown in ratio. The current registry marks this benchmark as not independently verified with evidence level B.

Family
TAU-bench Airline
Modality
text
Primary category
reasoning
Score direction
higher
LLMBoard eligible
No
Evaluation key
tau-bench-airline|llm-stats-current

Source-native results are preserved. Eligibility for the overall LLMBoard score is a separate policy decision.

FAQ

Common questions about TAU-bench Airline.

Which model scores highest on TAU-bench Airline?

Claude Sonnet 4.5 is currently ranked first with 70.0%.

What does TAU-bench Airline measure?

Part of τ-bench (TAU-bench), a benchmark for Tool-Agent-User interaction in real-world domains. The airline domain evaluates language agents' ability to interact with users through dynamic conversations while following domain-specific rules and using API tools. Agents must handle airline-related tasks and policies reliably.

Is a higher score better?

Yes. Higher values rank better for this benchmark.

How many models are compared?

23 unique published model results are currently shown.

Does this benchmark affect the overall score?

This benchmark is preserved as source-native evidence but is not eligible for the current overall score.

Rankings

OverallCodingText ArenaPricing

Modalities

Image GenerationVideo GenerationSpeech-to-TextEmbeddings

Benchmarks

All BenchmarksReasoningMathCoding

Vendors

All VendorsOpenAIAnthropicGoogle
llmboard.aiCopyright 2026 llmboard.ai