llmboard.ai
Benchmarks
CompareRankings
llmboard.ai
Benchmarks
CompareRankings
HomeModelsPhi 4 multimodal

Microsoft model product

Phi 4 multimodal

Phi-4-multimodal-instruct is a lightweight (5.57B parameters) open multimodal foundation model that leverages research and datasets from Phi-3.5 and 4.0. It processes text, image, and audio inputs to generate text outputs, supporting a 128K token context length. Enhanced via SFT, DPO, and RLHF for instruction following and safety.

Updated Aug 10, 2026. Default version: Phi-4-multimodal-instruct

Compare
LLMBoard scoreN/ANot scored
CoverageN/ANo published score
Context window128KTokens
Official input priceN/AOfficial price unavailable

On this page

  • Specification
  • Capability
  • Benchmarks
  • Arena
  • Pricing
  • Versions
  • About
  • Similar models
  • FAQ

Model specification

Structured fields from the published default version.

Version
Phi-4-multimodal-instruct
Released
Feb 1, 2025
Knowledge cutoff
Jun 1, 2024
Parameters
5.6B
Context window
128K
Max output
128K
Inputs
image, text
Outputs
text
Open weights
No
License
MIT

Capability profile

This profile uses the latest version under this unique model that has a calculated LLMBoard score. Arena and price are excluded.

No capability profile

No published version under this unique model currently has enough benchmark coverage to calculate a score.

Benchmark results

Published benchmark records for the scored version currently unavailable.

No benchmark rows

No published benchmark result is linked to the scored version.

Arena results

Preference and agent-evaluation signals from published Arena datasets.

No published Arena match

The default version has no published Arena rows, or its source alias has not been resolved.

Pricing

Official vendor API PAYG pricing is summarized first. The table then lists individual provider offerings without treating their minimum as the official price.

Official API
N/A
Official provider
N/A
Lowest third-party
From $0.07 input, $0.11 output per 1M via NanoGPT
Tracked offerings
1
1 rows
Columns

Show columns

NanoGPTphi-4-multimodal-instructglobal$0.07$0.11128KAug 7, 2026

Official prices use only the vendor's configured official Provider and positive standard USD PAYG rates. Third-party offers remain explicitly labeled.

Phi 4 multimodal versions

All published versions linked to this unique model. The score columns identify the version used by the current overall ranking.

1 rows
Columns

Show columns

Phi-4-multimodal-instructFeb 1, 2025N/A5.6B128K128KNoMIT

What is Phi 4 multimodal?

A concise description based on the published model registry.

Phi-4-multimodal-instruct is a lightweight (5.57B parameters) open multimodal foundation model that leverages research and datasets from Phi-3.5 and 4.0. It processes text, image, and audio inputs to generate text outputs, supporting a 128K token context length. Enhanced via SFT, DPO, and RLHF for instruction following and safety.

Use the benchmark, Arena and pricing sections above as separate evidence. A missing field means the current data snapshot does not support that claim.

Data snapshot: 2026-08-07. Editorial model content is not available in the backend.

Models similar to Phi 4 multimodal

Recommendations prioritize the same model type and family, then the closest published LLMBoard score.

#144
MI

Phi 3.5 mini

Microsoft

25.9 LLMBoard

DetailsCompare
#134
MI

Phi 4 Mini

Microsoft

31.5 LLMBoard

DetailsCompare
#121
MI

Phi 4

Microsoft

36.9 LLMBoard

DetailsCompare
#107
MI

Phi 3.5 MoE

Microsoft

42.4 LLMBoard

DetailsCompare
#92
MI

Phi 4 Reasoning

Microsoft

47.9 LLMBoard

DetailsCompare
#83
MI

Phi 4 Reasoning Plus

Microsoft

50.3 LLMBoard

DetailsCompare

FAQ

Common questions about Phi 4 multimodal.

When was Phi 4 multimodal released?

Phi 4 multimodal's default version was released on Feb 1, 2025.

How much does Phi 4 multimodal cost?

No official standard PAYG price is currently available for Phi 4 multimodal. The lowest tracked third-party offer starts at $0.07 input and $0.11 output via NanoGPT.

Who created Phi 4 multimodal?

Phi 4 multimodal is published under Microsoft in the model registry.

What is the context window for Phi 4 multimodal?

The default version has a 128K token context window.

Is Phi 4 multimodal open weight?

No. The default version is not marked as having publicly available weights.

How many API providers offer Phi 4 multimodal?

1 published provider offerings are linked to the default version.

What models should I compare Phi 4 multimodal with?

No nearby ranked alternatives are currently available.

Rankings

OverallCodingText ArenaPricing

Modalities

Image GenerationVideo GenerationSpeech-to-TextEmbeddings

Benchmarks

All BenchmarksReasoningMathCoding

Vendors

All VendorsOpenAIAnthropicGoogle
llmboard.aiCopyright 2026 llmboard.ai