#170
AC
Qwen3.5 0.8B
Alibaba Cloud / Qwen Team
0.3 LLMBoard
OpenBMB model product
MiniCPM-SALA (Sparse Attention and Linear Attention) is a 9B hybrid model built from a MiniCPM-4.0 checkpoint via continual training (~2T tokens, 25% of training-from-scratch cost). It interleaves 25% InfLLM-V2 sparse attention and 75% Lightning Attention layers, achieving up to 3.5x inference speed over dense baselines at 256K tokens. With HyPE (Hybrid Positional Encoding) and NoPE in sparse layers, the model extrapolates to 2048K tokens despite a 520K training length, enabling 1M-token inference on consumer GPUs like the RTX 5090.
Updated Aug 10, 2026. Default version: MiniCPM-SALA
Structured fields from the published default version.
This profile uses the latest version under this unique model that has a calculated LLMBoard score. Arena and price are excluded.
No published version under this unique model currently has enough benchmark coverage to calculate a score.
Published benchmark records for the scored version currently unavailable.
No published benchmark result is linked to the scored version.
Preference and agent-evaluation signals from published Arena datasets.
The default version has no published Arena rows, or its source alias has not been resolved.
Official vendor API PAYG pricing is summarized first. The table then lists individual provider offerings without treating their minimum as the official price.
The default version has no provider offering with current input or output token prices.
Official prices use only the vendor's configured official Provider and positive standard USD PAYG rates. Third-party offers remain explicitly labeled.
All published versions linked to this unique model. The score columns identify the version used by the current overall ranking.
| MiniCPM-SALA | N/A | 9.5B | N/A | N/A | No | Apache 2.0 |
A concise description based on the published model registry.
MiniCPM-SALA (Sparse Attention and Linear Attention) is a 9B hybrid model built from a MiniCPM-4.0 checkpoint via continual training (~2T tokens, 25% of training-from-scratch cost). It interleaves 25% InfLLM-V2 sparse attention and 75% Lightning Attention layers, achieving up to 3.5x inference speed over dense baselines at 256K tokens. With HyPE (Hybrid Positional Encoding) and NoPE in sparse layers, the model extrapolates to 2048K tokens despite a 520K training length, enabling 1M-token inference on consumer GPUs like the RTX 5090.
Use the benchmark, Arena and pricing sections above as separate evidence. A missing field means the current data snapshot does not support that claim.
Data snapshot: 2026-08-07. Editorial model content is not available in the backend.
Recommendations prioritize the same model type and family, then the closest published LLMBoard score.
Common questions about MiniCPM SALA.
MiniCPM SALA's default version was released on Feb 11, 2026.
No official standard PAYG price is currently available for MiniCPM SALA.
MiniCPM SALA is published under OpenBMB in the model registry.
The current registry does not publish a context window for the default version.
No. The default version is not marked as having publicly available weights.
No published provider offering is currently linked to the default version.
No nearby ranked alternatives are currently available.