Last updated: 2026-07-20 · Report outdated data on GitHub
MiMo API Pricing — July 2026
Frontier-class reasoning at 5–20% of closed-source pricing. V2.5-Pro list input cut to $0.42/M in May 2026; cache-hit input cut 99% to $0.0035/M. All prices in USD per 1 million tokens.
Bottom line: V2.5-Pro list ~$0.42/$0.83, cache-hit input 99% off
MiMo-V2.5-Pro list pricing is now $0.42/M input + $0.83/M output (¥3 / ¥6 per 1M CNY). For workloads that can benefit from prompt caching — chat history, RAG, code review, long-context retrieval — the cache-hit input rate is $0.0035/M (¥0.025/M, 99% off from the original 2025 launch rate). The May 2026 cut was specifically for cache-hit pricing; list output dropped 86% (V2.5-Pro) and 93% (V2.5 standard). MiMo-V2.5-Pro matches Claude Opus 4.5 on long-context reasoning and is ~97% cheaper on list input ($0.42 vs $15) and ~99% cheaper on cache-hit input. MiMo-7B is free under MIT (zero per-token cost when self-hosted).
All MiMo weights are MIT-licensed. Self-hosting eliminates per-token cost entirely for teams with GPU infrastructure.
→ Use the Token Cost Calculator (real-time USD/CNY)
1. The V2.5 lineup (full pricing)
All prices per 1 million tokens. List = pay-as-you-go rate; cache-hit = applies to repeated tokens when prompt caching is enabled. Verified July 20, 2026 against the official 2026-05-27 announcement.
| Model | Params | Context | Best for | List input / 1M | Cache-hit input / 1M | List output / 1M | Change since May 2026 |
|---|---|---|---|---|---|---|---|
| MiMo-V2.5-Pro | 1T+ MoE | 1M | Agent, long-context, reasoning | $0.42 | $0.0035 | $0.83 | Cache-hit −99% · Output −86% |
| MiMo-V2.5-Pro-UltraSpeed | 1T+ MoE (FP4+DFlash) | 1M | High-throughput coding, 1000+ tok/s | $1.25 | $0.0104 | $2.50 | Internal-test only (apply for access) |
| MiMo-V2.5 (standard) | ~309B MoE | 1M | General chat, content, lightweight agent | $0.14 | $0.0028 | $0.28 | Cache-hit −98% · Output −93% |
| MiMo-V2.5-Omni | TBD | TBD | Multimodal (text / image / video / audio) | TBD | TBD | TBD | MIT weights out; API TBD |
| MiMo-V2.5-TTS | TBD | N/A | Bilingual speech synthesis (CN/EN) | Commercial | license | license | Commercially licensed separately |
| MiMo-V2.5-ASR | TBD | N/A | Speech-to-text (CN/EN) | ¥0.5/hr | (audio input) | — | Open weights (Apache-2.0) |
| MiMo-7B | 7B dense | 32k | Edge / on-device (phones, cars, IoT) | $0.00 | $0.00 | $0.00 | MIT weights, free self-host |
| MiMo Code | Agent layer | Infinite | Terminal-native coding agent | Free app | + API usage | + API usage | Open source, MIT |
| MiMo-V2-Flash (deprecated) | 309B / 15B active | 56k | Replaced by V2.5 standard | — | — | — | Offline 2026-06-30 |
2. Side-by-side: same workload, 5 leading APIs
All prices list (USD per 1M tokens). Prompt caching and batch API rates are not included. Verified July 20, 2026 against provider list pages. Comparison baseline = MiMo-V2.5-Pro list input $0.42.
| Provider | Model | List input / 1M | List output / 1M | vs MiMo-V2.5-Pro list input |
|---|---|---|---|---|
| Xiaomi | MiMo-V2.5-Pro | $0.42 | $0.83 | — baseline — |
| Xiaomi | MiMo-V2.5 (standard) | $0.14 | $0.28 | ~67% cheaper input |
| DeepSeek | DeepSeek V4 Pro | $0.435 | $0.87 | ~3% more expensive input (similar tier) |
| DeepSeek | DeepSeek-R1 | $0.55 | $2.19 | ~31% more expensive input |
| Alibaba | Qwen3 Max | $0.60 | $2.40 | ~43% more expensive input |
| OpenAI | GPT-4o mini | $0.15 | $0.60 | ~64% cheaper input (smaller model) |
| OpenAI | GPT-5 | $1.25 | $10.00 | ~3× more expensive input |
| Anthropic | Claude Haiku 3.5 | $0.80 | $4.00 | ~90% more expensive input (smaller model) |
| Anthropic | Claude Opus 4.5 | $15.00 | $75.00 | ~36× more expensive input |
Reading the table: MiMo-V2.5-Pro at $0.42/$0.83 list matches Claude Opus 4.5 on long-context reasoning, beats Claude Haiku 3.5 on coding (73.4% vs 65% on SWE-Bench Verified), and is priced just below DeepSeek V4 Pro on input. The 1M context window (vs 128k for V4 Pro) is the deal-breaker for legal / scientific / code-review workloads. For prompt-cacheable workloads (chat history, RAG, repeated system prompts), the cache-hit input rate ($0.0035/M) makes MiMo-V2.5-Pro effectively free on input.
3. Why cache-hit pricing matters
Most real-world LLM applications send the same tokens multiple times: system prompts, retrieved documents in RAG, prior conversation history, code files under review. MiMo (and most modern APIs) offers prompt caching: a flag at request time, and the repeated prefix is billed at the much lower cache-hit rate. On MiMo-V2.5-Pro:
- A 50k-token system prompt + 200k-token document corpus sent on every request
- List rate: 250k × $0.42/M = $0.105 per request
- Cache-hit rate: 250k × $0.0035/M = $0.000875 per request — a 120× reduction on the cached portion
- Add a 1k-token user query (uncached) + 500-token answer (output $0.83/M = $0.00042) → total $0.0013 per request
This is the reason the May 2026 cut was framed as "99% off" — for the kind of repeated-token workloads that dominate production agent use, the effective per-request cost is genuinely near zero on the input side. See the calculator for your workload →
4. Pricing history (timeline)
Permanent cache-hit price cut (−98% to −99%). V2.5-Pro cache-hit input drops to ¥0.025/M ($0.0035), V2.5 standard cache-hit input to ¥0.02/M ($0.0028). List output also cut: V2.5-Pro ¥6/M (−86% from launch), V2.5 standard ¥2/M (−93%). Xiaomi cites "ecosystem expansion over short-term margin" as the rationale. Backed by SGLang HiCache + sliding-window attention reducing KV-cache movement to 1/7 of pre-optimization cost.
MiMo-V2 series decommissioned. V2-Flash, V2-Pro, V2-Omni all migrated or retired. The V2.5 line (released April 23, 2026) is now the only production family. Weights remain on HuggingFace for self-hosting.
V2.5 series open-sourced. V2.5 standard, V2.5-Pro (1T MoE, 1M context), V2.5-Omni, V2.5-TTS, V2.5-ASR — all MIT-licensed. Pro public beta initially at ¥5/¥15 list per 1M during a 60-day window.
V2-Flash GA. 309B/15B MoE, 56k context, launched at $0.50/$1.50 list per 1M. Already the cheapest MoE model at its parameter class at the time.
MiMo-7B open-weights release. MIT licensed. Free to use, modify, and self-host. Zero per-token cost on any hardware that fits (INT4 = 3.5 GB VRAM).
5. Where to buy the MiMo API
① Xiaomi MiMo Platform (direct)
Official platform, V2.5 standard + V2.5-Pro on the public API. OpenAI-compatible endpoint. Best price, official SLA.
② OpenRouter
One API key, all MiMo variants plus 200+ other models. Best for prototyping — switch models without changing code. Slight markup over list.
③ HuggingFace Pro
Open weights, Inference Endpoints, Pro tier for SLA-backed inference. Best for self-hosting at scale.
④ Quickstart guide
5-line Python install + first call. curl / Node / Go examples included. →
6. Pricing FAQ
Why did MiMo cut prices so aggressively in May 2026?
Specifically, the cache-hit input price for V2.5-Pro dropped 99% (to $0.0035/M, ¥0.025/M) and the cache-hit input for V2.5 standard dropped 98% (to $0.0028/M, ¥0.02/M). List output also dropped 86% (V2.5-Pro to $0.83/M, ¥6) and 93% (V2.5 standard to $0.28/M, ¥2). Xiaomi cited engineering wins (SGLang HiCache + sliding-window attention reducing KV-cache movement to 1/7 of pre-optimization) plus ecosystem strategy. The 99% number is for cache-hit input — the headline rate that matters for prompt-cacheable agent workloads. List input ($0.42/M) is a separate, larger number.
What's the difference between list and cache-hit pricing?
List = tokens billed at the standard pay-as-you-go rate when there's no cached prefix match. Cache-hit = tokens billed when a request reuses a prefix that the server already has in cache (you enable this per-request). For most agent and chat workloads, the bulk of input tokens (system prompt, retrieved docs, conversation history) are cache-eligible, so the effective input cost is closer to the cache-hit rate than the list rate.
Are there volume discounts?
Not at the official list rate. Enterprise contracts (1B+ tokens/month) are negotiated directly with Xiaomi's sales team — typically 20–40% off list. OpenRouter also offers pay-as-you-go without volume tiers. The Token Plan subscription (¥X/month for a token bundle) was also revised in May 2026 to give 5–8× more tokens for the same price.
Is self-hosting cheaper than the API?
Depends on volume. Break-even for MiMo-7B (the smallest self-hostable model) is around 50M tokens/month on a single A10G. For V2.5 standard and V2.5-Pro, you'd need multi-GPU infrastructure — break-even is 500M+ tokens/month. See deployment guide →
Is V2.5-Omni / TTS / ASR pricing finalized?
ASR is at ¥0.5/hr of input audio on the public API. Omni and TTS API tiers are TBD as of July 2026 — weights are MIT-licensed (Omni) or commercially licensed separately (TTS). Sign up for platform.xiaomimimo.com to get notified when API tiers launch.
Related pages
- Token Cost Calculator — real-time USD/CNY estimation vs 5 competitors
- MiMo-V2.5 Series deep dive — architecture, benchmarks, deployment
- MiMo vs DeepSeek — most popular head-to-head
- MiMo vs Claude Haiku 3.5 — budget frontier model comparison
- Benchmark research — SWE-Bench, AIME, MMLU scores
- API quickstart — 5-line install + first call
Unofficial community resource. Not affiliated with Xiaomi Inc. Prices last verified July 20, 2026 against the official 2026-05-27 announcement and provider list pages. For mission-critical budgeting, always confirm with the provider's pricing page.