Last updated: 2026-07-20 · Report outdated data on GitHub

MiMo API Pricing — July 2026

Frontier-class reasoning at 5–20% of closed-source pricing. V2.5-Pro list input cut to $0.42/M in May 2026; cache-hit input cut 99% to $0.0035/M. All prices in USD per 1 million tokens.

Bottom line: V2.5-Pro list ~$0.42/$0.83, cache-hit input 99% off

MiMo-V2.5-Pro list pricing is now $0.42/M input + $0.83/M output (¥3 / ¥6 per 1M CNY). For workloads that can benefit from prompt caching — chat history, RAG, code review, long-context retrieval — the cache-hit input rate is $0.0035/M (¥0.025/M, 99% off from the original 2025 launch rate). The May 2026 cut was specifically for cache-hit pricing; list output dropped 86% (V2.5-Pro) and 93% (V2.5 standard). MiMo-V2.5-Pro matches Claude Opus 4.5 on long-context reasoning and is ~97% cheaper on list input ($0.42 vs $15) and ~99% cheaper on cache-hit input. MiMo-7B is free under MIT (zero per-token cost when self-hosted).

All MiMo weights are MIT-licensed. Self-hosting eliminates per-token cost entirely for teams with GPU infrastructure.

Heads up: MiMo-V2-Flash was deprecated on 2026-06-30 along with the rest of the V2 series. The current production lineup is V2.5 standard, V2.5-Pro, and V2.5-Pro-UltraSpeed. V2.5-Omni / V2.5-TTS / V2.5-ASR weights are MIT-licensed; API pricing is TBD.

→ Use the Token Cost Calculator (real-time USD/CNY)

1. The V2.5 lineup (full pricing)

All prices per 1 million tokens. List = pay-as-you-go rate; cache-hit = applies to repeated tokens when prompt caching is enabled. Verified July 20, 2026 against the official 2026-05-27 announcement.

Model Params Context Best for List input / 1M Cache-hit input / 1M List output / 1M Change since May 2026
MiMo-V2.5-Pro 1T+ MoE 1M Agent, long-context, reasoning $0.42 $0.0035 $0.83 Cache-hit −99% · Output −86%
MiMo-V2.5-Pro-UltraSpeed 1T+ MoE (FP4+DFlash) 1M High-throughput coding, 1000+ tok/s $1.25 $0.0104 $2.50 Internal-test only (apply for access)
MiMo-V2.5 (standard) ~309B MoE 1M General chat, content, lightweight agent $0.14 $0.0028 $0.28 Cache-hit −98% · Output −93%
MiMo-V2.5-Omni TBD TBD Multimodal (text / image / video / audio) TBD TBD TBD MIT weights out; API TBD
MiMo-V2.5-TTS TBD N/A Bilingual speech synthesis (CN/EN) Commercial license license Commercially licensed separately
MiMo-V2.5-ASR TBD N/A Speech-to-text (CN/EN) ¥0.5/hr (audio input) — Open weights (Apache-2.0)
MiMo-7B 7B dense 32k Edge / on-device (phones, cars, IoT) $0.00 $0.00 $0.00 MIT weights, free self-host
MiMo Code Agent layer Infinite Terminal-native coding agent Free app + API usage + API usage Open source, MIT
MiMo-V2-Flash (deprecated) 309B / 15B active 56k Replaced by V2.5 standard — — — Offline 2026-06-30

2. Side-by-side: same workload, 5 leading APIs

All prices list (USD per 1M tokens). Prompt caching and batch API rates are not included. Verified July 20, 2026 against provider list pages. Comparison baseline = MiMo-V2.5-Pro list input $0.42.

Provider Model List input / 1M List output / 1M vs MiMo-V2.5-Pro list input
Xiaomi MiMo-V2.5-Pro $0.42 $0.83 — baseline —
Xiaomi MiMo-V2.5 (standard) $0.14 $0.28 ~67% cheaper input
DeepSeek DeepSeek V4 Pro $0.435 $0.87 ~3% more expensive input (similar tier)
DeepSeek DeepSeek-R1 $0.55 $2.19 ~31% more expensive input
Alibaba Qwen3 Max $0.60 $2.40 ~43% more expensive input
OpenAI GPT-4o mini $0.15 $0.60 ~64% cheaper input (smaller model)
OpenAI GPT-5 $1.25 $10.00 ~3× more expensive input
Anthropic Claude Haiku 3.5 $0.80 $4.00 ~90% more expensive input (smaller model)
Anthropic Claude Opus 4.5 $15.00 $75.00 ~36× more expensive input

Reading the table: MiMo-V2.5-Pro at $0.42/$0.83 list matches Claude Opus 4.5 on long-context reasoning, beats Claude Haiku 3.5 on coding (73.4% vs 65% on SWE-Bench Verified), and is priced just below DeepSeek V4 Pro on input. The 1M context window (vs 128k for V4 Pro) is the deal-breaker for legal / scientific / code-review workloads. For prompt-cacheable workloads (chat history, RAG, repeated system prompts), the cache-hit input rate ($0.0035/M) makes MiMo-V2.5-Pro effectively free on input.

3. Why cache-hit pricing matters

Most real-world LLM applications send the same tokens multiple times: system prompts, retrieved documents in RAG, prior conversation history, code files under review. MiMo (and most modern APIs) offers prompt caching: a flag at request time, and the repeated prefix is billed at the much lower cache-hit rate. On MiMo-V2.5-Pro:

This is the reason the May 2026 cut was framed as "99% off" — for the kind of repeated-token workloads that dominate production agent use, the effective per-request cost is genuinely near zero on the input side. See the calculator for your workload →

4. Pricing history (timeline)

May 27, 2026

Permanent cache-hit price cut (−98% to −99%). V2.5-Pro cache-hit input drops to ¥0.025/M ($0.0035), V2.5 standard cache-hit input to ¥0.02/M ($0.0028). List output also cut: V2.5-Pro ¥6/M (−86% from launch), V2.5 standard ¥2/M (−93%). Xiaomi cites "ecosystem expansion over short-term margin" as the rationale. Backed by SGLang HiCache + sliding-window attention reducing KV-cache movement to 1/7 of pre-optimization cost.

June 30, 2026

MiMo-V2 series decommissioned. V2-Flash, V2-Pro, V2-Omni all migrated or retired. The V2.5 line (released April 23, 2026) is now the only production family. Weights remain on HuggingFace for self-hosting.

April 23, 2026

V2.5 series open-sourced. V2.5 standard, V2.5-Pro (1T MoE, 1M context), V2.5-Omni, V2.5-TTS, V2.5-ASR — all MIT-licensed. Pro public beta initially at ¥5/¥15 list per 1M during a 60-day window.

December 2025

V2-Flash GA. 309B/15B MoE, 56k context, launched at $0.50/$1.50 list per 1M. Already the cheapest MoE model at its parameter class at the time.

June 2025

MiMo-7B open-weights release. MIT licensed. Free to use, modify, and self-host. Zero per-token cost on any hardware that fits (INT4 = 3.5 GB VRAM).

5. Where to buy the MiMo API

6. Pricing FAQ

Why did MiMo cut prices so aggressively in May 2026?

Specifically, the cache-hit input price for V2.5-Pro dropped 99% (to $0.0035/M, ¥0.025/M) and the cache-hit input for V2.5 standard dropped 98% (to $0.0028/M, ¥0.02/M). List output also dropped 86% (V2.5-Pro to $0.83/M, ¥6) and 93% (V2.5 standard to $0.28/M, ¥2). Xiaomi cited engineering wins (SGLang HiCache + sliding-window attention reducing KV-cache movement to 1/7 of pre-optimization) plus ecosystem strategy. The 99% number is for cache-hit input — the headline rate that matters for prompt-cacheable agent workloads. List input ($0.42/M) is a separate, larger number.

What's the difference between list and cache-hit pricing?

List = tokens billed at the standard pay-as-you-go rate when there's no cached prefix match. Cache-hit = tokens billed when a request reuses a prefix that the server already has in cache (you enable this per-request). For most agent and chat workloads, the bulk of input tokens (system prompt, retrieved docs, conversation history) are cache-eligible, so the effective input cost is closer to the cache-hit rate than the list rate.

Are there volume discounts?

Not at the official list rate. Enterprise contracts (1B+ tokens/month) are negotiated directly with Xiaomi's sales team — typically 20–40% off list. OpenRouter also offers pay-as-you-go without volume tiers. The Token Plan subscription (¥X/month for a token bundle) was also revised in May 2026 to give 5–8× more tokens for the same price.

Is self-hosting cheaper than the API?

Depends on volume. Break-even for MiMo-7B (the smallest self-hostable model) is around 50M tokens/month on a single A10G. For V2.5 standard and V2.5-Pro, you'd need multi-GPU infrastructure — break-even is 500M+ tokens/month. See deployment guide →

Is V2.5-Omni / TTS / ASR pricing finalized?

ASR is at ¥0.5/hr of input audio on the public API. Omni and TTS API tiers are TBD as of July 2026 — weights are MIT-licensed (Omni) or commercially licensed separately (TTS). Sign up for platform.xiaomimimo.com to get notified when API tiers launch.

Related pages

Unofficial community resource. Not affiliated with Xiaomi Inc. Prices last verified July 20, 2026 against the official 2026-05-27 announcement and provider list pages. For mission-critical budgeting, always confirm with the provider's pricing page.