Xiaomi MiMo-V2.6 Explained: 1M Context, Native Multimodality and Large-Scale Agentic RL

Xiaomi officially released the MiMo-V2.6 family on September 22, 2026. The launch includes MiMo-V2.6-Pro for maximum capability, MiMo-V2.6-Flash for a lower-cost balance of speed and intelligence, and a 9B distilled checkpoint intended as a starting point for agentic reinforcement-learning research. Xiaomi also published model weights, a technical report and supporting RL resources.
The headline is not just scale. Xiaomi trained one model family to process text, images, video and audio while working across coding, web, computer-use and cybersecurity environments. It supports a one-million-token context window, and Xiaomi exposed part of the training process through an official live RL dashboard that recorded steps, tokens, cost, restarts and changing evaluation results.
The practical conclusion is that MiMo-V2.6 is a serious open-weight agent-model candidate with unusually aggressive API pricing. It is not yet a settled winner. Most launch-day benchmark results were run or reported by Xiaomi, a 1M-token limit does not guarantee reliable retrieval across every prompt, and the original checkpoints still demand datacenter-class deployment resources.
The three-point summary
- Pro is a sparse MoE with 1.02T total and 42B activated parameters; Flash has 309B total and 15B activated parameters.
- Both support 1M tokens of context, up to 128K output tokens, native multimodal input and tool use.
- Hosted prices are exceptionally low, but self-hosting the full models requires extensive multi-GPU infrastructure and the performance claims still need independent replication.

MiMo-V2.6 is a family, not one model
Pro is the flagship for complex projects and long-horizon agent work. Xiaomi’s model card describes a 1.02T-parameter network that activates 42B parameters for each token. Flash uses 309B total and 15B activated parameters while retaining the same broad input categories and long-context limit at a lower API price.
MiMo-V2.6-Distill-Qwen-9B is different. It is not simply a consumer-sized copy of Pro. Xiaomi says it fine-tuned Qwen3.5-9B on MiMo-generated data covering code, general agents, visual coding and cybersecurity, then released the SFT checkpoint as an open starting point for agentic RL research.
The practical split is straightforward:
- Choose Pro when maximum capability on difficult, long tasks matters most.
- Choose Flash when throughput, iteration volume and predictable cost matter most.
- Choose Distill 9B when the goal is open experimentation or smaller-scale local research.
How a 1.02T model activates 42B parameters
Pro and Flash use a sparse mixture-of-experts architecture. The model does not calculate every parameter for every token. A router selects a small subset of experts: eight of 384 routed experts in Pro and eight of 256 in Flash.
This reduces active computation, but it does not make Pro equivalent to a dense 42B model on ordinary hardware. Active compute and the memory required to store and distribute all model weights are different constraints. Xiaomi’s own SGLang example uses tensor parallelism of 16, data parallelism of two and two nodes for Pro. Its Flash example uses tensor parallelism of eight and data parallelism of two. These are not reference setups for a typical desktop PC.
What the one-million-token context window means
Both hosted models advertise a 1M-token context window and up to 128K output tokens. That creates room for large repositories, multi-session tool traces, document collections and tokens derived from video or audio.
But supports 1M is not the same as remains equally accurate and economical at 1M. Long requests introduce several tradeoffs:
- Uncached input increases the bill.
- More material can increase retrieval difficulty and latency.
- Old and new instructions are more likely to conflict.
- Multimodal inputs consume tokens differently from plain text.
- Agents can accumulate irrelevant tool output over a long run.
Production systems still benefit from retrieval, summaries, checkpoints and context caching. A larger limit expands the operating range; it does not eliminate context engineering.
Native multimodality and computer use
Xiaomi’s model cards say Pro and Flash accept text, image, video and audio in one model. They use a 681M-parameter MiMo vision encoder, a 308M AudioTokenizer and a 127M audio-patch encoder. The hosted models also list deep thinking, tool calls, streaming, web search, structured output and context caching.
Xiaomi’s launch examples go well beyond chat:
- Building a 3D game scene and interaction logic from text, images or video
- Creating objects and scenes in Blender
- Operating office and productivity interfaces using visual feedback
- Checking results and revising later actions
- Retrieving literature and running tools for materials screening
- Orchestrating Figma, front-end, slide, video and music workflows
These demonstrations show the intended capability envelope, not guaranteed performance in every desktop environment. Computer use also depends on the capture layer, action tools, application permissions, retry logic and safeguards around the model.

What Xiaomi’s official live RL dashboard reveals
One of the most notable parts of the release is the official training dashboard. It exposed more than a polished final benchmark table. As of September 22, both Pro and Flash were marked ended after 30 steps.
The dashboard records:
- 1,568 prompts and 16 rollouts per prompt, or 25,088 trajectories per step for each model
- 752,640 cumulative trajectories per model after 30 steps
- About 75.0 billion cumulative training tokens for Pro and 81.4 billion for Flash
- Final estimated costs of about $2.62 million for Pro and $854,000 for Flash
- Fourteen recorded Pro restarts and five Flash restarts
The operator notices make the record more useful. Xiaomi disclosed a Pro GPU out-of-memory failure caused by expert-load imbalance, a network problem between the training cluster and grader deployment, and a Flash restart after an infrastructure error on one dataset went undetected. A notice also said the team removed a cyber dataset from a subsequent Pro run after observing problematic rollout patterns.
That transparency shows agentic RL as an engineering process involving data, distributed systems, graders and reward design—not a single clean training command. The dashboard still represents Xiaomi’s own logs rather than an external audit.
What “self-improvement” does and does not mean
Xiaomi describes the release as reinforcement learning toward self-improvement. This does not mean that a deployed model autonomously rewrites itself without control. The published method is closer to the following loop:
- Generate multiple solution trajectories in coding, general, visual and cybersecurity environments.
- Compare successful trajectories within a group instead of using only binary pass or fail.
- Shift reward toward higher-quality, shorter and more efficient successful paths.
- Use adversarial evaluation, anomaly detection and verifier cross-checks to resist reward hacking.
- Mix multiple agent harnesses so strategies can transfer to harnesses not seen during training.
The important innovation is a richer feedback and grading loop during controlled training. It should not be interpreted as a production service changing its weights secretly after deployment.
Strong benchmark numbers—with an important label
Xiaomi’s model card reports competitive agent results. Pro and Flash score 71.9 and 67.9 on DeepSWE v1.1, 76.9 and 73.6 on Toolathlon-Verified, 82.0 and 80.8 on OSWorld-Verified, and 72.3 and 71.5 on MiMo VisualCoding.
Flash staying close to Pro across those tasks is particularly notable. Flash even leads the reported CyberGym result, 95.1 to Pro’s 94.0. The figures still need careful framing:
- Xiaomi selected and ran the launch configurations.
- Some evaluations, including MiMo Code Bench and MiMo Cyber Bench, are internal.
- Model comparisons may not use identical latency, token and tool budgets.
- Launch day does not provide enough time for broad third-party replication or long-running reliability studies.
The responsible reading is that MiMo-V2.6 presents highly competitive initial results for an open-weight family—not that it has conclusively beaten every closed model.
U.S. API pricing
For overseas pay-as-you-go API use, Xiaomi lists these prices per million tokens:
- Pro: $0.0036 cached input, $0.435 uncached input and $0.87 output
- Flash: $0.0028 cached input, $0.14 uncached input and $0.28 output
- Pro UltraSpeed: $0.036 cached input, $4.35 uncached input and $8.70 output
One million uncached input tokens plus 100,000 output tokens would therefore cost about $0.522 on Pro or $0.168 on Flash before add-ons. Web search is billed separately at $5 per 1,000 search-tool calls overseas, and retrieved pages also increase input-token usage. Xiaomi’s Batch API discounts Pro and Flash real-time rates by 50%.
UltraSpeed targets up to 20 times faster Pro output, but its token price is also ten times the regular Pro rate. It is a latency-sensitive production option rather than a universally better tier.
How U.S. users can access it
Xiaomi lists four main paths:
- Generate an API key on the Xiaomi MiMo Open Platform and use the OpenAI- or Anthropic-compatible protocol.
- Use MiMo Desktop with a membership or personal API key.
- Use official tools such as MiMo Code, MiMo Studio or MiMo Claw.
- Download the Hugging Face weights and serve them with SGLang or vLLM.
The hosted API IDs are lowercase: mimo-v2.6-pro, mimo-v2.6-flash and mimo-v2.6-pro-ultraspeed.
OpenAI protocol compatibility means an application can often migrate by changing its base URL and model name. It does not mean MiMo usage is included with a ChatGPT or OpenAI subscription, nor does it make MiMo a native option in OpenAI’s apps. Clients should separately test reasoning fields, multimodal upload formats, streaming events and multi-turn tool-call history.
Is local deployment realistic?
The full Pro and Flash checkpoints are open, but they are not consumer-class models. Xiaomi’s reference commands already assume substantial GPU parallelism. Quantization can reduce storage and memory demand, but long-context KV cache, multimodal encoders and throughput requirements remain significant.
The 9B distilled checkpoint is a more practical starting point for individual developers. Its model page links to community quantizations for llama.cpp, Ollama and LM Studio. Those builds may differ from Xiaomi’s hosted models in output quality, maximum context and tool-call behavior.
What the MIT license covers—and what it does not
The Hugging Face pages for Pro, Flash and Distill 9B display the MIT license. That is a permissive foundation for research, modification and commercial use. It is not a universal waiver for every layer of a product.
- Review terms for base models and third-party components.
- Generated outputs can still create copyright, publicity and trademark risk.
- Hosted API use is governed by Xiaomi’s service terms and privacy policy, not only the weight license.
- Cybersecurity automation should be limited to systems you are authorized to test.
- High-impact decisions should not be automated from model output alone.
Commercial teams should re-check each repository’s license, notices, dependencies and service terms at the exact revision they deploy.
Privacy and enterprise-data questions
Before sending source code, customer files, employee records or unreleased research to a hosted endpoint, confirm retention, training-use policy, processing location, deletion controls, audit support and contract terms. A model card and price page do not settle those questions by themselves.
Self-hosting can reduce external data transfer, but it shifts responsibility to the operator. Access controls, secret filtering, log redaction, least-privilege tools, output validation and incident response remain essential. Computer-use and cybersecurity capabilities make narrow action permissions more important, not less.
Who should consider MiMo-V2.6
Flash is the most compelling option for teams that need high-volume coding agents, document processing or repeated tool calls under a tight cost ceiling. Pro is better aligned with difficult, long-horizon projects that combine visual material, tools and research. Organizations with GPU clusters or a need to study model internals also gain unusual flexibility from the open weights.
Teams in heavily regulated environments, teams that require independently validated performance, or users who need the full original model on a single personal computer have good reasons to wait. U.S. businesses should also verify contracting, privacy and support terms before routing sensitive production data to the hosted service.
Frequently asked questions
Is MiMo-V2.6 open source?
Xiaomi released Pro and Flash weights, a technical report, Distill 9B and supporting RL resources, and the model pages display the MIT license. Hosted services and third-party dependencies can have separate terms.
Should I use Pro or Flash?
Use Pro when maximum capability on difficult, long tasks is the priority. Use Flash when call volume, speed-cost balance and iteration are more important. Evaluate both on your own tasks with the same tool and token budgets.
Should I fill the entire 1M-token window?
Usually not. The limit is capacity, not a recommendation. Retrieval, summaries and caching are often faster, cheaper and easier to debug.
Can I run Pro or Flash on a Mac mini?
The original checkpoints’ official deployment configurations assume multi-GPU parallelism. A suitable quantization of the Distill 9B checkpoint is a more realistic personal-computer experiment.
How strong is English performance?
English and Chinese are listed on the model cards, and many published evaluations are English-centric. Real product quality still depends on domain language, tool integrations and long-session reliability, so a task-specific evaluation is necessary.
Is the live dashboard still training the models?
No. At the September 22 check, both runs had completed 30 steps and were marked ended. The site remains useful as a public record of metrics, evaluation changes and operator notices.
Bottom line
MiMo-V2.6 combines a very large sparse MoE, 1M context and native multimodal input with aggressive hosted pricing and open weights. Flash’s proximity to Pro on Xiaomi’s initial agent benchmarks is notable, as is the decision to expose cost, restarts and evaluation changes through an official training dashboard.
It is still too early to crown a winner from launch-day numbers. Independent replication, long-context accuracy, U.S. privacy and contracting terms, tool-call reliability and real workload latency need testing. The best next step is a small, representative evaluation of Pro and Flash that measures quality, total cost and recovery from tool failures before any broad rollout.
Trademarks and image notice
Xiaomi, MiMo, Qwen and related marks belong to their respective owners. This article is not sponsored, endorsed or approved by Xiaomi. The article images do not reproduce product screens, trademark logos, official launch artwork or third-party photographs.
Official sources
- Xiaomi MiMo: Official MiMo-V2.6 announcement
- Xiaomi MiMo: Official live RL dashboard
- Xiaomi MiMo: MiMo-V2.6-Pro model and pricing
- Xiaomi MiMo: MiMo-V2.6-Flash model and pricing
- Hugging Face: MiMo-V2.6-Pro-RL model card
- Hugging Face: MiMo-V2.6-Flash-RL model card
- Hugging Face: MiMo-V2.6-Distill-Qwen-9B
- Xiaomi MiMo: Batch API pricing
- Xiaomi MiMo: Web-search support and pricing



