This is the English edition. 한국어판 and 日本語版 are also available.

Xiaomi MiMo-V2.6 Explained: 1M Context, Native Multimodality and Large-Scale Agentic RL

2026-09-22 · AI · United States · Zoogom Editorial

#Xiaomi#MiMo-V2.6#Open-weight AI#Multimodal AI#AI agents#Reinforcement learning

A modular AI core connecting multiple input and tool pathways in an independent visualization of MiMo-V2.6

Xiaomi officially released the MiMo-V2.6 family on September 22, 2026. The launch includes MiMo-V2.6-Pro for maximum capability, MiMo-V2.6-Flash for a lower-cost balance of speed and intelligence, and a 9B distilled checkpoint intended as a starting point for agentic reinforcement-learning research. Xiaomi also published model weights, a technical report and supporting RL resources.

The headline is not just scale. Xiaomi trained one model family to process text, images, video and audio while working across coding, web, computer-use and cybersecurity environments. It supports a one-million-token context window, and Xiaomi exposed part of the training process through an official live RL dashboard that recorded steps, tokens, cost, restarts and changing evaluation results.

The practical conclusion is that MiMo-V2.6 is a serious open-weight agent-model candidate with unusually aggressive API pricing. It is not yet a settled winner. Most launch-day benchmark results were run or reported by Xiaomi, a 1M-token limit does not guarantee reliable retrieval across every prompt, and the original checkpoints still demand datacenter-class deployment resources.

The three-point summary

  1. Pro is a sparse MoE with 1.02T total and 42B activated parameters; Flash has 309B total and 15B activated parameters.
  2. Both support 1M tokens of context, up to 128K output tokens, native multimodal input and tool use.
  3. Hosted prices are exceptionally low, but self-hosting the full models requires extensive multi-GPU infrastructure and the performance claims still need independent replication.

A model-selection guide comparing MiMo-V2.6 Pro, Flash and the 9B research checkpoint

MiMo-V2.6 is a family, not one model

Pro is the flagship for complex projects and long-horizon agent work. Xiaomi’s model card describes a 1.02T-parameter network that activates 42B parameters for each token. Flash uses 309B total and 15B activated parameters while retaining the same broad input categories and long-context limit at a lower API price.

MiMo-V2.6-Distill-Qwen-9B is different. It is not simply a consumer-sized copy of Pro. Xiaomi says it fine-tuned Qwen3.5-9B on MiMo-generated data covering code, general agents, visual coding and cybersecurity, then released the SFT checkpoint as an open starting point for agentic RL research.

The practical split is straightforward:

How a 1.02T model activates 42B parameters

Pro and Flash use a sparse mixture-of-experts architecture. The model does not calculate every parameter for every token. A router selects a small subset of experts: eight of 384 routed experts in Pro and eight of 256 in Flash.

This reduces active computation, but it does not make Pro equivalent to a dense 42B model on ordinary hardware. Active compute and the memory required to store and distribute all model weights are different constraints. Xiaomi’s own SGLang example uses tensor parallelism of 16, data parallelism of two and two nodes for Pro. Its Flash example uses tensor parallelism of eight and data parallelism of two. These are not reference setups for a typical desktop PC.

What the one-million-token context window means

Both hosted models advertise a 1M-token context window and up to 128K output tokens. That creates room for large repositories, multi-session tool traces, document collections and tokens derived from video or audio.

But supports 1M is not the same as remains equally accurate and economical at 1M. Long requests introduce several tradeoffs:

Production systems still benefit from retrieval, summaries, checkpoints and context caching. A larger limit expands the operating range; it does not eliminate context engineering.

Native multimodality and computer use

Xiaomi’s model cards say Pro and Flash accept text, image, video and audio in one model. They use a 681M-parameter MiMo vision encoder, a 308M AudioTokenizer and a 127M audio-patch encoder. The hosted models also list deep thinking, tool calls, streaming, web search, structured output and context caching.

Xiaomi’s launch examples go well beyond chat:

These demonstrations show the intended capability envelope, not guaranteed performance in every desktop environment. Computer use also depends on the capture layer, action tools, application permissions, retry logic and safeguards around the model.

Xiaomi’s published live reinforcement-learning scale and model-card agent evaluations

What Xiaomi’s official live RL dashboard reveals

One of the most notable parts of the release is the official training dashboard. It exposed more than a polished final benchmark table. As of September 22, both Pro and Flash were marked ended after 30 steps.

The dashboard records:

The operator notices make the record more useful. Xiaomi disclosed a Pro GPU out-of-memory failure caused by expert-load imbalance, a network problem between the training cluster and grader deployment, and a Flash restart after an infrastructure error on one dataset went undetected. A notice also said the team removed a cyber dataset from a subsequent Pro run after observing problematic rollout patterns.

That transparency shows agentic RL as an engineering process involving data, distributed systems, graders and reward design—not a single clean training command. The dashboard still represents Xiaomi’s own logs rather than an external audit.

What “self-improvement” does and does not mean

Xiaomi describes the release as reinforcement learning toward self-improvement. This does not mean that a deployed model autonomously rewrites itself without control. The published method is closer to the following loop:

  1. Generate multiple solution trajectories in coding, general, visual and cybersecurity environments.
  2. Compare successful trajectories within a group instead of using only binary pass or fail.
  3. Shift reward toward higher-quality, shorter and more efficient successful paths.
  4. Use adversarial evaluation, anomaly detection and verifier cross-checks to resist reward hacking.
  5. Mix multiple agent harnesses so strategies can transfer to harnesses not seen during training.

The important innovation is a richer feedback and grading loop during controlled training. It should not be interpreted as a production service changing its weights secretly after deployment.

Strong benchmark numbers—with an important label

Xiaomi’s model card reports competitive agent results. Pro and Flash score 71.9 and 67.9 on DeepSWE v1.1, 76.9 and 73.6 on Toolathlon-Verified, 82.0 and 80.8 on OSWorld-Verified, and 72.3 and 71.5 on MiMo VisualCoding.

Flash staying close to Pro across those tasks is particularly notable. Flash even leads the reported CyberGym result, 95.1 to Pro’s 94.0. The figures still need careful framing:

The responsible reading is that MiMo-V2.6 presents highly competitive initial results for an open-weight family—not that it has conclusively beaten every closed model.

U.S. API pricing

For overseas pay-as-you-go API use, Xiaomi lists these prices per million tokens:

One million uncached input tokens plus 100,000 output tokens would therefore cost about $0.522 on Pro or $0.168 on Flash before add-ons. Web search is billed separately at $5 per 1,000 search-tool calls overseas, and retrieved pages also increase input-token usage. Xiaomi’s Batch API discounts Pro and Flash real-time rates by 50%.

UltraSpeed targets up to 20 times faster Pro output, but its token price is also ten times the regular Pro rate. It is a latency-sensitive production option rather than a universally better tier.

How U.S. users can access it

Xiaomi lists four main paths:

  1. Generate an API key on the Xiaomi MiMo Open Platform and use the OpenAI- or Anthropic-compatible protocol.
  2. Use MiMo Desktop with a membership or personal API key.
  3. Use official tools such as MiMo Code, MiMo Studio or MiMo Claw.
  4. Download the Hugging Face weights and serve them with SGLang or vLLM.

The hosted API IDs are lowercase: mimo-v2.6-pro, mimo-v2.6-flash and mimo-v2.6-pro-ultraspeed.

OpenAI protocol compatibility means an application can often migrate by changing its base URL and model name. It does not mean MiMo usage is included with a ChatGPT or OpenAI subscription, nor does it make MiMo a native option in OpenAI’s apps. Clients should separately test reasoning fields, multimodal upload formats, streaming events and multi-turn tool-call history.

Is local deployment realistic?

The full Pro and Flash checkpoints are open, but they are not consumer-class models. Xiaomi’s reference commands already assume substantial GPU parallelism. Quantization can reduce storage and memory demand, but long-context KV cache, multimodal encoders and throughput requirements remain significant.

The 9B distilled checkpoint is a more practical starting point for individual developers. Its model page links to community quantizations for llama.cpp, Ollama and LM Studio. Those builds may differ from Xiaomi’s hosted models in output quality, maximum context and tool-call behavior.

What the MIT license covers—and what it does not

The Hugging Face pages for Pro, Flash and Distill 9B display the MIT license. That is a permissive foundation for research, modification and commercial use. It is not a universal waiver for every layer of a product.

Commercial teams should re-check each repository’s license, notices, dependencies and service terms at the exact revision they deploy.

Privacy and enterprise-data questions

Before sending source code, customer files, employee records or unreleased research to a hosted endpoint, confirm retention, training-use policy, processing location, deletion controls, audit support and contract terms. A model card and price page do not settle those questions by themselves.

Self-hosting can reduce external data transfer, but it shifts responsibility to the operator. Access controls, secret filtering, log redaction, least-privilege tools, output validation and incident response remain essential. Computer-use and cybersecurity capabilities make narrow action permissions more important, not less.

Who should consider MiMo-V2.6

Flash is the most compelling option for teams that need high-volume coding agents, document processing or repeated tool calls under a tight cost ceiling. Pro is better aligned with difficult, long-horizon projects that combine visual material, tools and research. Organizations with GPU clusters or a need to study model internals also gain unusual flexibility from the open weights.

Teams in heavily regulated environments, teams that require independently validated performance, or users who need the full original model on a single personal computer have good reasons to wait. U.S. businesses should also verify contracting, privacy and support terms before routing sensitive production data to the hosted service.

Frequently asked questions

Is MiMo-V2.6 open source?

Xiaomi released Pro and Flash weights, a technical report, Distill 9B and supporting RL resources, and the model pages display the MIT license. Hosted services and third-party dependencies can have separate terms.

Should I use Pro or Flash?

Use Pro when maximum capability on difficult, long tasks is the priority. Use Flash when call volume, speed-cost balance and iteration are more important. Evaluate both on your own tasks with the same tool and token budgets.

Should I fill the entire 1M-token window?

Usually not. The limit is capacity, not a recommendation. Retrieval, summaries and caching are often faster, cheaper and easier to debug.

Can I run Pro or Flash on a Mac mini?

The original checkpoints’ official deployment configurations assume multi-GPU parallelism. A suitable quantization of the Distill 9B checkpoint is a more realistic personal-computer experiment.

How strong is English performance?

English and Chinese are listed on the model cards, and many published evaluations are English-centric. Real product quality still depends on domain language, tool integrations and long-session reliability, so a task-specific evaluation is necessary.

Is the live dashboard still training the models?

No. At the September 22 check, both runs had completed 30 steps and were marked ended. The site remains useful as a public record of metrics, evaluation changes and operator notices.

Bottom line

MiMo-V2.6 combines a very large sparse MoE, 1M context and native multimodal input with aggressive hosted pricing and open weights. Flash’s proximity to Pro on Xiaomi’s initial agent benchmarks is notable, as is the decision to expose cost, restarts and evaluation changes through an official training dashboard.

It is still too early to crown a winner from launch-day numbers. Independent replication, long-context accuracy, U.S. privacy and contracting terms, tool-call reliability and real workload latency need testing. The best next step is a small, representative evaluation of Pro and Flash that measures quality, total cost and recovery from tool failures before any broad rollout.

Trademarks and image notice

Xiaomi, MiMo, Qwen and related marks belong to their respective owners. This article is not sponsored, endorsed or approved by Xiaomi. The article images do not reproduce product screens, trademark logos, official launch artwork or third-party photographs.

Official sources

Source: Xiaomi MiMo · Includes original screenshots or graphics