In a post on X, @XiaomiMiMo announces Xiaomi MiMo-V2.6 Pro and Flash as two omnimodal models developed with scaled reinforcement learning. Xiaomi describes them as models built to work across multiple modalities and reports stronger coding, computer-use, 3D-reasoning and creative capabilities.

The announcement says the open release includes model weights, a technical report, reinforcement-learning environments and training code. It does not specify the model sizes, licences, access instructions, deployment requirements or benchmark methodology.

What Xiaomi announced

MiMo-V2.6 has two announced variants: Pro and Flash. The post says Pro performs on par with Claude Opus 5 and GPT-5.6 Sol across most agent benchmarks. That comparison is Xiaomi’s characterization of the results, and the announcement does not define which benchmarks are included in “most” or what “on par” means.

Xiaomi also reports a score of 46 on the Artificial Analysis Intelligence Index and describes it as the highest score among open-source models. The announcement does not specify the index’s methodology, comparison date or underlying evaluation conditions, so this is a reported ranking rather than an independently verified result.

The term omnimodal generally describes a model designed to handle several kinds of input or output, such as text, images or other media. The post does not provide a separate technical breakdown of which modalities either model supports.

Reported benchmark results

The comparison published with the announcement lists scores for MiMo-V2.6 Pro, MiMo-V2.6 Flash, the previous-generation MiMo-V2.5 Pro and several named comparison models. The figures below reproduce that comparison as reported by Xiaomi. A dash indicates that the image does not show a value.

Area

Benchmark

MiMo-V2.6 Pro

MiMo-V2.6 Flash

MiMo-V2.5 Pro

Claude Opus 5

GPT-5.6 Sol

Fable 5

Code agent

DeepSWE v1.1

71.9

67.9

19.0

74.0

73.0

70.0

Code agent

ProgramBench

26.5

26.0

12.5

37.0

25.0

33.0

Code agent

MiMo Code Bench

63.2

61.2

40.4

68.6

59.3

General agent

AutomationBench v1.0.0.6

53.1

52.3

16.0

50.3

45.8

46.2

General agent

Toolathlon-Verified

76.9

73.6

49.1

80.6

74.9

77.9

General agent

GDPval-AA 2.1

1673

1107

1708

1588

1595

General agent

Agents’ Last Exam

31.6

27.6

13.2

31.6

30.8

25.7

General agent

Terminal Bench 4.0

34.9

28.8

1.5

49.0

39.9

42.4

General agent

Terminal Bench 2.1

89.9

87.6

65.2

89.1

88.8

84.3

General agent

OSWorld-Verified

82.0

80.8

61.5†

83.4

83.0

86.0

General agent

JobBench

62.0

61.2

25.0

65.7

45.4

57.4

Cybersecurity

CyberGym

94.0

95.1

40.0

Cybersecurity

MiMo Cyber Bench

80.2

77.2

0.0

Cybersecurity

ExploitGym

17.8

6.0

0.2

22.1

30.3

28.4

Cybersecurity

ExploitBench

47.9

25.3

16.6

70.0

78.5

78.0

Cybersecurity

SEC Bench Pro

66.3

47.5

17.7

79.1

Visual agent

MiMo Visual Coding

72.3

71.5

70.0

73.4

69.1

† The source image says this OSWorld-Verified result is from MiMo-V2.5. The image does not explain how that qualification affects comparison with the other entries.

Coding results show improvement over MiMo-V2.5

On all three listed code-agent benchmarks, both MiMo-V2.6 variants are above MiMo-V2.5 Pro in Xiaomi’s comparison. Pro remains below Claude Opus 5 on each benchmark, while its results are closer to GPT-5.6 Sol: Pro is below GPT-5.6 Sol on DeepSWE v1.1, but above it on ProgramBench and MiMo Code Bench. Flash is slightly below Pro on all three rows.

The announcement does not provide the prompting, tool, sampling or evaluation details needed to determine whether the model comparisons were conducted under matching conditions.

General-agent results are mixed by task

The general-agent rows show improvement over MiMo-V2.5 Pro for both new variants on the listed benchmarks. Pro is ahead of Flash on AutomationBench, Toolathlon-Verified, Agents’ Last Exam, both Terminal Bench versions, OSWorld-Verified and JobBench. Flash has no reported GDPval-AA 2.1 value in the image, so that result cannot be compared for the two variants.

The comparison with the other named models varies by benchmark. MiMo-V2.6 Pro is listed above Claude Opus 5 on AutomationBench and Terminal Bench 2.1, while Claude Opus 5 is higher on Toolathlon-Verified, GDPval-AA 2.1, Terminal Bench 4.0, OSWorld-Verified and JobBench. Pro and Claude Opus 5 have the same reported score on Agents’ Last Exam.

The image also shows MiMo-V2.6 Pro below Fable 5 on OSWorld-Verified and Terminal Bench 4.0, but above Fable 5 on AutomationBench, Agents’ Last Exam and Terminal Bench 2.1.

These differences matter because “agent” benchmarks cover different tasks. The announcement does not establish a single overall ranking across computer use, tool use, terminal work and job-related evaluation.

Cybersecurity results do not point in one direction

Xiaomi reports large improvements over MiMo-V2.5 Pro on the listed cybersecurity benchmarks. Flash is shown at 95.1 on CyberGym, slightly above Pro’s 94.0, while Pro is higher on MiMo Cyber Bench, ExploitGym, ExploitBench and SEC Bench Pro.

The comparison models are only present on some cybersecurity rows. On ExploitGym, both MiMo-V2.6 scores are below Claude Opus 5, GPT-5.6 Sol and Fable 5. On ExploitBench, Pro and Flash are also below all three named comparison models shown. SEC Bench Pro lists a score for GPT-5.6 Sol but not for Claude Opus 5 or Fable 5. The announcement does not explain the different comparison coverage or evaluation conditions.

Cybersecurity benchmark scores should not by themselves be read as proof of real-world security capability. The announcement does not specify the test setup, task limitations, safety controls or methodology behind these figures.

Visual coding is reported separately

The visual-agent section contains one benchmark, MiMo Visual Coding. Xiaomi lists MiMo-V2.6 Pro at 72.3 and Flash at 71.5. The image does not show a MiMo-V2.5 Pro result for this row. Pro is listed below GPT-5.6 Sol but above Claude Opus 5 and Fable 5; Flash is above Claude Opus 5 and Fable 5 but below GPT-5.6 Sol.

This result is narrower than a general claim about visual reasoning or multimodal performance. The supplied announcement does not include examples of the tasks, the supported inputs or outputs, or a separate measurement for the broader 3D-reasoning and creative capabilities it mentions.

What Xiaomi says the open release includes

Xiaomi says the open release includes four categories of material:

  • open model weights;

  • a technical report;

  • reinforcement-learning environments; and

  • training code.

The announcement does not specify the licence, release timing, download location, model sizes, hardware requirements, API availability or whether every listed component is available under the same terms. “Open model weights” therefore describes what Xiaomi says is included without establishing a particular licence or access model.

What the results establish—and what they do not

The reported comparison shows MiMo-V2.6 Pro and Flash scoring above MiMo-V2.5 Pro on most of the rows where both generations are listed. It also shows a mixed relationship with the named frontier models: the newer Xiaomi models are competitive on some rows and behind on others. Pro generally has higher reported scores than Flash in the listed code and general-agent rows, while Flash is higher on CyberGym.

Those observations describe Xiaomi’s published comparison, not an independent benchmark study. The announcement does not specify evaluation prompts, tools, model settings, sample sizes, uncertainty, hardware, dates or whether every model was tested under matching conditions.

For now, MiMo-V2.6 is best understood as an announcement of two Xiaomi models and an accompanying open-release package, with company-reported comparisons across coding, agent, cybersecurity and visual-coding tasks. The licence, release timing, download or API access, model sizes, hardware requirements and benchmark methodology remain unspecified.

Sources