In a post on X, @arena reports that Xiaomi’s MiMo-V2.6-Pro and MiMo-V2.6-Flash have entered Agent Arena with different reported strengths. MiMo-V2.6-Pro ranks #5 among “open-source models,” while MiMo-V2.6-Flash ranks #9 among “open models.” The supplied evidence does not explain whether those two groups use the same inclusion criteria.
Arena presents these figures as measurements from real-world agentic sessions. They should be read as Arena-reported results, not as independently verified benchmark findings.

Image credit: @arena on X
MiMo-V2.6-Pro ranks higher in Arena’s reported results
Arena reports that MiMo-V2.6-Pro recorded a +3.17% net improvement across more than 8,100 real-world agentic sessions. It places the model fifth among the open-source models included in the reported ranking.
The post text compares that result with MiMo-V2.5-Pro, reporting a #13 position and -7.23% net improvement. However, the supplied Arena graphic displays MiMo-V2.5-Pro at #14 with approximately -7.2%. The source materials therefore conflict on the prior model’s rank, so the reported nine-place improvement should not be treated as settled. The post text does support its reported 10.4 percentage-point difference in net improvement, but the supplied material does not explain the metric’s calculation.
Confirmed Success is Pro’s strongest reported signal
MiMo-V2.6-Pro’s strongest result in the announcement is its Confirmed Success score. Arena defines Confirmed Success as explicit user feedback that the task worked. The model recorded +7.35% on that measure and ranked #2 among the open-source models.
The supplied post does not provide the task mix, feedback volume or uncertainty information behind that score. Confirmed Success is therefore a reported user-feedback measure, not a complete assessment of the model’s reliability across tasks.
MiMo-V2.6-Flash emphasizes lower reported cost
Arena places MiMo-V2.6-Flash at #9 among open models, with a -0.57% net improvement across more than 13,000 real-world agentic sessions. Its session count is higher than Pro’s reported 8,100-plus sessions, so the two figures should not be treated as results from identical sample sizes.
Arena reports a median cost of $0.04 per task for MiMo-V2.6-Flash, which it says is 56% lower than MiMo-V2.6-Pro’s median cost. The post does not provide Pro’s absolute median cost per task.

Image credit: @arena on X
What the Pareto-frontier result means
Arena says MiMo-V2.6-Flash landed on the Agent Arena Pareto frontier. A Pareto frontier shows options for which improving one measured dimension—here, cost or net improvement—would require giving up some of the other. In this announcement, that explains why Arena highlights Flash’s cost-and-performance position rather than simply calling it the highest-ranked model.
Flash’s reported position pairs its $0.04 median cost per task with a -0.57% net-improvement figure. That does not mean Flash ranked above Pro overall: Arena’s reported rankings place Pro at #5 and Flash at #9, while the frontier view highlights Flash’s lower reported cost.
Arena links to its Agent Arena Pareto leaderboard for the frontier view. The linked page did not provide readable text in the supplied source material, so it cannot independently establish the announcement’s rankings, costs, session counts or frontier placement.
How to interpret the two results
The announcement presents two different reported trade-offs:
MiMo-V2.6-Pro: a higher reported ranking, +3.17% net improvement across 8.1K-plus sessions and a +7.35% Confirmed Success score that ranks second among the open-source models.
MiMo-V2.6-Flash: a lower reported ranking, -0.57% net improvement across 13K-plus sessions and a $0.04 median cost per task, which Arena places on the Pareto frontier.
Together, these figures do not present one model as uniformly better. Pro has the stronger reported ranking and Confirmed Success result, while Flash has the lower reported cost in Arena’s comparison. The supplied announcement does not establish broader reliability, general model quality or pricing outside this measurement.





0 comments
No approved comments yet. You can start the conversation.
Leave a comment