MiMo-V2.5 is a native omnimodal model by Xiaomi. It delivers Pro-level agentic performance at roughly half the inference cost, while surpassing MiMo-V2-Omni in multimodal perception across image and video understanding tasks. Its 1M context window supports complete documents, extended conversations, and complex task contexts in a single pass, making it ideal for integration with agent frameworks where strong reasoning, rich perception, and cost efficiency all matter.
| $0.14 | $0.28 | $0.0028 | 2.70s | 27 tps | ||
| $0.40 | $2.00 | $0.08 | 2.70s | 16 tps | ||
| $0.168 | $0.336 | $0.00336 | 1.75s | 28 tps | ||
| $0.168 | $0.336 | $0.0034 | 3.42s | 27 tps | ||
15% off | $0.14$0.119 | $0.28$0.238 | $0.003$0.00255 | 2.80s | 25 tps |
P50, best across providers
P50, best provider
When an error occurs in an upstream provider, we can recover by routing to another healthy provider, if your request filters allow it. You can access per-provider uptime data programmatically through the Endpoints API. Learn more about our load balancing and customization options.
