MiniMax M3: Native Multimodal MoE Model with 1M-Token Context
MiniMax M3 is a native multimodal model from MiniMax with a 1M-token context window. It has about 428B total parameters with about 23B activated per token, and it takes text, image, and video input. MiniMax built it for long-horizon coding and agentic work, and reports frontier-level results across long-horizon agentic benchmarks in both coding and cowork tasks.
Key Features
- Native Multimodality - Trained on mixed text, image, and video data from the very first step, for deeper fusion across modalities than an adapter added later
- 1M-Token Context - Long documents, large repositories, and long agent trajectories fit in a single context
- MiniMax Sparse Attention (MSA) - A sparse attention operator built for million-token contexts; MiniMax reports 9x faster prefill and 15x faster decode than MiniMax M2 at 1M context, with per-token compute cut to 1/20
- Sparse Mixture-of-Experts - Only about 23B of the 428B parameters run on any given token
- Three Reasoning Modes - Thinking can be enabled, disabled for lower latency and higher throughput, or left adaptive so the model decides when extra reasoning helps
- Coding and Cowork - Strong on long-horizon agentic benchmarks for both software engineering and general cowork tasks
Use Cases
- Long-horizon agentic coding and repository-level software engineering
- Multi-step agent workflows with tool calling
- Image and video understanding alongside text
- Long-document and multi-document analysis across very long contexts
- Assistants that switch between fast direct answers and step-by-step reasoning per request
Architecture and Design
MiniMax M3 is a Mixture-of-Experts language model paired with a vision encoder for image and video input. The language model has 60 layers: the first 3 are dense and the remaining 57 route each token through 4 of 128 experts plus one shared expert. MiniMax Sparse Attention replaces standard grouped-query attention in the sparse layers, cutting the attention compute and memory cost of long contexts while preserving model quality. The model weights and model card are on Hugging Face.
Inference
The reasoning mode is chosen per request as enabled, adaptive, or disabled, with adaptive as the default. MiniMax recommends temperature 1.0 and top-p 0.95.
Deploy MiniMax M3 on Vast.ai for multimodal reasoning, agentic coding, and long-context work on flexible GPU infrastructure.