MiniMax M-3: open multimodal with sparse attention
Published 20 July 2026 · Editorial update 6 September 2026 · 2 min read · Enternovate

This archived article retains its title and date, not a promise of current model availability. Earlier parameter counts and speed multipliers need checked primary sources before reuse. For a deployment, begin with the exact checkpoint or API endpoint and its documentation.
Multimodal support is a pipeline property. Confirm which inputs the provider accepts and which formats the agent actually sends. An image-capable model does not prove that a particular integration can analyse video or a scanned document end to end.
Build evaluation cases with known answers. Include low-quality scans, tables, misleading captions and unsupported files. Require the agent to say when it cannot inspect a source rather than presenting an inference as an observation.
Long inputs can increase latency, memory requirements and cost. Measure the whole workflow, including extraction and tool calls. Sparse attention is an architectural technique, not a blanket guarantee of practical local inference.
Xavani can coordinate supported providers and tools, while Nyarhi can organise approved results as knowledge. Decide what may be stored, how long it is retained and who can retrieve it. Validate the integration and licence before introducing client documents.