Together AI Masters MiniMax M3 Inference

Together AI details engineering feats enabling efficient MiniMax M3 inference, unlocking 1M-token context and multimodality.

Together AI logo next to MiniMax logo with abstract AI graphics
Together AI partners with MiniMax for efficient M3 model inference.· Together AI
Visual TL;DR
MiniMax M3 DemandsDriver
From the article 5 mentionsMiniMax M3, designed for advanced coding, agentic workflows, and multimodal reasoning, presents unique serving demands, particularly with its extended context length and rich input processing requirements.
Extreme ContextDriver
From the article 3 mentionsThe company has detailed significant engineering breakthroughs enabling efficient MiniMax M3 inference, unlocking the model's ambitious 1 million token context window and native multimodal capabilities.
Native MultimodalityDriver
rich input processing requirements for diverse data
From the articleThe company has detailed significant engineering breakthroughs enabling efficient MiniMax M3 inference, unlocking the model's ambitious 1 million token context window and native multimodal capabilities.
MiniMax Sparse AttentionCore
novel mechanism reducing computational burden of long contexts
From the article 2 mentionsThe core of MiniMax M3's efficiency challenge lies in its novel MiniMax Sparse Attention (MSA) mechanism.
Together AI PlatformCore
From the article 6 mentionsTogether AI is positioning itself as the go-to platform for demanding large language models, announcing its role as the preferred cloud partner for MiniMax's latest M3 model.
Efficient InferenceEffect
enabling complex systems challenges for cutting-edge AI
From the article 2 mentionsThe company has detailed significant engineering breakthroughs enabling efficient MiniMax M3 inference, unlocking the model's ambitious 1 million token context window and native multimodal capabilities.
Advanced AI UnlockedOutcome
powering demanding large language models
From the articleMiniMax M3, designed for advanced coding, agentic workflows, and multimodal reasoning, presents unique serving demands, particularly with its extended context length and rich input processing requirements.

Together AI is positioning itself as the go-to platform for demanding large language models, announcing its role as the preferred cloud partner for MiniMax's latest M3 model. The company has detailed significant engineering breakthroughs enabling efficient MiniMax M3 inference, unlocking the model's ambitious 1 million token context window and native multimodal capabilities.

This collaboration highlights Together AI's commitment to tackling complex systems challenges for cutting-edge AI. MiniMax M3, designed for advanced coding, agentic workflows, and multimodal reasoning, presents unique serving demands, particularly with its extended context length and rich input processing requirements.

Engineering for Extreme Context and Multimodality

The core of MiniMax M3's efficiency challenge lies in its novel MiniMax Sparse Attention (MSA) mechanism. This architecture reduces the computational burden of long contexts by limiting the tokens each query attends to, a critical departure from quadratic scaling. Together AI's team developed a KV-Block-Major sparse attention kernel to optimize this, improving arithmetic intensity by reorganizing data flow.

Further enhancing long-context handling, Together AI integrated MSA with paged attention. This allows for dynamic KV cache management, crucial for variable request lengths, and reportedly yielded a 5% boost in decode throughput.

The model's multimodal capabilities necessitated a dedicated preprocessing pipeline. A new Rust-based Serving Model Gateway (SMG) now handles image and video decoding, resizing, and patching on the CPU. This offloads GPU resources, ensuring the inference engine focuses on generation.

These optimizations collectively resulted in performance improvements of 81% to 125% across various concurrency levels for agentic-style workloads, according to Together AI's internal benchmarks.

Together AI will host the open-weights MiniMax M3 model as a developer endpoint upon its public release.

© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Daniel Singer

Written by

Daniel Singer

Editor, StartupHub.ai

Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.

More from Daniel Singer