Why million token context AI agents matter

MiniMax M3 wagers agents need 1M-token memory, native vision, and sparse attention to make long tool traces cheap and usable.

S
StartupHub.ai Staff
3 min read
MiniMax M3 sparse attention diagram showing 1M token context for AI agents
MiniMax M3 pairs sparse attention with native multimodality for long agentic traces· AI Engineer
Contents(3)

MiniMax built M3 around a simple premise: million token context AI agents need a full million tokens to survive tool loops, not just to summarize a book.

StartupHub data

Companies working on this

Profiles of the companies named in this story, with funding and a one-liner from our database.

MiniMax
$2.0B
MiniMax is a leading global technology company and one of the pioneers of large language models (LLMs) in Asia, developing proprietary LLMs across various modalities.
DeepSeek
$45.0B
Hugging Face
$4.5B
Hugging Face is the leading AI community and platform for machine learning collaboration, enabling developers to build, share, and deploy models, datasets, and applications.
Why million token context AI agents matter - AI Engineer
Why million token context AI agents matter, from AI Engineer

Olive Song, a former NYU PhD in Yann LeCun's lab who now works at MiniMax, told Hugging Face co-founder Thomas Wolf on the AI Engineer stage that the 400-billion-parameter M3 keeps vision, video and text together from step one.

How the sparse attention actually works

MSA splits the job into two branches. An index branch flags which blocks matter, and a sparse branch only computes attention on those blocks.

Think of it like a librarian who marks the relevant shelves before you read, so you never scan every page in a 10-million-token library.

Song said MiniMax already tested 10 million tokens in M1 and 01 for non-agentic review tasks, but M3 makes the 1 million window functional for multi-round tool use.

What this gets right, and where it could still break

Short contexts hurt agents because every tool response, image and video frame eats the window until the model forgets the goal.

MiniMax claims native multimodality from step zero avoids the adapter trap that drags down text performance, and says interleaved data with careful cleaning kept the vision tower from collapsing.

That collapse risk is real, and Song admitted earlier labs saw training diverge after a few steps when modalities were mixed too early.

The efficiency pitch holds up. MiniMax reports MiniMax Sparse Attention cuts per-token compute at 1M tokens to 1/20 of prior generation with 9x prefill and 15x decode speedups over M2. M3 activates only about 20 to 23 billion parameters, but cheap inference does not mean a trillion-token future is practical on current hardware.

Wolf noted M3 is the only top-five open model with working multimodal and long context today, which gives Hugging Face immediate distribution that DeepSeek, Moonshot's Kimi and GLM do not have in the same package.

One detail most recaps missed: the sparse attention design came from an intern, a hint that MiniMax lets anyone propose and ship architecture after releases instead of gating research.

Song said builders should stress test long video plus tool use now and send failures back, especially on unstructured PowerPoints and hour-long videos where retrieval alone fails.

If agents must actually watch tutorials to use tools, a million tokens is a starting point, not a ceiling.

© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
S

Written by

StartupHub.ai Staff

Editorial team

The staff writers of StartupHub.ai, ranging from investment analysts to avid AI tool users, early adopters and critical enthusiasts. Backgrounds span engineering, business and the arts. We hold every piece to rigorous standards of research and review.