Unlocking Ultra-Long Context for LLMs
MiniMax Sparse Attention breaks the context window barrier for LLMs, enabling millions of tokens with significant compute reduction and practical speedups.
Visual TL;DR
quadratic cost of standard softmax attention hinders ultra-long context
agentic workflows, code reasoning, persistent memory require millions of tokens
From the articleThe insatiable demand for ultra-long context capabilities in frontier LLMs, spanning agentic workflows, repository-scale code reasoning, and persistent memory, is currently stymied by the quadratic cost of standard softmax attention.
From the article 2 mentionsTo surmount this challenge, the researchers introduce MiniMax Sparse Attention (MSA), a novel blockwise sparse attention mechanism built upon Grouped Query Attention (GQA).
From the article 2 mentionsMSA employs a lightweight Index Branch to score and select a Top-k subset of key-value blocks for each GQA group, enabling group-specific sparse retrieval.
From the article 2 mentionsThe Main Branch then executes exact block-sparse attention exclusively over these selected blocks.
enables millions of tokens with significant compute reduction
From the articleThis computational barrier renders models untenable at deployment scale for contexts stretching into the millions of tokens.
optimized for GPU execution and efficient deployment across architectures
From the articleCrucially, when paired with its optimized kernel, MSA delivers substantial wall-clock speedups: 14.2x for prefill and 7.6x for decoding on H800 hardware.
© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Written by
Daniel SingerEditor, StartupHub.ai
Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.