When Elon Musk's xAI released Grok 4 Heavy in July 2025, it became the first AI system to score above 40% on Humanity's Last Exam, a graduate-level benchmark spanning science, law, and mathematics. The result did not come from a larger model or a longer context window. It came from an architecture that runs up to 16 instances of Grok 4 simultaneously on a single query, has each instance reason independently, and then synthesizes their answers before producing a final response. (Bloomberg, July 10, 2025) Thirteen months later, xAI's Grok Bot extended that same multi-agent design from a benchmark result to a commercial product for professional workflows. (Bloomberg, August 11, 2026)
Sixteen agents, one answer: how Grok 4 Heavy works
The standard Grok 4 model, released July 10, 2025, operates as a single-instance reasoning system. The Heavy tier runs a different process: the system spawns up to 16 parallel instances of Grok 4, each working through the same query independently. Once the instances finish reasoning, they compare their outputs and converge on a final answer through a debate-style synthesis. xAI describes the design as analogous to a study group where participants work separately, then compare notes. (The Information)
The benchmark results from this approach were specific. Grok 4 Heavy scored 100% on AIME 2025 (competitive mathematics), 96.7% on HMMT 2025 (another mathematics tournament), 88.4% on GPQA Diamond (graduate-level science reasoning), and 44.4% on Humanity's Last Exam with tools enabled. The HLE score was the first time any publicly available model crossed 40% on that benchmark, which was designed to resist saturation by current AI systems. On Artificial Analysis's Omniscience evaluation, the model recorded a 78% non-hallucination rate, the highest posted by any frontier model at the time of its release. (LLM Stats)
The architecture carries a compute cost: 16 parallel inference passes per query uses substantially more GPU capacity than a single pass. Grok 4 Heavy is priced at a premium to the standard model tier and available through Super Grok subscriptions and the xAI API. That overhead shapes where xAI can deploy the design commercially, and the product decisions that followed reflect it.
From Grok 4.5 to Grok Bot: applying the architecture to products
On July 8, 2026, xAI and Cursor jointly released Grok 4.5, a model built for finance, legal, and coding tasks. The release came weeks after SpaceX announced plans to acquire Cursor, the AI coding assistant, in a deal Bloomberg reported valued the startup at $60 billion. The two companies described Grok 4.5 as offering lower per-token costs than comparable models from OpenAI and Anthropic, with the design optimized for sustained professional use rather than one-off queries. (Bloomberg, July 8, 2026)
