Thinking Machines Lab released Inkling on July 15, a 975-billion-parameter open-weight model trained from scratch on 45 trillion tokens of text, image, audio, and video, arriving 22 months after Mira Murati's departure from OpenAI and four months after a gigawatt-scale compute deal with Nvidia.
A 975-Billion-Parameter Model That Knows Its Own Limits
Inkling is a mixture-of-experts system with 975 billion total parameters, of which roughly 41 billion are active on any given query. It was trained on 45 trillion tokens spanning text, image, audio, and video, supports a context window of up to one million tokens, and is available for download through Hugging Face, according to the company's launch post and reporting by TechCrunch and Bloomberg.
The company was deliberate in framing what Inkling is not. Its own launch statement reads: "Inkling is not the strongest overall model available today, open or closed. Instead, a combination of qualities makes it a good open-weights base for customization: multimodal capabilities, efficient thinking, and availability on Tinker for fine-tuning." That framing sets Murati against the benchmark-maximization culture she helped build at OpenAI, positioning customizability over headline scores. Users can dial thinking effort up or down to trade depth for speed, and the model is designed to flag uncertainty rather than fabricate an answer.
Murati announced the release on X with characteristic brevity: "Our first model, Inkling. Trained from scratch, weights are open, fine-tunable on Tinker today." The open-weight release distinguishes Thinking Machines from Anthropic and OpenAI, aligning it more closely with Meta's LLaMA series as a base for enterprise fine-tuning rather than a consumer product.
The Compute Stack: Nvidia Vera Rubin and Google Cloud GB300
Inkling was not trained in a vacuum. In March 2026, Nvidia announced a long-term strategic partnership with Thinking Machines Lab that included an undisclosed equity investment and a multi-year agreement to deploy one gigawatt of Vera Rubin computing capacity, per the Nvidia blog and TechCrunch. One gigawatt is a substantial commitment; for context, Jensen Huang's Nvidia reported a $1 trillion order backlog across all customers during his 2026 keynotes.
