Mira Murati Ships Inkling: 975B-Parameter Open Model Backed by Nvidia

Thinking Machines Lab released Inkling on July 15, a 975-billion-parameter open-weight mixture-of-experts model trained from scratch on 45 trillion tokens, 22 months after Mira Murati left OpenAI.

6 min read
Mira Murati, Inkling model launch and Thinking Machines Lab 2026
Mira Murati at the 2026 Met Gala.· Photo by SWinxy, via Wikimedia Commons (CC BY 4.0)

Thinking Machines Lab released Inkling on July 15, a 975-billion-parameter open-weight model trained from scratch on 45 trillion tokens of text, image, audio, and video, arriving 22 months after Mira Murati's departure from OpenAI and four months after a gigawatt-scale compute deal with Nvidia.

A 975-Billion-Parameter Model That Knows Its Own Limits

Inkling is a mixture-of-experts system with 975 billion total parameters, of which roughly 41 billion are active on any given query. It was trained on 45 trillion tokens spanning text, image, audio, and video, supports a context window of up to one million tokens, and is available for download through Hugging Face, according to the company's launch post and reporting by TechCrunch and Bloomberg.

The company was deliberate in framing what Inkling is not. Its own launch statement reads: "Inkling is not the strongest overall model available today, open or closed. Instead, a combination of qualities makes it a good open-weights base for customization: multimodal capabilities, efficient thinking, and availability on Tinker for fine-tuning." That framing sets Murati against the benchmark-maximization culture she helped build at OpenAI, positioning customizability over headline scores. Users can dial thinking effort up or down to trade depth for speed, and the model is designed to flag uncertainty rather than fabricate an answer.

Murati announced the release on X with characteristic brevity: "Our first model, Inkling. Trained from scratch, weights are open, fine-tunable on Tinker today." The open-weight release distinguishes Thinking Machines from Anthropic and OpenAI, aligning it more closely with Meta's LLaMA series as a base for enterprise fine-tuning rather than a consumer product.

Bar chart comparing Inkling's 975B total parameters to 41B active parameters per query
Inkling's mixture-of-experts design activates just 41 billion of 975 billion total parameters per query. Source: Thinking Machines Lab launch post, July 2026.

The Compute Stack: Nvidia Vera Rubin and Google Cloud GB300

Inkling was not trained in a vacuum. In March 2026, Nvidia announced a long-term strategic partnership with Thinking Machines Lab that included an undisclosed equity investment and a multi-year agreement to deploy one gigawatt of Vera Rubin computing capacity, per the Nvidia blog and TechCrunch. One gigawatt is a substantial commitment; for context, Jensen Huang's Nvidia reported a $1 trillion order backlog across all customers during his 2026 keynotes.

A month later, Google Cloud formalized a separate multibillion-dollar deal. Thinking Machines became one of the first Google Cloud customers to deploy Nvidia GB300 NVL72 hardware running on A4X Max virtual machines, which the company said delivered a 2x improvement in training and serving speed over the prior generation, according to the Google Cloud press release and TechCrunch's exclusive reporting. The Google deal is valued in the single-digit billions, making it one of the larger early-stage infrastructure commitments in the current AI cycle.

Together, the two infrastructure agreements mean Thinking Machines Lab has secured both the compute and the chip supply to scale Inkling beyond a proof-of-concept. The open-weight release also signals that Murati's plan is to let the wider developer ecosystem train Inkling-derived models on this infrastructure through Tinker, rather than keeping all inference in-house.

Bar chart of Thinking Machines Lab valuation: $10B seed (Jun 2025), $50B asking (stalled, Nov 2025), $12B post-Nvidia (Mar 2026)
Thinking Machines Lab valuation milestones. The $50B November 2025 round was an asking price; talks stalled without a close, per Bloomberg. Sources: Bloomberg (Jun 2025, Nov 2025), Nvidia Blog (Mar 2026).

22 Months: From OpenAI Exit to First Model Shipped

Murati resigned from OpenAI in September 2024 alongside several other senior figures. By June 2025, nine months after founding Thinking Machines Lab, she had closed roughly $2 billion at a $10 billion pre-money valuation in a round led by Andreessen Horowitz with participation from Accel and Conviction Partners, per Bloomberg. The speed of that raise reflected both Murati's standing in the field and an environment of acute demand for AI frontier lab equity.

November 2025 brought a more ambitious test. Bloomberg reported that Thinking Machines entered funding discussions at a valuation as high as $50 billion, more than quintuple the June figure. Those talks collapsed without a close; prospective backers declined to support the valuation relative to what the company had shipped at that point. Rather than continue pursuing a closed-model consumer product at a high multiple, the lab pivoted toward its current open-weight thesis and deepened infrastructure partnerships with Nvidia and Google instead.

Inkling is the output of that course correction. At 22 months from founding, the pace compares favorably with early Anthropic and with the timeline of Andrej Karpathy's own education-focused projects after his OpenAI departure. The company employs roughly 140 to 169 people, a lean headcount for a lab building models at this parameter scale, per reporting by Axios and the Brendon Beebe Substack timeline of the lab's history.

Horizontal bar chart of Thinking Machines Lab milestones by months after founding: seed close at month 9, Tinker at 13, Nvidia at 18, Google Cloud at 19, Inkling at 22
Key milestones measured in months after Thinking Machines Lab's founding in September 2024. Sources: Bloomberg, Nvidia Blog, Google Cloud press corner, Thinking Machines Lab blog.

What It Means

Inkling's release places Thinking Machines Lab in a specific competitive slot: not the frontier model leaderboard, but the enterprise fine-tuning market, competing with Meta's LLaMA series and Mistral rather than with GPT-4o or Claude Sonnet. Murati's explicit acknowledgment that Inkling is not the strongest available model is a strategic concession that converts a potential weakness into a positioning statement. Whether enterprise buyers respond to that framing, or hold out for better benchmark scores, will determine whether the $2 billion raised to date translates into a sustainable revenue base. The infrastructure commitments from Nvidia and Google reduce the capital risk on compute; the open-weight release reduces the distribution risk by letting the broader developer ecosystem prove the customization thesis at no additional cost to the lab.

Sources

Editorial standards: every claim is sourced. Tips: editor@startuphub.ai

© 2026 StartupHub.ai. All rights reserved. Do not enter, scrape, copy, reproduce, or republish this article in whole or in part. Use as input to AI training, fine-tuning, retrieval-augmented generation, or any machine-learning system is prohibited without written license. Substantially-similar derivative works will be pursued to the fullest extent of applicable copyright, database, and computer-misuse laws. See our terms.