On July 15, 2026, Thinking Machines Lab published Inkling, a 975-billion-parameter open-weight model, releasing full weights under the Apache 2.0 licence and positioning it as a structural alternative to the proprietary APIs sold by OpenAI and Anthropic. The move crystallises a bet that former OpenAI CTO Mira Murati has been building toward since founding the company in early 2025: that enterprises will pay more for customisable infrastructure than for commodity intelligence delivered by the token.
The open-weight wager: 975 billion parameters, zero licensing fees
Inkling is a Mixture-of-Experts (MoE) model: 975 billion total parameters, of which 41 billion are active per token. The company trained it on 45 trillion tokens spanning text, images, audio, and video; the model processes all four modalities as input and returns text, with a 1-million-token context window. A smaller companion model, Inkling-Small, was previewed alongside the main release: 276 billion total parameters, 12 billion active per token.
Unlike recent flagship releases from OpenAI, Anthropic, and Google, Thinking Machines Lab made no claim to state-of-the-art performance on standard benchmarks. The company stated plainly at launch that Inkling "is not the strongest overall model available today," framing it instead as a purpose-built base for downstream fine-tuning, per TechCrunch. The full weights are on Hugging Face under Apache 2.0, with no licensing restrictions on commercial use or derivative works.
The company published a Bridgewater Associates case study alongside the release: a version of Inkling fine-tuned on Bridgewater's proprietary financial data scored 84.7 percent on internal financial-reasoning tests, outperforming the proprietary APIs it replaced, while costing roughly one-fourteenth as much to operate on an ongoing basis. That ratio is the core of Thinking Machines Lab's commercial argument.
Tinker as the real product: customisation over commodity AI
Thinking Machines Lab does not sell intelligence by the token. Revenue runs through Tinker: companies upload proprietary data, fine-tune Inkling or another open-weight base, and receive weights they can download or host on Tinker's infrastructure. The enterprise keeps the IP. Thinking Machines Lab charges a platform fee and takes a cut of hosted compute, not a per-token inference margin. That distinction matters structurally because inference margins have compressed sharply industry-wide since 2025.
