What Mira Murati Said and Shipped at Thinking Machines in 2026

In June 2026, Mira Murati gave her first major interview since leaving OpenAI, outlining a new class of real-time AI called interaction models. A month later, Thinking Machines Lab shipped Inkling, a 975-billion-parameter open-weight model aimed at enterprise customization. Here is the full account of what she said and what it signals.

7 min read
Mira Murati, Thinking Machines Lab CEO, 2026 milestones recap
Mira Murati at the 2026 Met Gala, May 2026.· Photo by SWinxy, via Wikimedia Commons (CC BY 4.0)
Contents(5)

In June 2026, Mira Murati gave her first substantial press interview since leaving OpenAI as CTO in September 2024, sitting down with Bloomberg's Emily Chang at the Bloomberg Technology Summit in San Francisco. Eighteen months of near-silence followed by two concrete product moves in 60 days: a research preview of what the company calls interaction models, and then Inkling, a 975-billion-parameter open-weight model released on July 15. (Bloomberg, Thinking Machines Lab)

The Bloomberg Appearance: Interaction Models and a New Architecture

The June 4 Bloomberg Summit session was Murati's first extended on-record interview since the OpenAI departure. TechCrunch described the appearance as "stepping back into the spotlight, carefully," and the framing matched the substance: she did not relitigate the OpenAI chapter, focusing instead on what Thinking Machines Lab is building and why the underlying architecture differs from existing systems.

The central product argument she laid out was for interaction models, a class of system the company had previewed in a research release in May 2026. (MarkTechPost, May 13, 2026) The design differs from conventional prompt-and-response dynamics in a specific way: rather than waiting for a user turn, the interaction model processes continuous streams of audio, video, and text in roughly 200-millisecond intervals. A companion background model handles reasoning and tool use asynchronously, so latency-sensitive response and computationally heavy inference run on separate tracks simultaneously.

The distinction matters because it changes the interface metaphor. Conventional chat systems are still fundamentally structured as exchanges, one message sent and one received. The interaction model is designed to remain live, listening and responding within a conversation rather than between turns. The Bloomberg appearance positioned that shift as the company's central technical bet, not an incremental refinement of existing products.

Inkling: 975 Billion Parameters, 41 Billion Active, and a Customization Pitch

One month after the Bloomberg appearance, on July 15, Thinking Machines Lab released Inkling, its first publicly available model. The architecture is a mixture-of-experts system with 975 billion total parameters, of which about 41 billion are active for any given inference request. It was trained on 45 trillion tokens spanning text, image, audio, and video natively, meaning modality-switching is built into the base model rather than added via separate adapters. (Bloomberg, July 15, 2026)

The headline benchmark numbers are competitive but not top-of-table. Inkling scored 97.1 on AIME 2026, 87.2 on GPQA Diamond, 77.6 on SWE-Bench Verified, and 74.1 on MCP Atlas. On SimpleQA Verified, it came in at 43.9. Thinking Machines positioned the release explicitly as a starting point rather than a finished product. In the company's efficiency testing, Inkling matched Nvidia Nemotron 3 Ultra's Terminal Bench 2.1 score while generating approximately one-third as many tokens, a meaningful difference in inference cost at scale. (TechCrunch, July 15, 2026)

The release is open-weight: developers and companies can download the model and fine-tune it directly. Murati has framed this alongside Tinker, Thinking Machines' model-customization platform, which is designed to let enterprises adapt Inkling to specific domains, policies, and datasets without starting from scratch. VentureBeat reported that the company explicitly listed resistance to censorship as a design property of the release, a feature choice with clear implications for international enterprise customers operating in jurisdictions where model behavior is subject to government restriction. (VentureBeat)

The Enterprise Logic Behind Running Lean and Open

The strategic context for both moves, the interaction models research preview and the Inkling release, is the enterprise market. Murati's thesis, as she has described it across the Bloomberg interview and in company communications, is that the next phase of AI adoption turns on customization. Enterprises in knowledge-intensive verticals, including healthcare, legal, financial services, and enterprise software, need models that can absorb organization-specific data, apply internal policies, and adapt to proprietary processes. A model you download and retrain on your own infrastructure is fundamentally different in kind from a model you rent via API, both in what you can do with it and in the compliance properties it can satisfy.

Thinking Machines Lab is competing in a specific corridor of that market: open-weight models with native multimodal capability and an enterprise-grade toolchain. Its nearest comparables on the open-weight side include Mistral AI, which has 1,220 employees and its own enterprise customization stack, and Meta's Llama series, which is free but without a formal enterprise support layer. On the proprietary side sit OpenAI, Anthropic, and Google DeepMind, which offer managed API access rather than downloadable weights.

What is notable about Thinking Machines Lab's current operating profile is the deliberate restraint in headcount. StartupHub.ai data shows the company employs 212 people as of August 2026. Mistral AI, a close strategic peer in the open-weight enterprise segment, employs roughly six times as many at 1,220. The contrast is striking given both companies are competing for the same enterprise customization market: Murati is running a research-and-product concentrated team rather than scaling sales and operations ahead of revenue. That profile matches the go-to-market trajectory she described at Bloomberg: build the model and the toolchain first, then grow the customer base on top of a validated product.

What It Means

Murati's two public moves in summer 2026 define Thinking Machines Lab's position more clearly than the two years of build in private: a real-time interaction layer above conventional chat, and an open-weight model that enterprises can own and adapt. The combination is a specific bet that the frontier of enterprise AI value will shift from which model scores highest on a leaderboard to which model integrates most cleanly into a specific organization's infrastructure and data. Whether that thesis plays out depends on enterprise willingness to take on the operational cost of running and fine-tuning a 975-billion-parameter model internally, a non-trivial ask. The Tinker platform is the company's answer to that friction. The next signal to watch is whether enterprise contracts follow the June-July product announcements at a pace that matches the $12 billion valuation Thinking Machines Lab carries.

Sources

Editorial standards: every claim is sourced. Tips: [email protected]

© 2026 StartupHub.ai. All rights reserved. Do not enter, scrape, copy, reproduce, or republish this article in whole or in part. Use as input to AI training, fine-tuning, retrieval-augmented generation, or any machine-learning system is prohibited without written license. Substantially-similar derivative works will be pursued to the fullest extent of applicable copyright, database, and computer-misuse laws. See our terms.