#Thinking Machines Lab
19 articles with this tag

What Mira Murati Said and Shipped at Thinking Machines in 2026
In June 2026, Mira Murati gave her first major interview since leaving OpenAI, outlining a new class of real-time AI called interaction models. A month later, Thinking Machines Lab shipped Inkling, a 975-billion-parameter open-weight model aimed at enterprise customization. Here is the full account of what she said and what it signals.

Sam Altman's OpenAI Alumni Now Lead Companies Worth $1 Trillion
Former OpenAI executives now lead companies worth over $1 trillion combined. Here is how Sam Altman's alumni network became his biggest rivals in the 2026 AI landscape.

Mira Murati Bets on Open Weights: Inkling vs. the Closed AI Stack
Thinking Machines Lab launched Inkling, a 975-billion-parameter open-weight model, under Apache 2.0 in July 2026. Here is how Mira Murati's open customisation strategy differs from OpenAI and Anthropic's proprietary approaches, and what the Nvidia gigawatt deal means for the lab's compute ambitions.

Thinking Machines Lab cuts costs with Inkling-Small
Thinking Machines Lab launches Inkling-Small, a 276B parameter model that delivers comparable performance to its larger predecessor at a fraction of the cost.

Mira Murati Ships Inkling: 975B-Parameter Open Model Backed by Nvidia
Thinking Machines Lab released Inkling on July 15, a 975-billion-parameter open-weight mixture-of-experts model trained from scratch on 45 trillion tokens, 22 months after Mira Murati left OpenAI.

Inkling Model Lands on Databricks
Thinking Machines Lab's Inkling model is now accessible on Databricks via Unity AI Gateway, enhancing enterprise AI development for coding and agentic tasks.

Together AI adds Inkling multimodal model
Together AI integrates Inkling, a new multimodal AI model from Thinking Machines Lab, offering text, image, and audio processing with controllable reasoning.

Mira Murati's Interaction Model: Full-Duplex AI at 0.4 Seconds
Thinking Machines Lab's TML-Interaction-Small processes audio and video in 200ms chunks, responds in 0.4 seconds, and runs full-duplex without a VAD harness. Here is how the architecture works and what Murati said at Bloomberg Tech.

Mira Murati’s Thinking Machines: $2B Raised, $50B Stalled, Shipping
Thinking Machines Lab raised $2 billion at a $12 billion seed valuation in July 2025, then sought $50 billion from investors four months later before backers passed. Here is the full financial arc: valuation milestones, Nvidia and Google infrastructure deals, and two shipped products through June 2026.

AI's Next Leap: Interaction Models
Thinking Machines Lab introduces 'interaction models' for AI, enabling real-time, multimodal collaboration that mirrors human conversation.
Architectural Interactivity, Linguistic Interpretability, and Molecular Synthesis: The Frontier of Native AI
Three organisations now define the frontier of native AI: Thinking Machines is rebuilding human-AI collaboration as a low-latency interaction model, the Effable movement wants interpretable safety frameworks like SafetyAnalyst, and Isomorphic Labs is converting AlphaFold into an end-to-end drug design engine. The common thread is moving from AI as a layer of abstraction toward AI as a fundamental component of human and biological systems.
Thinking Machines Lab Wants to Replace OpenAI Realtime With a Model That Listens While It Speaks
Mira Murati's lab published its first technical paper, arguing that real-time interactivity should be a native model capability rather than scaffolding bolted around turn-based language models, and it ships benchmarks where GPT Realtime-2 scores near zero.

Thinking Machines, NVIDIA Forge Gigawatt AI Pact
Thinking Machines Lab and NVIDIA announce a gigawatt-scale partnership for AI training, including a significant investment from NVIDIA.

Tinker launches OpenAI API compatibility, challenging vendor lock-in.

Tinker Call for Projects: Thinking Machines Lab Seeks ML Innovators

On-Policy Distillation LLMs Redefine Post-Training Efficiency

Solving LLM Nondeterminism: A Breakthrough for Reproducible AI

The Real Reason for LLM Inference Nondeterminism
The true cause of LLM inference nondeterminism is not random GPU math, but a systemic failure of "batch invariance" tied to unpredictable server load.