AI's Next Leap: Interaction Models

Thinking Machines Lab introduces 'interaction models' for AI, enabling real-time, multimodal collaboration that mirrors human conversation.

Abstract visualization of interconnected nodes representing AI interaction model data streams.
Conceptual illustration of Thinking Machines' interaction models.
Visual TL;DR
Collaboration BottleneckDriver
From the article 3 mentionsThe core idea is to address what the lab calls the 'collaboration bottleneck.' Current AI systems, often optimized for autonomous tasks, struggle with human-in-the-loop workflows.
Thinking Machines LabCore
From the article 2 mentionsThinking Machines Lab is pushing the boundaries of human-AI collaboration with its research preview of interaction models.
Improved AI WorkflowsOutcome
overcomes limitations of single-thread AI processing
From the articleThe core idea is to address what the lab calls the 'collaboration bottleneck.' Current AI systems, often optimized for autonomous tasks, struggle with human-in-the-loop workflows.
Interaction ModelsContext
AI embeds interactivity directly, not external systems
From the article 9+ mentionsThinking Machines Lab is pushing the boundaries of human-AI collaboration with its research preview of interaction models.
Real-time MultimodalContext
From the article 2 mentionsThe models are designed to process audio, video, and text continuously, enabling real-time thinking, responding, and acting.
Meet Humans WhereEffect
From the articleThe goal is to enable AI interfaces that meet humans where they are, facilitating natural interaction through speaking, listening, seeing, and interjecting.
Fluid Human-AIEffect
enables working with AI as natural as people
From the article 2 mentionsThis new approach embeds interactivity directly into the AI, rather than relying on external systems, aiming to make working with AI as fluid as collaborating with another person.

Thinking Machines Lab is pushing the boundaries of human-AI collaboration with its research preview of interaction models. This new approach embeds interactivity directly into the AI, rather than relying on external systems, aiming to make working with AI as fluid as collaborating with another person. The models are designed to process audio, video, and text continuously, enabling real-time thinking, responding, and acting.

The core idea is to address what the lab calls the 'collaboration bottleneck.' Current AI systems, often optimized for autonomous tasks, struggle with human-in-the-loop workflows. Users can't always specify needs upfront, and interfaces often push humans out, despite their value in clarifying and providing feedback. The goal is to enable AI interfaces that meet humans where they are, facilitating natural interaction through speaking, listening, seeing, and interjecting.

The Collaboration Bottleneck

Existing AI models operate on a single thread, waiting for user input to complete before processing new information. This turn-based system creates a narrow communication channel, limiting the nuances of human knowledge, intent, and judgment that can be conveyed. It's akin to resolving a critical disagreement over email instead of an in-person conversation.

Thinking Machines argues that interactivity must scale with intelligence and be an intrinsic part of the AI model itself. This contrasts with current methods that stitch together external components to simulate real-time capabilities. The lab's research suggests that models trained from scratch with a multi-stream, micro-turn design achieve state-of-the-art performance in both intelligence and responsiveness.

Capabilities

Building interactivity into the model unlocks several new capabilities:

  • Seamless dialog management: The model inherently understands user cues like thinking, yielding, or self-correction without needing a separate component.
  • Verbal and visual interjections: The AI can interrupt or interject contextually, mirroring natural human conversation flow.
  • Simultaneous speech: Both the user and the AI can speak concurrently, useful for applications like live translation.
  • Time-awareness: The model has a direct understanding of elapsed time.
  • Simultaneous tool calls, search, and generative UI: While conversing, the AI can concurrently perform searches, call tools, or generate user interfaces, seamlessly integrating results back into the dialogue.

These continuous interactions foster an experience that feels more like a true collaboration than a series of prompts and responses. It's a significant shift from the segmented, turn-based interactions that currently define much of our engagement with AI.

© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Daniel Singer

Written by

Daniel Singer

Editor, StartupHub.ai

Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.