# AI's Next Leap: Interaction Models _Thinking Machines Lab introduces 'interaction models' for AI, enabling real-time, multimodal collaboration that mirrors human conversation._ **Updated:** 2026-08-22 **Published:** 2026-05-13 **Source:** https://www.startuphub.ai/ai-news/technology/2026/ai-s-next-leap-interaction-models --- Thinking Machines Lab is pushing the boundaries of human-AI collaboration with its research preview of [interaction models](https://thinkingmachines.ai/blog/interaction-models/). This new approach embeds interactivity directly into the AI, rather than relying on external systems, aiming to make working with AI as fluid as collaborating with another person. The models are designed to process audio, video, and text continuously, enabling real-time thinking, responding, and acting. Collaboration BottleneckDriver From the article 3 mentionsThe core idea is to address what the lab calls the 'collaboration bottleneck.' Current AI systems, often optimized for autonomous tasks, struggle with human-in-the-loop workflows.Thinking Machines LabCoreFrom the article 2 mentionsThinking Machines Lab is pushing the boundaries of human-AI collaboration with its research preview of interaction models.Improved AI WorkflowsOutcomeovercomes limitations of single-thread AI processingFrom the articleThe core idea is to address what the lab calls the 'collaboration bottleneck.' Current AI systems, often optimized for autonomous tasks, struggle with human-in-the-loop workflows.developsInteraction ModelsContextAI embeds interactivity directly, not external systemsFrom the article 9+ mentionsThinking Machines Lab is pushing the boundaries of human-AI collaboration with its research preview of interaction models.Real-time MultimodalContextFrom the article 2 mentionsThe models are designed to process audio, video, and text continuously, enabling real-time thinking, responding, and acting.Meet Humans WhereEffectFrom the articleThe goal is to enable AI interfaces that meet humans where they are, facilitating natural interaction through speaking, listening, seeing, and interjecting.leads toFluid Human-AIEffectenables working with AI as natural as peopleFrom the article 2 mentionsThis new approach embeds interactivity directly into the AI, rather than relying on external systems, aiming to make working with AI as fluid as collaborating with another person. The core idea is to address what the lab calls the 'collaboration bottleneck.' Current AI systems, often optimized for autonomous tasks, struggle with human-in-the-loop workflows. Users can't always specify needs upfront, and interfaces often push humans out, despite their value in clarifying and providing feedback. The goal is to enable AI interfaces that meet humans where they are, facilitating natural interaction through speaking, listening, seeing, and interjecting. ## The Collaboration Bottleneck Existing AI models operate on a single thread, waiting for user input to complete before processing new information. This turn-based system creates a narrow communication channel, limiting the nuances of human knowledge, intent, and judgment that can be conveyed. It's akin to resolving a critical disagreement over email instead of an in-person conversation. [Thinking](https://thinkingmachines.ai/blog/index.xml) Machines argues that interactivity must scale with intelligence and be an intrinsic part of the AI model itself. This contrasts with current methods that stitch together external components to simulate real-time capabilities. The lab's research suggests that models trained from scratch with a multi-stream, micro-turn design achieve state-of-the-art performance in both intelligence and responsiveness. ## Capabilities Building interactivity into the model unlocks several new capabilities: - **Seamless dialog management:** The model inherently understands user cues like thinking, yielding, or self-correction without needing a separate component. - **Verbal and visual interjections:** The AI can interrupt or interject contextually, mirroring natural human conversation flow. - **Simultaneous speech:** Both the user and the AI can speak concurrently, useful for applications like live translation. - **Time-awareness:** The model has a direct understanding of elapsed time. - **Simultaneous tool calls, search, and generative UI:** While conversing, the AI can concurrently perform searches, call tools, or generate user interfaces, seamlessly integrating results back into the dialogue. These continuous interactions foster an experience that feels more like a true collaboration than a series of prompts and responses. It's a significant shift from the segmented, turn-based interactions that currently define much of our engagement with AI. --- Original analysis from [startuphub.ai](https://www.startuphub.ai), the #1 AI startup directory.