AI Achieves Clinician-Level Video Consults

AMIE (Video), a Gemini-based AI, achieves clinician-level performance in real-time video consultations, surpassing text-only models and matching human physicians in key assessment areas.

Illustration of an AI system interacting with a patient via video consultation.
AMIE (Video) represents a new frontier in AI's ability to engage with the nuances of patient-physician interactions.
Visual TL;DR
Text AI limitationsDriver
discards vital non-verbal cues, disadvantaging patients unable to articulate symptoms in writing
From the article 3 mentionsLimitations persist in fine anatomical precision, discerning subtle affective nuances, and processing high-frequency movements, indicating areas for future refinement.
Early AV AIDriver
showed feasibility but fell short of clinician-level performance in medical assessments
From the articleEarly attempts to imbue medical AI with audio-visual capabilities have shown feasibility but have fallen short of clinician-level performance.
AMIE (Video)Core
a new Gemini-based multi-agent system integrating dialogue, reasoning, and real-time perception
From the article 5 mentionsA new Gemini-based multi-agent system, AMIE (Video), is now demonstrating expert-level performance in real-time clinical video consultations.
Real-time AV perceptionContext
From the article 2 mentionsThis system integrates low-latency dialogue, clinical reasoning, and crucially, real-time audio-visual perception.
Expert assessmentEffect
demonstrates expert-level assessment across modalities, including subtle non-verbal cues
Clinician-level performanceOutcome
achieves clinician-level performance in real-time video consultations, matching human physicians
From the article 2 mentionsEarly attempts to imbue medical AI with audio-visual capabilities have shown feasibility but have fallen short of clinician-level performance.
Surpasses text modelsOutcome
surpasses text-only models in key assessment areas, bridging the sensory gap

The standard for patient-physician interaction has always been audio-visual, a rich modality that captures subtle non-verbal cues essential for diagnosing illness. Text-based AI, while promising, inherently discards these vital perceptual dimensions and disadvantages patients unable to articulate symptoms in writing. Early attempts to imbue medical AI with audio-visual capabilities have shown feasibility but have fallen short of clinician-level performance.

Bridging the Sensory Gap in Clinical AI

A new Gemini-based multi-agent system, AMIE (Video), is now demonstrating expert-level performance in real-time clinical video consultations. This system integrates low-latency dialogue, clinical reasoning, and crucially, real-time audio-visual perception. The researchers developed a taxonomy and automated evaluations specifically for clinical audio-visual cues in telehealth settings to guide its development. In a randomized Objective Structured Clinical Examination (OSCE) study involving 30 primary care physicians (PCPs), 15 patient actors, and 100 clinical scenarios, AMIE (Video) was directly compared against its text-only counterpart, AMIE (Text), and human PCPs conducting video consultations.

Expert-Level Assessment Across Modalities

The results are striking. Clinical evaluators rated AMIE (Video) as on par with or superior to PCPs in history-taking, diagnosis, management, and physical observation and examination. Patient actors expressed a preference for AMIE's approach to assessing and explaining conditions, finding it more effective than human physicians in these specific areas. While patient actors favored AMIE (Video) for its communicative effectiveness, convenience, and sense of being understood over text chat interfaces, they did note a preference for PCPs in rapport and partnership building. Limitations persist in fine anatomical precision, discerning subtle affective nuances, and processing high-frequency movements, indicating areas for future refinement.

© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Daniel Singer

Written by

Daniel Singer

Editor, StartupHub.ai

Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.