Hippocratic AI: 200M Patient Calls Show AI's Healthcare Promise

Hippocratic AI's Vivek Muppalla details how their AI has conducted 200M+ patient calls, achieving 99.89% "no harm" accuracy through advanced architecture and rigorous evaluation.

7 min read
Vivek Muppalla presenting on stage at AI Engineer World's Fair
AI Engineer
Visual TL;DR
Healthcare ScarcityDriver
system built on scarcity of clinicians, time, and money, necessitating triage
From the article 4 mentionsSpeaking at the AI Engineer World's Fair, Muppalla highlighted the transformative potential of AI in healthcare, moving from a system historically built on scarcity to one of abundance.
Hippocratic AI MissionContext
From the article 6 mentionsHippocratic AI's mission is to build "clinically safe abundance for all," not by replacing clinicians, but by empowering them to reach more patients than ever before.
200M+ Patient CallsCore
AI has conducted over 200 million patient calls, demonstrating scalability
Intelligence-Latency FlywheelContext
optimizing for speed and quality in AI-driven clinical conversations
ASR & Contextual UnderstandingCore
advanced speech recognition and deep understanding of clinical context
Orchestration & VerifiersCore
power of orchestrating AI models with human-like verification steps
From the articleFor critical functions like scheduling, verifiers are used to confirm accuracy, achieving a 99.49% success rate.
99.89% No HarmOutcome
achieving high accuracy through advanced architecture and rigorous evaluation
From the article 3 mentionsThey boast a 99.89% accuracy rating for "no harm" and an 8.5 out of 10 patient satisfaction score across over 60 health systems.
Healthcare AbundanceEffect
From the article 5 mentionsSpeaking at the AI Engineer World's Fair, Muppalla highlighted the transformative potential of AI in healthcare, moving from a system historically built on scarcity to one of abundance.
Contents(7)

Vivek Muppalla, head of engineering at Hippocratic AI, shared insights into the company's groundbreaking work in using AI for clinical patient conversations. Speaking at the AI Engineer World's Fair, Muppalla highlighted the transformative potential of AI in healthcare, moving from a system historically built on scarcity to one of abundance.

Hippocratic AI: 200M Patient Calls Show AI's Healthcare Promise - AI Engineer
Hippocratic AI: 200M Patient Calls Show AI's Healthcare Promise — from AI Engineer

Shifting from Scarcity to Abundance in Healthcare

Muppalla opened by noting the stark reality that most people have never received a proactive healthcare call from their provider, attributing this to a system built on scarcity of clinicians, time, and money, necessitating triage. However, he posited that advancements in technology and AI, capable of clinically safe conversations at dropping costs, are flipping this equation.

Hippocratic AI's mission is to build "clinically safe abundance for all," not by replacing clinicians, but by empowering them to reach more patients than ever before. The company's commitment to safety and patient well-being is underscored by an oath every employee takes, "Do No Harm, Patient First, Access for All." This ethos is reflected in their product, which has facilitated over 200 million clinical interactions with no significant safety incidents. They boast a 99.89% accuracy rating for "no harm" and an 8.5 out of 10 patient satisfaction score across over 60 health systems.

The Intelligence-Latency Flywheel

Muppalla delved into the technical challenges of creating AI that is both intelligent and fast enough for real-time conversation. He explained that many existing models offer high intelligence but suffer from significant latency, making them unusable for dialogue. Conversely, faster models often lack the necessary clinical accuracy. Hippocratic AI's solution lies in building a vertically integrated stack from the ground up, optimizing each component to achieve both high intelligence and low latency.

The company has developed a proprietary "constellation" architecture, named Polaris, which orchestrates multiple models. The system comprises three main parts: input (hearing), the "brain" (reasoning), and output (talking back). The input stage handles speech detection, including bilingual switching and background noise detection. The "brain" is not a single model but a network of 31 models, with a primary model managing the conversation and 30 specialist models covering areas like labs, medications, and scheduling. The output stage generates HD quality voice, custom personalities, and a clinical documentation engine that feeds information back into health systems.

ASR and Contextual Understanding

A significant portion of the talk focused on their Automatic Speech Recognition (ASR) system, designed to overcome the challenges of real-world audio. Muppalla highlighted that many perceived model reasoning failures are actually due to misheard audio. To combat this, Hippocratic AI employs a decoder-only audio LLM, fine-tuned on millions of clinical conversations, incorporating contextual biasing and domain knowledge. This approach significantly reduces word error rates, particularly for medical terms and phonetic nuances.

The system uses techniques like contextual biasing and single-word correction to improve accuracy. For instance, providing the model with the patient's medical history or the current task context helps narrow down potential medication names, reducing guesswork. They also employ a secondary scoring round for single-word responses, using conversational context to ensure accuracy. These optimizations have led to a 50% reduction in word error rate compared to off-the-shelf models and a P99 latency three times faster.

The Power of Orchestration and Verifiers

Muppalla explained that running 31 models simultaneously is achieved through a counterintuitive parallel processing approach, where each model first quickly checks if its input is relevant to the conversation. This short-circuiting mechanism helps manage latency. Additionally, asynchronous run models perform background verifications, ensuring the accuracy of information and tool calls. For critical functions like scheduling, verifiers are used to confirm accuracy, achieving a 99.49% success rate.

Optimizing for Speed and Quality

The company has implemented several inference stack optimizations to achieve both speed and quality without compromise. These include 4-bit quantization, speculative decoding, and KV cache compression. These techniques allow their models to deliver answers in under half a second with zero loss in clinical accuracy.

Rigorous Evaluation and Safety

Muppalla emphasized the critical role of evaluations, particularly in healthcare where a 1% error rate can have severe consequences. For a scheduling use case with 10,000 calls a day, a 1% error rate translates to 100 incorrect appointments daily. To address this, Hippocratic AI relies on a combination of over 7,500 trained clinicians and 775,000+ test calls. They grade model outputs not just on correctness but also on harm levels, from "no harm" to "severe harm" and "death." Their Polaris system has achieved a 99.89% accuracy for "no harm."

The Importance of Empathy and People

Beyond safety and accuracy, Hippocratic AI recognizes the importance of empathy in patient interactions. They have developed a benchmark called HEART to measure the empathy of their AI systems, ensuring that patients feel comfortable opening up. Muppalla concluded by highlighting the company's investment in people, with programs like the Agent Deployment Residency and AI/LLM Residency designed to train and attract top talent. He extended an invitation for others to join them in building the future of healthcare AI.

© 2026 StartupHub.ai. All rights reserved. Do not enter, scrape, copy, reproduce, or republish this article in whole or in part. Use as input to AI training, fine-tuning, retrieval-augmented generation, or any machine-learning system is prohibited without written license. Substantially-similar derivative works will be pursued to the fullest extent of applicable copyright, database, and computer-misuse laws. See our terms.