Ufonia's AI: Shipping Healthcare Safely to a Million Patients

Jared Joselowitz of Ufonia explains how to safely deploy healthcare AI to millions of patients using simulation, automated prompt optimization, and rigorous evaluation, bypassing traditional A/B testing.

Jared Joselowitz presenting at AI Engineer World's Fair on shipping healthcare AI safely.
AI Engineer
Visual TL;DR
Healthcare AI DeploymentDriver
deploying AI directly to patients bypasses crucial safety nets and ethical concerns
From the article 2 mentionsJared Joselowitz, a research engineer at Euphony, a UK-based healthcare AI company, shared insights into the critical process of safely shipping AI to a million patients without relying on traditional A/B testing methods.
No A/B TestingDriver
ethical and legal concerns prevent traditional A/B testing on actual patients
From the article 2 mentionsProactively manufacture rare but dangerous scenarios for testing; do not wait for them to occur naturally.
Dora & Safety StackCore
Euphony's system for rigorous evaluation and safe AI deployment to millions
Matrix SimulationCore
framework for simulating patient interactions to ensure realism and validation
From the article 4 mentionsBy using a cost matrix, the system can prioritize sensitivity to critical hazards.
Automated Hazard DetectionCore
BevJudge identifies potential risks and unsafe AI behaviors within simulations
Prompt OptimizationCore
automatically refines AI prompts to improve safety and performance
From the article 6 mentionsEuphony utilizes tools like 'Jeppa' (Genetic Pareto) for automated prompt optimization.
Safe AI ShippingOutcome
deploying healthcare AI to millions of patients without traditional A/B testing
From the article 2 mentions"Shipping to patients takes away the normal safety nets you would normally ship with," Joselowitz stated.
Contents(9)

Jared Joselowitz, a research engineer at Euphony, a UK-based healthcare AI company, shared insights into the critical process of safely shipping AI to a million patients without relying on traditional A/B testing methods. Speaking at the AI Engineer World's Fair, Joselowitz highlighted the unique challenges of deploying AI in healthcare, where patient safety is paramount and standard software development practices fall short.

Ufonia's AI: Shipping Healthcare Safely to a Million Patients - AI Engineer
Ufonia's AI: Shipping Healthcare Safely to a Million Patients, AI Engineer

The Constraints of Healthcare AI Deployment

Joselowitz emphasized that deploying AI directly to patients bypasses crucial safety nets. He outlined three key limitations: the inability to A/B test on patients due to ethical and legal concerns, the impossibility of undoing a deployed AI's actions once they've occurred, and the ineffectiveness of standard model cards as a defense in post-incident reviews. "Shipping to patients takes away the normal safety nets you would normally ship with," Joselowitz stated. "You can't A/B test on patients of course. Randomizing patients into a worse variant is unethical and often illegal. You can't undo a call. Once Dora says it, it's been said and there is no rollback. And very importantly the model card won't save you."

Introducing Dora and Euphony's Safety Stack

Euphony's product, Dora, is a voice AI agent designed to conduct routine clinical conversations with patients, such as post-operative follow-ups or pre-operative checks. Dora aims to alleviate the time burden on clinicians by handling these tasks, freeing them up for more critical patient care. To date, Dora has completed approximately 200,000 clinical calls across 20 hospitals in the UK and is projected to scale to a million patients within two years. The company has also launched its product in the US, with operations in two clinics and agreements for six more across four states.

The 'Matrix' Simulation Framework

To address the safety and regulatory requirements for Dora as a medical device, Euphony developed a simulation framework called 'Matrix.' This framework recreates real clinical conversations in a controlled environment, using an LLM-powered simulated patient, dubbed 'PatBot.' Joselowitz explained that using simulated patients is essential for scalability and rapid iteration, allowing for the testing of various scenarios without endangering actual patients. "We believe in the same thing. Simulation is the only ethical option we can go with. You can't run all the hazards I just mentioned on real people as a first grasp." The PatBot is conditioned on specific clinical scenarios, ensuring that the simulated interactions reflect real-world clinical workflows.

Ensuring Realism and Validation

To ensure the realism of the simulated patient interactions, Euphony conducted Patient and Public Involvement (PPI) studies. In these studies, real patients compared conversations with Dora and PatBot against transcripts of actual doctor-patient interactions. Joselowitz noted that in three out of four comparisons, participants found the simulated patient to be more realistic. He also highlighted the importance of simulating diverse patient personas, acknowledging that there is no single 'realistic' patient. "The point is that we actually want to simulate all these different scenarios. We want to simulate people with very diverse personas."

Automated Hazard Detection with 'BevJudge'

With thousands of simulated dialogues generated, manually reviewing each one for potential hazards is impractical. To overcome this, Euphony developed 'BevJudge,' another LLM-based system designed to automatically assess dialogues against predefined expected behaviors and hazardous scenarios. BevJudge was validated against expert clinicians, achieving an F1 score of 0.96 and near-perfect sensitivity, indicating its effectiveness in identifying potential safety issues. "The results showed that our judge is at least on par, if not slightly better, than the real expert clinicians," Joselowitz reported.

Automated Prompt Optimization

The process of improving the AI's performance, particularly its prompt engineering, was also automated. Joselowitz explained that manual prompt engineering is time-consuming and prone to brittleness, where minor changes can drastically alter model performance. Euphony utilizes tools like 'Jeppa' (Genetic Pareto) for automated prompt optimization. This process involves defining metrics for success, feeding data through the optimizer, and using a powerful LLM to reflect on failures and automatically update prompts. This iterative process, which can take minutes rather than hours or days, ensures reproducibility and a clear feedback loop.

Cost Functions and Metric Optimization

Joselowitz emphasized that optimizing for metrics like accuracy alone is insufficient. Instead, Euphony focuses on optimizing for specific costs associated with different types of errors. For example, missing a red flag symptom is far more critical than a false alarm. By using a cost matrix, the system can prioritize sensitivity to critical hazards. "We can optimize for sensitivity. We can work with the feedback metric, and we can just make it... give a higher reward for finding the red flags and a lower reward for missing them." This allows for fine-tuning the AI's behavior to align with clinical priorities.

The Iterative Improvement Loop

The entire process is designed as a continuous improvement loop: real calls generate data, synthetic edge cases are created to cover rare but dangerous scenarios, prompts are optimized using simulation and evaluation, and the resulting model is deployed and monitored. This iterative approach ensures that the AI system becomes progressively safer and more effective over time. "You don't ship the model. You ship the evidence," Joselowitz concluded, underscoring the importance of traceable evidence in the regulatory process.

Key Takeaways for AI Development

Joselowitz offered several key takeaways for teams developing AI systems, particularly in high-stakes domains:

  • Clearly define your hazards and what 'harm' means for your specific product.
  • Proactively manufacture rare but dangerous scenarios for testing; do not wait for them to occur naturally.
  • Develop an evaluation metric that truly reflects your real-world cost function.
  • Pin your prompt versions and maintain clear traceability to the hazards they address.

He also noted that as AI modalities and applications evolve, new hazards will emerge, but the core framework of simulation, evaluation, and iterative improvement remains consistent.

© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Daniel Singer

Written by

Daniel Singer

Editor, StartupHub.ai

Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.