Ufonia's AI: Shipping Healthcare Safely to a Million Patients

Jared Joselowitz of Ufonia explains how to safely deploy healthcare AI to millions of patients using simulation, automated prompt optimization, and rigorous evaluation, bypassing traditional A/B testing.

10 min read
Jared Joselowitz presenting at AI Engineer World's Fair on shipping healthcare AI safely.
AI Engineer

Visual TL;DR. Healthcare AI Deployment leads to No A/B Testing. Healthcare AI Deployment requires Dora & Safety Stack. No A/B Testing necessitates Dora & Safety Stack. Dora & Safety Stack uses Matrix Simulation. Matrix Simulation enables Automated Hazard Detection. Automated Hazard Detection informs Prompt Optimization. Dora & Safety Stack achieves Safe AI Shipping. Matrix Simulation enables Safe AI Shipping. Prompt Optimization contributes to Safe AI Shipping.

  1. Healthcare AI Deployment: deploying AI directly to patients bypasses crucial safety nets and ethical concerns
  2. No A/B Testing: ethical and legal concerns prevent traditional A/B testing on actual patients
  3. Dora & Safety Stack: Euphony's system for rigorous evaluation and safe AI deployment to millions
  4. Matrix Simulation: framework for simulating patient interactions to ensure realism and validation
  5. Automated Hazard Detection: BevJudge identifies potential risks and unsafe AI behaviors within simulations
  6. Prompt Optimization: automatically refines AI prompts to improve safety and performance
  7. Safe AI Shipping: deploying healthcare AI to millions of patients without traditional A/B testing
Visual TL;DR
Visual TL;DR, startuphub.ai Healthcare AI Deployment requires Dora & Safety Stack. Dora & Safety Stack uses Matrix Simulation. Dora & Safety Stack achieves Safe AI Shipping. Matrix Simulation enables Safe AI Shipping requires uses achieves enables Healthcare AI Deployment Dora & Safety Stack Matrix Simulation Safe AI Shipping From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai Healthcare AI Deployment requires Dora & Safety Stack. Dora & Safety Stack uses Matrix Simulation. Dora & Safety Stack achieves Safe AI Shipping. Matrix Simulation enables Safe AI Shipping requires uses achieves enables Healthcare AIDeployment Dora & SafetyStack Matrix Simulation Safe AI Shipping From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai Healthcare AI Deployment requires Dora & Safety Stack. Dora & Safety Stack uses Matrix Simulation. Dora & Safety Stack achieves Safe AI Shipping. Matrix Simulation enables Safe AI Shipping requires uses achieves enables Healthcare AI Deployment deploying AI directly to patients bypassescrucial safety nets and ethical concerns Dora & Safety Stack Euphony's system for rigorous evaluationand safe AI deployment to millions Matrix Simulation framework for simulating patientinteractions to ensure realism andvalidation Safe AI Shipping deploying healthcare AI to millions ofpatients without traditional A/B testing From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai Healthcare AI Deployment requires Dora & Safety Stack. Dora & Safety Stack uses Matrix Simulation. Dora & Safety Stack achieves Safe AI Shipping. Matrix Simulation enables Safe AI Shipping requires uses achieves enables Healthcare AIDeployment deploying AIdirectly topatients bypasses… Dora & SafetyStack Euphony's systemfor rigorousevaluation and safe… Matrix Simulation framework forsimulating patientinteractions to… Safe AI Shipping deployinghealthcare AI tomillions of… From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai Healthcare AI Deployment leads to No A/B Testing. Healthcare AI Deployment requires Dora & Safety Stack. No A/B Testing necessitates Dora & Safety Stack. Dora & Safety Stack uses Matrix Simulation. Matrix Simulation enables Automated Hazard Detection. Automated Hazard Detection informs Prompt Optimization. Dora & Safety Stack achieves Safe AI Shipping. Matrix Simulation enables Safe AI Shipping. Prompt Optimization contributes to Safe AI Shipping leads to requires necessitates uses enables informs achieves enables contributes to Healthcare AI Deployment deploying AI directly to patients bypassescrucial safety nets and ethical concerns No A/B Testing ethical and legal concerns preventtraditional A/B testing on actual patients Dora & Safety Stack Euphony's system for rigorous evaluationand safe AI deployment to millions Matrix Simulation framework for simulating patientinteractions to ensure realism andvalidation Automated Hazard Detection BevJudge identifies potential risks andunsafe AI behaviors within simulations Prompt Optimization automatically refines AI prompts toimprove safety and performance Safe AI Shipping deploying healthcare AI to millions ofpatients without traditional A/B testing From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai Healthcare AI Deployment leads to No A/B Testing. Healthcare AI Deployment requires Dora & Safety Stack. No A/B Testing necessitates Dora & Safety Stack. Dora & Safety Stack uses Matrix Simulation. Matrix Simulation enables Automated Hazard Detection. Automated Hazard Detection informs Prompt Optimization. Dora & Safety Stack achieves Safe AI Shipping. Matrix Simulation enables Safe AI Shipping. Prompt Optimization contributes to Safe AI Shipping leads to requires necessitates uses enables informs achieves enables contributes to Healthcare AIDeployment deploying AIdirectly topatients bypasses… No A/B Testing ethical and legalconcerns preventtraditional A/B… Dora & SafetyStack Euphony's systemfor rigorousevaluation and safe… Matrix Simulation framework forsimulating patientinteractions to… Automated HazardDetection BevJudge identifiespotential risks andunsafe AI behaviors… PromptOptimization automaticallyrefines AI promptsto improve safety… Safe AI Shipping deployinghealthcare AI tomillions of… From startuphub.ai · The publishers behind this format

Jared Joselowitz, a research engineer at Euphony, a UK-based healthcare AI company, shared insights into the critical process of safely shipping AI to a million patients without relying on traditional A/B testing methods. Speaking at the AI Engineer World's Fair, Joselowitz highlighted the unique challenges of deploying AI in healthcare, where patient safety is paramount and standard software development practices fall short.

Ufonia's AI: Shipping Healthcare Safely to a Million Patients - AI Engineer
Ufonia's AI: Shipping Healthcare Safely to a Million Patients — from AI Engineer

The Constraints of Healthcare AI Deployment

Joselowitz emphasized that deploying AI directly to patients bypasses crucial safety nets. He outlined three key limitations: the inability to A/B test on patients due to ethical and legal concerns, the impossibility of undoing a deployed AI's actions once they've occurred, and the ineffectiveness of standard model cards as a defense in post-incident reviews. "Shipping to patients takes away the normal safety nets you would normally ship with," Joselowitz stated. "You can't A/B test on patients of course. Randomizing patients into a worse variant is unethical and often illegal. You can't undo a call. Once Dora says it, it's been said and there is no rollback. And very importantly the model card won't save you."

Introducing Dora and Euphony's Safety Stack

Euphony's product, Dora, is a voice AI agent designed to conduct routine clinical conversations with patients, such as post-operative follow-ups or pre-operative checks. Dora aims to alleviate the time burden on clinicians by handling these tasks, freeing them up for more critical patient care. To date, Dora has completed approximately 200,000 clinical calls across 20 hospitals in the UK and is projected to scale to a million patients within two years. The company has also launched its product in the US, with operations in two clinics and agreements for six more across four states.

The 'Matrix' Simulation Framework

To address the safety and regulatory requirements for Dora as a medical device, Euphony developed a simulation framework called 'Matrix.' This framework recreates real clinical conversations in a controlled environment, using an LLM-powered simulated patient, dubbed 'PatBot.' Joselowitz explained that using simulated patients is essential for scalability and rapid iteration, allowing for the testing of various scenarios without endangering actual patients. "We believe in the same thing. Simulation is the only ethical option we can go with. You can't run all the hazards I just mentioned on real people as a first grasp." The PatBot is conditioned on specific clinical scenarios, ensuring that the simulated interactions reflect real-world clinical workflows.

Ensuring Realism and Validation

To ensure the realism of the simulated patient interactions, Euphony conducted Patient and Public Involvement (PPI) studies. In these studies, real patients compared conversations with Dora and PatBot against transcripts of actual doctor-patient interactions. Joselowitz noted that in three out of four comparisons, participants found the simulated patient to be more realistic. He also highlighted the importance of simulating diverse patient personas, acknowledging that there is no single 'realistic' patient. "The point is that we actually want to simulate all these different scenarios. We want to simulate people with very diverse personas."

Automated Hazard Detection with 'BevJudge'

With thousands of simulated dialogues generated, manually reviewing each one for potential hazards is impractical. To overcome this, Euphony developed 'BevJudge,' another LLM-based system designed to automatically assess dialogues against predefined expected behaviors and hazardous scenarios. BevJudge was validated against expert clinicians, achieving an F1 score of 0.96 and near-perfect sensitivity, indicating its effectiveness in identifying potential safety issues. "The results showed that our judge is at least on par, if not slightly better, than the real expert clinicians," Joselowitz reported.

Automated Prompt Optimization

The process of improving the AI's performance, particularly its prompt engineering, was also automated. Joselowitz explained that manual prompt engineering is time-consuming and prone to brittleness, where minor changes can drastically alter model performance. Euphony utilizes tools like 'Jeppa' (Genetic Pareto) for automated prompt optimization. This process involves defining metrics for success, feeding data through the optimizer, and using a powerful LLM to reflect on failures and automatically update prompts. This iterative process, which can take minutes rather than hours or days, ensures reproducibility and a clear feedback loop.

Cost Functions and Metric Optimization

Joselowitz emphasized that optimizing for metrics like accuracy alone is insufficient. Instead, Euphony focuses on optimizing for specific costs associated with different types of errors. For example, missing a red flag symptom is far more critical than a false alarm. By using a cost matrix, the system can prioritize sensitivity to critical hazards. "We can optimize for sensitivity. We can work with the feedback metric, and we can just make it... give a higher reward for finding the red flags and a lower reward for missing them." This allows for fine-tuning the AI's behavior to align with clinical priorities.

The Iterative Improvement Loop

The entire process is designed as a continuous improvement loop: real calls generate data, synthetic edge cases are created to cover rare but dangerous scenarios, prompts are optimized using simulation and evaluation, and the resulting model is deployed and monitored. This iterative approach ensures that the AI system becomes progressively safer and more effective over time. "You don't ship the model. You ship the evidence," Joselowitz concluded, underscoring the importance of traceable evidence in the regulatory process.

Key Takeaways for AI Development

Joselowitz offered several key takeaways for teams developing AI systems, particularly in high-stakes domains:

  • Clearly define your hazards and what 'harm' means for your specific product.
  • Proactively manufacture rare but dangerous scenarios for testing; do not wait for them to occur naturally.
  • Develop an evaluation metric that truly reflects your real-world cost function.
  • Pin your prompt versions and maintain clear traceability to the hazards they address.

He also noted that as AI modalities and applications evolve, new hazards will emerge, but the core framework of simulation, evaluation, and iterative improvement remains consistent.

© 2026 StartupHub.ai. All rights reserved. Do not enter, scrape, copy, reproduce, or republish this article in whole or in part. Use as input to AI training, fine-tuning, retrieval-augmented generation, or any machine-learning system is prohibited without written license. Substantially-similar derivative works will be pursued to the fullest extent of applicable copyright, database, and computer-misuse laws. See our terms.