AI Clinical Notes: The Silent Errors We Miss

Sebastian Fox of Componso discusses the critical but often unseen errors in AI-generated clinical notes and proposes a new approach to AI evaluation.

4 min read
Sebastian Fox presenting on AI clinical note errors at AI Engineer World's Fair.
AI Engineer

In the rapidly evolving world of AI in healthcare, the accuracy and reliability of AI-generated clinical notes are paramount. Sebastian Fox, a medical doctor and CEO of Componso, a company specializing in AI evaluation systems, recently highlighted the significant risks associated with these systems. He presented a stark warning: the most dangerous AI errors are often the ones that appear perfectly fine on the surface.

AI Clinical Notes: The Silent Errors We Miss - AI Engineer
AI Clinical Notes: The Silent Errors We Miss — from AI Engineer

The Subtle Dangers of AI in Clinical Notes

Fox opened his presentation with a compelling example: an AI-generated clinical note for a patient with a headache. The note, which read like a routine case, recommended paracetamol for a tension-type headache. However, it omitted a crucial detail the patient mentioned: jaw ache when chewing. This omission, coupled with the patient's age (over 50), pointed to a potentially critical condition, giant cell arteritis, which, if untreated, could lead to blindness within days. The AI's failure to capture this nuance resulted in a note that was technically correct but dangerously incomplete, treating a potential emergency as a common ailment.

He contrasted this with another, more overt error where an AI incorrectly diagnosed a young man with diabetes and prescribed non-existent medications, leading to him being invited for diabetic eye screening for a condition he did not have. While such glaring errors are often caught, Fox emphasized that the more insidious, subtle errors are the ones that slip through unnoticed and can cause greater harm.

The Scale of the Problem: Errors in Production

Fox cited a significant real-world study indicating that approximately 1 in 20 AI-generated clinical notes contained an error serious enough to potentially cause significant harm to a patient. Furthermore, nearly 1 in 5 notes had an important omission, and over 1 in 10 contained a hallucination. This issue is particularly concerning given the rapid adoption of AI in healthcare, with services like ambient scribes already in use in about a third of U.S. practices.

The lack of comprehensive adverse event reporting for most AI systems means these errors often go undetected, creating a "flying blind" scenario where the impact on patient care is unknown but potentially severe. Fox stressed that this problem extends beyond healthcare to any high-stakes application of AI.

Why Current Evaluation Systems Fail

Fox explained that Large Language Models (LLMs) are now highly capable, and their errors are less about factual inaccuracies and more about subtle misinterpretations or omissions. He presented a scatter plot illustrating AI failures, showing that while some obvious errors are caught, many critical ones, particularly omissions and subtle misinterpretations, slip through automated checks.

The core issue, according to Fox, lies in the AI's inability to grasp the concept of "what matters" in a given context. This "taste" is tacit, contextual, and constantly evolving, making it difficult to codify into a fixed rubric for automated evaluation. While AI can process information, it lacks the human judgment to discern the critical importance of certain details over others.

The Proposed Solution: A Continuous Evaluation Loop

Fox proposed a solution centered on a continuous feedback loop for AI evaluation. Instead of relying solely on pre-defined rubrics or models that quickly become outdated, he advocated for a system that:

  • Discover: Identify failure modes from real-world outputs, not just theoretical ones.
  • Capture: Collect expert judgment and reasoning on these failures.
  • Calibrate: Assemble context-specific standards for each output by referencing past judgments, corrections, and guidelines.

This approach involves keeping "taste" as examples themselves, allowing the system to learn and adapt. Fox emphasized that the easiest way to start is by having experts provide free-form comments on real outputs, which then serve as the raw material for this iterative improvement process.

The Future of AI Evaluation

Fox concluded by reiterating that evaluation is not a static process but an ongoing activity. By focusing on continuous discovery, capture, and calibration, organizations can build more robust and trustworthy AI systems, ensuring that the "taste" of what truly matters is integrated into their core functionality. This continuous loop is essential for any application where "confidently wrong" decisions can have significant consequences.

© 2026 StartupHub.ai. All rights reserved. Do not enter, scrape, copy, reproduce, or republish this article in whole or in part. Use as input to AI training, fine-tuning, retrieval-augmented generation, or any machine-learning system is prohibited without written license. Substantially-similar derivative works will be pursued to the fullest extent of applicable copyright, database, and computer-misuse laws. See our terms.