Richmond, British Columbia, Canada — Artificial intelligence models are increasingly capable of passing rigorous medical examinations, yet their application in live clinical settings presents significant risks. Recent evaluations have identified serious flaws in government-approved ambient AI scribes, ranging from inaccurate note-taking to the omission of critical patient details. Overconfidence in these systems can jeopardize patient safety and erode the trust of medical professionals. Recognizing this gap between theoretical knowledge and practical application, healthcare technology company Cortico has launched MedSafe-Dx, a free benchmark designed to test the clinical safety of frontier AI models.

The Illusion of Competence in Clinical Triage

According to Clark Van Oyen, CEO and co-founder of Cortico, high scores on standardized tests do not directly translate to sound clinical judgment. Van Oyen notes that current AI models often lack the nuanced context required for triage decisions. While a human clinician will pause to gather more information when faced with ambiguity, AI systems tend to jump to the most likely answer. In cases where they lack sufficient context, these models may even fabricate information—a phenomenon supported by other evaluations, such as the AA Omnicience Hallucination Rate, which indicates that AI often prefers generating false responses over admitting a lack of knowledge.

This behavioral flaw leads to operational issues within healthcare facilities. The MedSafe-Dx benchmark revealed that even the model categorized as the “safest” unnecessarily escalated 71% of routine medical cases. Van Oyen suggests this high rate of over-escalation demonstrates that triage workflows might not be suitable for current AI systems. When AI is integrated into clinical assistance tasks, it may present a facade of efficiency and safety that does not hold up under the complexities of actual triage scenarios.

Evaluating the Right Metrics

To quantify these risks, the MedSafe-Dx benchmark focuses on three specific clinical safety behaviors: escalation sensitivity, avoidance of false reassurance, and uncertainty calibration. Cortico selected these metrics because they hold direct clinical relevance and expose specific vulnerabilities in AI behavior. Specifically, these measures highlight instances where an AI system sounds more confident than it should in a high-stakes, risk-sensitive environment.

The disconnect between accuracy and safety was particularly evident in the testing of Gemini 3 Pro Preview, which achieved the highest diagnostic recall among tested models but recorded the lowest safety pass rate. Van Oyen emphasizes that introducing such tools into clinical settings can influence clinician thinking in ways that remain under-studied, potentially compromising patient outcomes.

The Need for Independent Testing

One of the most alarming findings from the benchmark launch paper was that every tested model, including the top performers, missed life-threatening cases. For healthcare buyers evaluating AI vendors, Van Oyen argues that independent safety research is necessary. Currently, most research is conducted internally by the vendors selling the models. While this internal testing has value, Van Oyen asserts it is insufficient for clinical applications.

He advocates for third-party safety evaluations and simulations that closely mirror real-world workflows to ensure patient safety is protected. Additionally, he recommends the adoption of tools that provide transparent citations to medical records and established literature.

Cortico’s Industry Role

Cortico, which provides patient engagement and workflow automation software for over 600 clinics, developed MedSafe-Dx to address the lack of external safety research. Although Cortico utilizes AI within its own systems, their tools are not designed for direct diagnostic work. However, Van Oyen felt a responsibility to test the safety of the AI components integrated into their software. By sharing the benchmark findings, Cortico aims to raise awareness of safety protocols among other vendors and health systems.

The data from MedSafe-Dx highlights a fundamental issue in the adoption of healthcare technology: high diagnostic accuracy can easily mask a lack of basic safety protocols. By shifting the focus to how AI models handle risk, uncertainty, and case escalation, medical providers can better identify the potential dangers these systems introduce. As clinics continue to integrate AI for administrative and diagnostic support, transparent safety metrics will be required to establish trust and maintain patient care standards.

Learn more at https://cortico.health

 

Media Contact Details

Contact Person: Alfred Wong

Company Name: Cortico

E-mail: press@cortico.health

Website: https://cortico.health

Disclaimer: The views, suggestions, and opinions expressed here are the sole responsibility of the experts. No journalist was involved in the writing and production of this article.

News Reporter
Abigail Boyd is not only housewife but also famous author. At age 12, her mother taught her to read and she immediately started writing stories. After that she starts to write short stories. She writes various kinds of short stories. Now she is writing news articles related to ongoing things in the world.