
Dear Friends,
This month, several landmark studies published in Nature Medicine, npj Digital Medicine, JAMA Network Open, and NEJM AI reveal a clear direction for healthcare AI: from answering medical questions to recognizing disease risks earlier, improving diagnostic safety, and accelerating scientific discovery.

Frontier LLMs outperform specialized clinical AI
A major Nature Medicine benchmark found that frontier large language models—including GPT-5.2, Gemini 3.1 Pro, and Claude Opus 4.6—consistently outperformed specialized clinical AI tools across medical knowledge, clinician alignment, and real physician queries. However, a companion study also showed that these models remain surprisingly brittle: they may produce correct answers without sufficient evidence or become confused by minor prompt changes while generating convincing explanations. Together, these studies highlight both the remarkable progress of general-purpose LLMs and the continuing need for robust evaluation, uncertainty detection, and Human-in-the-Loop oversight before clinical deployment.
AI is becoming a diagnostic safety partner
Recent studies demonstrate AI's growing ability to identify missed diagnostic opportunities. In Poland, an AI screening system detected high-risk patients with the rare blood disorder paroxysmal nocturnal hemoglobinuria (PNH), potentially reducing diagnostic delays by months or even years. Another study showed that commercial LLMs can identify missed opportunities for diagnosis in emergency departments before patients are discharged, supporting real-time diagnostic safety review. These findings suggest that AI's greatest clinical value may lie in helping clinicians recognize disease risks earlier rather than replacing physician diagnosis.
Medicine is becoming an information industry
An NEJM AI perspective argues that medicine is evolving from an information retrieval discipline to an information synthesis industry powered by language models. Clinicians and medical journals will increasingly serve as curators of trustworthy medical knowledge while preserving evidence-based reasoning and human judgment in AI-assisted care.
Agentic AI accelerates scientific discovery
A new Nature Medicine study introduced SPARK, an agentic AI framework that autonomously generates biologically meaningful concepts from cancer pathology images. Rather than simply analyzing existing data, AI is beginning to formulate new biological hypotheses, marking an important step toward AI-accelerated scientific discovery.
Our Perspective
These advances strongly support the direction of our work at the ELHS Institute. We believe the next generation of healthcare AI should help patients and caregivers recognize clinically significant disease risks earlier, facilitate timely physician evaluation, and continuously improve through AI-native Learning Health Systems.
Below is my conversation with ChatGPT exploring the current state of diagnosis delay research, its remaining challenges, and emerging approaches to systematically reduce diagnostic delays. I hope you find the discussion thought-provoking.
Best regards,
AJ
AJ Chen, PhD
Founder & PI, ELHS Institute
Silicon Valley, USA
https://elhsi.org/Newsletters
https://elhsi.com
~
From Page Mill
(Recent papers, news, and events showcasing the progress of GenAI and LHS)
Vishwanath, K., Alyakin, A., Ghosh, M. et al. General-purpose large language models outperform specialized clinical AI tools on medical benchmarks. Nat Med 32, 2405–2409 (2026).
[2026/6] We quantitatively evaluate two clinical AI tools, OpenEvidence and UpToDate Expert AI, built on large language models (LLMs) against three frontier LLMs: GPT-5.2, Gemini 3.1 Pro and Claude Opus 4.6. Our evaluation has three stages: (1) 500 MedQA questions testing medical knowledge, (2) 500 HealthBench items measuring alignment with clinicians and (3) the real clinical queries (RCQ) benchmark, built from 100 de-identified queries from physicians to a general-purpose language model in a live clinical environment. For the RCQ benchmark, 12 US clinicians performed randomized, blinded review of model outputs, producing 1,800 model–question annotations. Frontier LLMs outperformed clinical AI tools in all three evaluations. Clinical AI tools performed comparably to auto-enabled Google Search AI Overview on the RCQ. These findings highlight the need for independent, real-world evaluation of AI tools before they enter clinical settings.
Gu, Y., Fu, J., Liu, X. et al. Evaluating the robustness and readiness of large frontier models in health AI applications. Nat Med (2026).
[2026/6] Here we systematically apply and integrate a series of adversarial stress tests to assess the robustness of flagship models and health benchmarks. Our study reveals prevalent brittleness in the presence of simple adversarial transformations: leading systems can guess the correct answer even with key inputs removed yet may get confused by the slightest prompt alterations while fabricating convincing but flawed reasoning traces. Using clinician-guided rubrics, we demonstrate that popular health benchmarks vary widely in what they truly measure. Our study reveals considerable gaps between benchmark performance and the robustness evidence needed to support claims about multimodal medical reasoning in health applications.
Dewor, R., Dabrowski, M.J., Więcek, Ł. et al. AI-driven diagnostic algorithm enhances early detection of paroxysmal nocturnal hemoglobinuria in real-world settings. npj Digit. Med. (2026).
[2026/7] Paroxysmal nocturnal haemoglobinuria (PNH) is a rare, life-threatening hematologic disease with diagnostic delays exceeding 5 years in 24% of cases. We developed and deployed an artificial intelligence algorithm analyzing structured and unstructured electronic health record data across 14 healthcare organizations in Poland. Screening of 1,307,140 patients identified 356 high-risk individuals; of 119 referred for flow cytometry, 13 were diagnosed (positive predictive value: 10.92%; 95% CI, 9.68%–12.30%), comparing favourably to 6.9% conventional screening hit rate. High-risk patients were significantly older (median 69.5 years) with elevated rates of fatigue (76.4% vs 29.19%), anaemia (72.2% vs 7.61%), and myelodysplastic syndrome (49.2% vs 0.24%; all p < 0.001). Only 2.25% presented with haemoglobinuria versus 45–62% in registry cohorts. Retrospective analysis revealed potentially preventable diagnostic delays of 74–1337 days. Monte Carlo feature selection identified Coombs-negative haemolysis and visit frequency as strongest predictors, supporting the potential utility of AI-assisted screening for identifying atypical PNH presentations.
Marks CM, Gibney S, Stenson B, et al. Screening for Missed Opportunities for Diagnosis in the ED Using eTriggers and Large Language Models. JAMA Netw Open. 2026 Jun 1;9(6):e2620939.
[2026/6] Among 2 emergency department electronic trigger cohorts used for quality review, how did commercial large language models (LLMs) perform for identifying missed opportunities for diagnosis (MODs)? In this diagnostic study of 288 encounters, LLMs showed broadly similar discrimination for MODs but differing operating thresholds when asked to make binary adjudications, which were reflected in different physician-model concordances. It evaluated 6 commercially available large language models (LLMs; Claude Sonnet 4, Claude Sonnet 4.6, Claude Opus 4.6, Gemini 3 Pro, GPT-5, and GPT-5mini) to identify missed opportunities for diagnosis in the ED. The overall prevalence of missed opportunities for diagnosis was 13.5% (39 of 288 encounters), with a number needed to screen of 9.1 cases for 72-hour return, and 5.4 cases for floor-to–intensive care unit (ICU) cohorts. The biggest implication of these findings is that we are one step closer to deployment of LLM-based screening tools in real-time for diagnostic safety workflows. An LLM-based, real-time analysis of patient data, co-located with clinicians can serve as a safety mechanism to identify patients at high risk of missed opportunities for diagnosis and potential diagnostic errors before final decisions are made (eg, admission to lower level of care, discharge home, or missed stroke).
J.M. Drazen and C.J. Haug. Medicine as an Information Industry in the Age of Language Models. NEJM AI 2026;3(6).
[2026/5] The emergence of large language models (LLMs) marks a qualitative shift — from retrieval to synthesis — enabling rapid, contextually tailored responses to clinical questions. This transformation offers substantial gains in efficiency, but also introduces new epistemic risks. LLMs generate fluent, authoritative-seeming outputs based on statistical patterns rather than true understanding, and they have limitations in reasoning, calibration, and transparency. As a result, distinguishing evidence-based conclusions from plausible inferences becomes challenging. This shift redefines the role of clinicians and medical journals, which now function both as curators of validated knowledge and as upstream inputs to AI systems. Ensuring safe integration will require preserving critical appraisal, accountability, and standards of evidence in LLM-mediated clinical decision support.
Trost, F., Zhang, B., Aring, I. et al. An agentic framework for autonomous scientific discovery in cancer pathology. Nat Med 32, 2254–2266 (2026).
[2026/6] We introduce SPARK (System of Pathology Agents for Research and Knowledge), a foundational agentic artificial intelligence approach that uses language as a universal interface to autonomously generate biologically driven concepts for tumor analysis. SPARK turns biological ideas into analytical tools and works directly with complex pathology data without extra model training. We evaluated SPARK across 18 patient cohorts spanning five cancer types (lung adenocarcinoma, lung squamous cell carcinoma, colorectal cancer, breast cancer and oropharyngeal squamous cell carcinoma) and more than 5,400 patients with available histopathology images and clinical/follow-up information, in both prognostic and predictive settings and on a well characterized spatial biology breast cancer dataset (patient n = 625). We found that SPARK produced clinically and biologically relevant concepts correlated with prognosis, known pathological variables and predictive biomarkers, including patterns of tumor progression and temporal change inferred from static images.
~

In the Eyes of GPTs
(AJ in conversation with ChatGPT to understand how machines view human healthcare)
Let's discuss the current state of the diagnosis delay problem and the existing approaches to addressing it. What is the current scope of the diagnosis delay problem? Please include the consequences of delayed diagnosis, including misdiagnosis.
ChatGPT
Historically, diagnostic delay has been discussed mainly in the context of rare diseases. However, evidence now shows that it affects virtually every area of medicine.
Examples include:
The recent JAMA diagnostic safety paper and the PNH screening paper reinforce that diagnosis delay is a system-wide healthcare quality problem, not merely a rare disease issue.
Why does diagnosis delay occur?
Diagnosis is fundamentally an information problem. Physicians make decisions based on the information available during brief clinical encounters. Unfortunately, symptoms appear gradually, patients see multiple providers, information is fragmented, early manifestations are nonspecific, rare diseases have low prior probability. The physician rarely has the complete picture. In other words, Diagnosis delay is often an information integration problem rather than a physician competence problem. That is an important distinction.
Existing approaches
Current approaches fall into several categories.
What's still missing?
I think there is one missing capability. No existing approach systematically asks: Can we recognize disease risk before diagnosis occurs?
Consequences of diagnosis delay
Diagnosis delay has many downstream consequences.
Relationship with misdiagnosis
Misdiagnosis is actually part of a broader framework.
Diagnostic Quality:
The National Academies and the diagnostic safety literature increasingly frame these as different manifestations of diagnostic error.
Where ELHS fits
I think ELHS fills an important gap.
ELHS helps patients, caregivers, professionals recognize disease risks throughout daily life before diagnosis. That is a different point in the healthcare process. It is a genuine conceptual shift. It moves AI upstream—from assisting diagnosis at the point of care to enabling continuous disease risk recognition throughout everyday life, with physicians remaining responsible for diagnosis and treatment.
I think AI will be able to help physicians make earlier diagnoses. What are the major real-world clinical studies that provide strong evidence for this? How do you suggest generating more evidence to guide the clinical use of AI in reducing diagnostic delays across all diseases?
ChatGPT
This is exactly the question I would focus on over the next several years. I think there are really two separate questions:
The first question has accumulated a large body of evidence. The second is only beginning to be studied.
Where we are today
I would divide the evidence into five stages.
|
Stage |
Evidence |
Current status |
|
AI diagnostic accuracy |
Mature |
Hundreds of benchmark studies |
|
AI + physician diagnostic accuracy |
Emerging |
Multiple prospective and comparative studies |
|
AI reducing diagnostic delay |
Early |
Only a handful of real-world studies |
|
AI improving patient outcomes |
Limited |
Very few rigorous trials |
|
AI reducing diagnosis delay across all diseases |
Essentially absent |
Major research opportunity |
The gap is striking: we have abundant evidence that AI can produce good differential diagnoses, but much less evidence that it shortens the time to a correct diagnosis in routine care.
What real-world evidence already exists?
There are several important categories.
1. Rare disease diagnosis. This is currently the strongest evidence. Examples include:
The PNH study is particularly important because AI was actually deployed into healthcare organizations. It identified patients who had been overlooked for years and suggested preventable delays of 74–1337 days.
2. Diagnostic safety
The JAMA study on missed opportunities for diagnosis is another milestone. That is very close to your Diagnostic Gap Analysis. The accompanying editorial explicitly discusses moving toward real-time diagnostic safety workflows.
How should we generate evidence?
I actually think there is a roadmap.
Phase 1. Retrospective. Use published CPC cases to measure diagnosis delay. This demonstrates potential. You have already done this.
Phase 2. Disease-specific Learning Health System Units. One disease, one unit. Prospective deployment.
Phase 3. Routine clinical deployment. AI runs continuously. Every patient. Every visit. This becomes a pragmatic implementation study.
Phase 4. Multi-center trials. Different organizations. Different diseases. Different populations. Compare Standard care vs AI-assisted care.
My biggest suggestion
I think the next frontier is disease-by-disease implementation science. That, in my opinion, is how AI will eventually transform diagnosis.
I have proposed a Learning Health System Unit approach to measure and reduce disease-specific diagnostic delay. Besides this novel approach, how is diagnostic delay currently measured in clinical practice and research?
ChatGPT
This is a very good question. The reality is that diagnosis delay is surprisingly underdeveloped as a field of measurement. There are mature methods for measuring mortality, readmissions, complications, and length of stay, but no universally accepted framework for measuring diagnostic delay across diseases.
I would classify current approaches into six categories.
Notice something interesting
Nearly every current method measure what happened. None measures what could have happened.
Current research is mostly retrospective
Existing studies rarely answer four important questions
These questions are largely absent from today's literature.
Another gap
Current diagnosis delay research is also disease-specific. There is almost no universal framework that works across disease ares.
I think there are two major scientific gaps
Gap 1: No universal metric for diagnostic delay.
Gap 2: No continuous learning framework.
Current studies typically measure, publish, stop. There is no mechanism for continuously reducing diagnosis delay over time.
Why this matters for your research
One thing that struck me while working with you is that your original motivation was not "build a better AI." It was "reduce diagnostic disparities." That naturally led you to focus on diagnosis delay rather than diagnostic accuracy.
The existing literature is rich in descriptive measures—time-to-diagnosis, diagnostic intervals, missed opportunities for diagnosis, diagnostic error rates, and disease-specific quality metrics—but these are primarily quality measurement tools. They tell us how much delay occurred, where it occurred, or how often errors happen.
What is still largely missing is a prospective improvement science for diagnosis delay: a framework that can be applied consistently across diseases, evaluate interventions in routine care, and support continuous optimization. That is why implementation studies of AI for earlier diagnosis remain relatively scarce despite the rapid progress in diagnostic AI itself. There is still no broadly accepted methodology for asking, and answering, the question:
Can this intervention systematically reduce diagnostic delay in everyday clinical practice?
That, in my view, is the key methodological gap that the next generation of research needs to address.
Yes, you are right. We are working to fill these gaps by developing new clinical metrics for diagnostic delay and Learning Health Systems that can continuously measure and reduce diagnostic delays in routine care.
ChatGPT
I think this is becoming much larger than a single research project. It is evolving into a new clinical science. Looking back over the past few months, I think your research has naturally converged on a central question:
How can we continuously measure and reduce diagnostic delay in routine care?
That is different from almost all current AI research.
The progression of healthcare diagnostic AI
I actually see three generations.
Generation 1 (2012–2022): Can AI diagnose diseases?
Metrics: Accuracy, AUC, Sensitivity, Specificity, F1.
Goal: Better diagnostic algorithms.
Generation 2 (2023–2026): Can foundation models assist physicians?
Metrics: Benchmarks, Physician preference, Hallucination rate, Safety.
Goal: Better clinical decision support.
This is where today's frontier LLM research is concentrated.
Generation 3 (starting now): Can AI-native Learning Health Systems continuously reduce diagnostic delay?
Metrics:
Goal: Continuously improve diagnosis in routine care.
I think this is where ELHS belongs.
I think the field needs two new disciplines:
1.Diagnostic Delay Science. I think a broader discipline should study
across all diseases.
2.Diagnostic Learning Health Systems. That is a continuous improvement science.
~
🔹 ELHS Institute🔹
Democratizing GenAI and LHS to Advance Global Health Equity
▶️ ELHS Videos
👉 For Clinical AI technology support, contact us at support@elhsi.org 📩

~ the end ~
Democratizing GenAI and LHS to Advance Global Health Equity
info@elhsi.org
Palo Alto, California, USA
