Dear Friends,
This month’s publications show continued progress toward responsible and clinically useful healthcare AI—but they also raise an increasingly important question: as AI becomes more capable and more deeply embedded in medicine and science, how do we ensure that it strengthens rather than weakens human judgment?

AI safety is becoming a broader question than whether an AI system produces a correct answer. A Nature feature examines the much-debated possibility of catastrophic risks from increasingly powerful AI. Closer to everyday medicine and research, a NEJM AI Perspective, “Where Are the Prepared Minds? The Impact of AI on Scientific Thinking,” raises a more immediate concern: routine reliance on AI could reduce opportunities for scientists and trainees to develop independent reasoning. Writing, interpreting evidence, questioning assumptions, and struggling through difficult problems are not merely ways to communicate scientific knowledge—they are part of how scientific judgment is formed. The authors therefore argue for deliberate, staged use of AI that preserves opportunities for independent thought.
At the same time, healthcare AI is progressing from technical performance toward prospective clinical evidence. Two new evaluation frameworks in NEJM AI and npj Digital Medicine emphasize that retrospective accuracy alone is insufficient. Clinical AI should advance through progressively stronger evidence—from technical and external validation to prospective evaluation, real-world implementation, monitoring, and mechanisms for responding to performance drift or safety signals. A prospective JAMA Cardiology study provides a practical example: AI-guided focused ultrasound performed by novice operators, combined with automated interpretation and targeted expert review, showed promise for scalable aortic stenosis screening.
AI is also moving beyond diagnosis toward decision support for treatment selection. A Nature Medicine study in non-small cell lung cancer developed multimodal AI models to predict immunotherapy outcomes. Importantly, its clinical usability study found that both expert and nonexpert physicians improved their predictions when assisted by an explainable AI tool. This points toward an important next stage of clinical AI: not simply predicting disease, but helping clinicians make better individualized treatment decisions.
Finally, progress must be equitable. A npj Digital Medicine Perspective highlights the widening global healthcare AI divide and examines privacy-enhancing technologies that could enable institutions to collaboratively develop AI without centralizing sensitive patient data. Such approaches may be especially important for bringing diverse populations and resource-limited health systems into the evidence-generation process rather than leaving them behind.
Together, these studies suggest that the next phase of healthcare AI should not be defined simply by more powerful models, but by better evidence, safer integration, better clinical decisions, and broader participation. They also leave us with a fundamental question: What happens to physicians’ and scientists’ own cognitive abilities when AI becomes a routine partner in learning, reasoning, writing, diagnosis, and decision-making?
In our discussion this month, In the Eyes of GPTs, we will begin exploring this question, with particular attention to the cognitive impact of AI on medical students and clinicians.
Best regards,
AJ
AJ Chen, PhD
Founder & PI, ELHS Institute
Silicon Valley, USA
https://elhsi.org/Newsletters
https://elhsi.com
~
From Page Mill
(Recent papers, news, and events showcasing the progress of GenAI and LHS)
Will AI really kill us all? The science behind the hype. 2026/9
Nature examines whether the technology might spell the end for humans, and why AI companies are calling for a slowdown.
Where Are the Prepared Minds? The Impact of AI on Scientific Thinking. NEJM AI 2026;3(10). DOI: 10.1056/AIe2601110
[20266/9] Preserving opportunities for independent thought is particularly important because scientific writing and interpretation are themselves mechanisms for organizing knowledge and developing judgment. As AI becomes more deeply embedded in research, scientific education should emphasize deliberate, staged use of these tools so that trainees can benefit from AI without compromising the cognitive skills required to become independent scientists.
From Technical Performance to Clinical Readiness: A Phase-Based Framework for Evidence Standards in Clinical Artificial Intelligence. NEJM AI 2026;3(10). DOI: 10.1056/AIp2600684
[2026/9] We propose a phase-based framework for evidentiary maturity in clinical AI, using cardiovascular AI as a representative use case. The framework spans model development and representation learning, internal validation, external validation, prospective evaluation of decision impact and implementation readiness, and postdeployment monitoring. Each phase is linked to a core question, expected evidence, and the type of claim that can reasonably be supported. By aligning claims with phase-appropriate evidence, this framework is intended to help investigators, reviewers, regulators, health systems, and clinicians distinguish technical feasibility from clinical readiness and identify when additional prospective evaluation, monitoring, or governance is needed.
Ye, Z., Chen, Y., Huang, X. et al. A five-phase evaluation framework for diagnostic and predictive medical artificial intelligence. npj Digit. Med. 9, 678 (2026).
[2026/9] This study proposes a five-phase evaluation framework for medical AI, supported by a dynamic evaluation architecture reflecting the nonlinear, iterative nature of AI systems. The framework integrates technical validation, operational robustness validation, controlled interaction validation, clinical evidence validation, and real-world integration validation, while incorporating phase-gating criteria and local and systemic fall-back triggers. These mechanisms enable re-entry into earlier phases based on drift, version updates, or safety signals, and accommodate parallel activities such as implementation research informing clinical trials. By systematically mapping multicenter external validation, shadow-mode testing, human–AI comparison and cooperation studies, randomized controlled trials, real-world evaluations, and adaptive designs into a coherent lifecycle pathway, the framework addresses persistent gaps between laboratory performance and clinical benefit. It provides researchers, clinical institutions, and regulators with an operational, scalable approach aligned with evolving regulatory expectations, supporting trustworthy, ethically aligned, and lifecycle-based evidence generation for medical AI systems.
Lee E, Naser JA, Kane CJ, et al. Artificial Intelligence–Enabled Acquisition and Interpretation for Screening Aortic Stenosis. JAMA Cardiol. Published online August 28, 2026. doi:10.1001/jamacardio.2026.3829
[2026/8] In this diagnostic study, a deep learning echocardiography algorithm demonstrated high diagnostic performance across validation cohorts and maintained high diagnostic accuracy during prospective evaluation using AI-guided focused cardiac ultrasound examinations acquired by novice operators. A 2-step workflow incorporating expert review of AI-positive and uninterpretable examinations substantially improved positive predictive value while requiring review of only approximately 10% of examinations. AI-guided focused cardiac ultrasound combined with automated interpretation may provide a scalable triage and possibly even screening strategy for moderate or greater aortic stenosis in settings with limited access to comprehensive echocardiography.
Prelaj A, Miskovic V, Sacco M, et al. Clinical usability of an explainable AI decision support tool and evaluation of multimodal models in NSCLC. Nat Med. 2026 Sep;32(9):3235-3247. doi: 10.1038/s41591-026-04488-2.
[2026/9] Despite a decade in, immunotherapy (IO) treatment selection in non-small cell lung cancer (NSCLC) remains largely guided by subgroup analyses and imperfect programmed death ligand 1 (PD-L1) and clinical scores. We integrated real-world clinical and blood (CB) data, computed tomography (CT) images, digital pathology (DP), and genomics into machine learning early fusion (MLEF) and deep learning intermediate fusion (DLIF) models. Machine learning (ML) and deep learning (DL) CB-only models achieved consistent performance across outcomes with area under the curve (AUC) up to 0.77 in the test (TEST) set. Performance drop in external validation (EXVAL) likely reflects population differences (AUC range: 0.55–0.72). AI models significantly surpassed PD-L1, Eastern Cooperative Oncology Group performance status (ECOG PS), neutrophil-to-lymphocyte ratio (NLR), lactate dehydrogenase (LDH) and Lung Immune Prognostic Index (LIPI) score in the independent TEST set. The clinical usability study showed that lung expert and nonexpert physicians improved their prediction with the explainable AI (XAI) ML CB-only based tool. Our analysis demonstrated that the AI-based predictive models using only routine CB data outperformed single biomarkers (PD-L1, ECOG PS, NLR and LDH) and the composite LIPI score.
Li, S., Yan, C., Ong, J.C.L. et al. Enabling equitable global health AI with privacy‑enhancing technologies. npj Digit. Med. (2026).
[2026/9] The global AI divide in healthcare is widening, with low-resource settings underrepresented in training data and underserved by infrastructure. Privacy-enhancing technologies (PETs) offer a path forward by enabling collaborative model development without centralizing sensitive data, but many remain difficult to implement in resource-constrained settings. In this Perspective, we examine the PET landscape, assess feasibility, address emerging challenges with large language models, and propose four priorities for equitable implementation.
~

In the Eyes of GPTs
(AJ in conversation with ChatGPT to understand how machines view human healthcare)
As physicians increasingly use AI decision support in routine clinical care, could they gradually become over-reliant on AI? How might this dependence develop, and would clinicians recognize when it is happening?
ChatGPT
Yes. I think this is likely to become one of the central human-factors safety issues of clinical AI. It is related to the NEJM AI concern about scientific thinking, but in clinical care the mechanism is even more direct: repeated exposure to useful AI can gradually change how clinicians allocate attention, effort, and responsibility for reasoning.
A useful way to think about the progression is:
AI assistance → demonstrated usefulness → increasing trust → cognitive offloading → reduced independent verification → automation bias → over-reliance.
Initially, clinicians may use AI appropriately as a second opinion. If it repeatedly identifies diagnoses they missed, summarizes complex records well, or recommends reasonable treatments, clinicians learn that consulting it saves time and cognitive effort. Under routine clinical pressures—high workload, limited time, information overload—the rational temptation is to let AI perform more of the initial reasoning.
The important point is that over-reliance does not require clinicians to believe AI is infallible. A physician can fully understand that an LLM makes mistakes and still become behaviorally dependent on it.
Several mechanisms could drive this:
Would clinicians know that they are becoming over-reliant?
This is the especially interesting question. Probably not reliably.
Clinicians can recognize obvious dependence—“I always ask the AI before deciding”—but the more important cognitive changes may be difficult to perceive. Over-reliance is partly a metacognitive problem: the clinician has to recognize that the process by which they are reasoning has changed.
Consider two physicians who both ultimately make the correct diagnosis:
Physician A: independently develops a differential → reviews AI → incorporates useful suggestions → makes final judgment.
Physician B: reads the AI diagnosis first → evaluates the patient largely within that frame → agrees with AI → makes the same final judgment.
Their observable clinical decision may be identical, but their cognitive processes are very different. Physician B may sincerely feel that they exercised independent judgment because they reviewed and approved the recommendation. Yet their reasoning was already anchored by the AI.
This becomes particularly important for medical students and residents. Experienced physicians already possess internal disease representations developed through years of independent reasoning. Trainees are still constructing them. If AI supplies the differential diagnosis, interpretation, and management plan before the learner struggles with the case, the learner may acquire knowledge while getting fewer opportunities to develop diagnostic reasoning itself.
That suggests an important distinction for your newsletter discussion:
The safety question is not only “Does AI make clinicians choose the wrong answer?”
It is also:
“Does long-term use of AI change clinicians' ability to reach a good answer without AI?”
Those are fundamentally different endpoints. Most current clinical AI evaluations measure the first; comparatively little prospective evidence addresses the second.
This also suggests a strong LHS research direction. We could measure not only AI-assisted clinical performance, but also human cognitive dependence over time: frequency of independent differential formation before AI exposure, acceptance of intentionally incorrect AI suggestions, ability to detect AI errors, changes in diagnostic breadth, confidence calibration, and performance on matched cases when AI is temporarily unavailable.
In other words, clinical AI safety may eventually need two kinds of surveillance: surveillance of the AI and surveillance of the human–AI system. The model can remain technically unchanged while clinician behavior around it changes substantially. That human adaptation may be one of the most important—and least visible—long-term effects of embedding AI into routine care.
How can we prevent AI from negatively affecting physicians’ independent clinical reasoning and cognitive skills? What specific actions should medical educators, healthcare organizations, AI developers, and clinicians take?
ChatGPT
Yes. The goal should not be to reduce AI use, but to design AI use so that clinicians continue doing the cognitive work that matters. The most effective interventions are likely to operate at the workflow, interface, education, and monitoring levels.
A useful principle is:
Think first → ask AI → compare → challenge → decide.
Specific actions could include:
This could become a design principle for ELHS
I think this fits particularly well with a Learning Health System. Instead of treating cognitive safety as a one-time AI usability question, it can become a continuous learning outcome:
Independent physician decision → AI recommendation → final physician decision → clinical outcome → longitudinal learning.
That structure provides three very informative observations:
Before AI: What did the clinician think independently?
After AI: How did AI change the clinician's thinking?
Outcome: Was that change beneficial or harmful?
From these data, you could distinguish appropriate reliance from over-reliance. For example, changing an initially wrong diagnosis after receiving a correct AI suggestion is beneficial reliance. Changing an initially correct diagnosis because of an incorrect AI suggestion is a particularly important cognitive-safety event.
This leads to what I think could be a powerful concept for your newsletter:
The solution to AI over-reliance is not simply telling physicians, “Use your own judgment.” Clinical AI should be designed to preserve independent judgment.
And perhaps an even broader principle:
A successful clinical AI system should improve clinicians' decisions when AI is present without degrading their ability to reason when AI is absent.
That gives you a measurable definition of cognitive safety that goes considerably beyond conventional model accuracy and could be tested prospectively.
Should governments regulate AI research and applications when they pose significant risks to individuals or society? What can we learn from the regulation of other powerful technologies, such as human cloning?
ChatGPT
There is a strong case for government regulation of high-risk AI applications and AI research involving people, especially in medicine. But the analogy to human cloning is useful mainly for illustrating the need for societal guardrails—not for treating AI as one uniformly prohibited technology.
Human cloning presents a relatively identifiable intervention that societies can prohibit or tightly constrain. AI is a general-purpose technology spanning low-risk activities such as drafting text through high-stakes activities such as diagnosis, treatment selection, autonomous medical decisions, biological research, and critical infrastructure. A single “regulate AI” rule would therefore be too crude.
A better model is risk-based, lifecycle regulation:
Low risk → freedom to innovate
Increasing risk → increasing evidence and oversight
Potentially catastrophic or irreversible risk → strong restrictions or prohibition
Healthcare already illustrates this approach. The FDA regulates qualifying AI-enabled medical-device functions while distinguishing some clinical decision-support software that falls outside the device definition. WHO similarly argues that health AI should protect human autonomy, safety, transparency, accountability, inclusion and equity, with humans retaining control over medical decisions.
Research and application should also be treated differently
AI research itself should generally remain open, because excessive restrictions on algorithms, model development, or basic research could impede beneficial discovery. But additional oversight becomes justified when research crosses certain boundaries: experiments involving patients, identifiable health data, deployment into clinical workflows, manipulation of human behavior, dangerous biological capabilities, or other substantial risks to people.
This distinction already exists in biomedical research. We do not require government permission to formulate a biological hypothesis, but research involving human participants receives ethics oversight. WHO's September 2026 report specifically argues that existing ethics mechanisms may not adequately cover novel AI risks and recommends strengthening oversight throughout the AI health-research lifecycle.
Your cognitive-safety question exposes an important regulatory gap
Traditional medical-device regulation tends to ask questions such as:
Is the AI accurate? Is it safe? Does it improve clinical performance?
But increasingly we may also need to ask:
What happens to clinicians after using it for five years?
Suppose an AI decision-support system improves diagnostic accuracy from 85% to 92%. That looks beneficial. But imagine that after several years of routine use, physicians become increasingly dependent on it, their unaided diagnostic performance declines, trainees fail to develop some reasoning skills, and clinicians become less capable of recognizing the system's rare but consequential errors.
The AI itself might still pass every conventional accuracy test.
The human–AI system has nevertheless changed.
WHO's existing governance framework already recognizes the underlying concern: transferring decision-making power to machines can undermine human autonomy, and healthcare professionals need sufficient information to use AI safely and effectively. But long-term cognitive effects on clinicians deserve much more explicit study.
This suggests that regulation should not simply say “a human must remain in the loop.” A physician clicking Accept does not necessarily constitute meaningful human oversight.
A stronger standard might eventually require evidence that clinical AI:
That last point is especially important.
Regulation should therefore protect human capability, not just humans from erroneous machines
The cloning analogy ultimately leads somewhere interesting. Society regulates some biomedical technologies because technical feasibility alone does not determine whether or how they should be used. AI raises the same principle, but across a much broader spectrum.
For healthcare AI, I would frame the emerging question this way:
Government should not regulate AI simply because it is AI. Regulation should follow the magnitude and reversibility of potential harm—and should evaluate the human–AI system, not only the algorithm.
And your discussion about medical education adds another dimension. We may eventually need to recognize preservation of human cognitive capability as an AI safety outcome.
That would connect the September newsletter papers very naturally:
AI safety → prospective clinical evidence → human-AI interaction → cognitive safety → responsible adoption.
The recent WHO report is particularly timely for your newsletter because it was released September 21, 2026 and explicitly calls for stronger ethics oversight of AI-related health research.
~
🔹 ELHS Institute🔹
Democratizing GenAI and LHS to Advance Global Health Equity
▶️ ELHS Videos
👉 For Clinical AI technology support, contact us at support@elhsi.org 📩
~ the end ~
Democratizing GenAI and LHS to Advance Global Health Equity
info@elhsi.org
Palo Alto, California, USA
