Newsletters

Your E-mail *

ELHS Newsletter | August 2026

Beyond AI vs. Doctors: Building the Evidence for the Future of Medicine

Subscribe

 

 

Dear Friends,

This month, several important studies highlight how rapidly healthcare AI is moving toward real-world applications. LLMs are being tested for hospital referral triage, transforming unstructured clinical notes into structured data, and automating complex EHR analyses. New AI models are improving prediction of treatment outcomes across diseases, while federated approaches are making it possible for medical AI to learn across healthcare organizations without sharing patient data.

 

 

At the same time, an influential JAMA Perspective argues that autonomous AI may eventually provide better cognitive medical care than either physicians alone or physicians assisted by AI. This provocative prediction has generated considerable debate (see Dr. Zeke Emanuel vs AMA CEO Dr. John Whyte debate video).

I believe it is still too early to draw such a broad conclusion. Most available evidence comes from benchmarks, retrospective studies, and simulated clinical tasks, while prospective evidence from routine clinical care remains limited. Rather than asking simply whether AI or humans will be better, we need real-world evidence to determine which clinical tasks should be performed by AI, physicians, or human-AI teams—and under what conditions.

A second JAMA Perspective raises another important question: Can AI radically shorten the journey from scientific breakthrough to implementation at scale?

I believe this may become one of the defining questions for Learning Health Systems. As AI rapidly gains capabilities to identify patients, predict outcomes, structure clinical data, and analyze real-world evidence, we need healthcare systems that can continuously evaluate these capabilities, integrate those that work into routine care, measure their impact, and learn from every patient.

The next breakthrough in healthcare AI may therefore be not AI alone, but AI-native Learning Health Systems that can translate rapidly advancing AI into better care—faster and at scale.

I hope you enjoy this month’s papers and discussion.

 

Best regards,

AJ

AJ Chen, PhD
Founder & PI, ELHS Institute
Silicon Valley, USA

https://elhsi.org/Newsletters
https://elhsi.com

 

~

 

From Page Mill

(Recent papers, news, and events showcasing the progress of GenAI and LHS) 

Emanuel EJ, Baker-Butler A, Khosla N, Khosla V. Will Autonomous AI Exceed AI-Aided Physicians as the Best Medical Care? JAMA. Published online August 17, 2026. doi:10.1001/jama.2026.15380

[2028/6] In cognitive medical functions, AI-alone medical care is likely to be better than physician-only or physician-AI hybrid care. Data from medicine and other fields suggest that when AI-alone performance is consistently superior to human-alone performance, AI alone surpasses human-AI hybrids. Paradoxically, hybrid care in which humans are in (or on) the loop to correct AI errors is likely to worsen rather than improve AI performance. Review of all published articles on AI in medicine since January 1, 2024, shows that medicine is rapidly approaching the transition point at which AI alone will exceed physicians and physician-AI hybrids in providing the best care at 5 fundamental cognitive medical tasks. Even when AI alone is superior, significant barriers to implementation remain. Nonetheless, superior autonomous AI will likely be ready to be deployed for real-world cognitive medical tasks in some, maybe many, workflows by 2030. Consequently, physicians, policymakers, and others need to urgently devise approaches to workflow, liability, regulation, reimbursement, and medical education.

Beckman AL, Chokshi DA. From Breakthrough to Follow-Through—A Public Health Agenda for AI. JAMA. Published online August 13, 2026. doi:10.1001/jama.2026.16748

[2026/8] Could generative artificial intelligence (AI) help with radically compressing the time frame from breakthrough to follow-through? Artificial intelligence offers the ability to automate case finding, augment outreach, and assist with patient navigation. If directed to support longitudinal human relationships, AI might meaningfully benefit the public’s health. Realizing this vision is not easy, however, and requires the union of 4 components: government leadership, technology, cross-sector partnerships, and capital.

Kang, B., Son, M., Jeong, W. et al. Large language model-assisted referral triage automation in a tertiary hospital. npj Digit. Med. (2026).

[202/8] Manual referral triage in tertiary hospitals is resource-intensive and prone to inconsistency. This study evaluates the feasibility of an on-premises large language model (LLM), using a roster-embedded prompting strategy, to assist in subspecialty triage of referral letters. Utilizing a dataset of 6624 electronic referral letters from Samsung Medical Center, we deployed an open-source LLM (Qwen-2.5-32B) within a secure infrastructure, incorporating real-time clinician availability. In a hold-out test set (n = 680), the LLM achieved a baseline accuracy of 75.4% (95% CI, 72.278.7%) compared to human coordinators, which improved to 84.7% (95% CI, 81.987.4%) after expert adjudication of discordant cases. Performance of frequently referred subspecialties was accuracy of 86.1%, and performance of infrequent ones was 74.4%, which were adjudicated. Error analysis revealed that misclassifications primarily occurred between clinically adjacent departments rather than at random, and a small subset of referrals (5.9%) lacked sufficient information for unique assignment. These findings demonstrate that our roster-embedded LLM can provide highly valid subspecialty assignments and serve a complementary role in referral workflows. By integrating such models into human-in-the-loop systems, tertiary hospitals may significantly enhance operational efficiency and reduce administrative burdens.

Shen W, Moon I, Nguyen TH, Li MM, Huang Y, Nair N, Marbach D, Zitnik M. Generalizable AI predicts immunotherapy outcomes across cancers and treatments. Nat Med. 2026 Aug;32(8):3010-3022. doi:10.1038/s41591-026-04502-7.

[2026/8] Here we present COMPASS, a pan-cancer foundation model that predicts immunotherapy response from bulk tumor transcriptomes using a concept bottleneck transformer. COMPASS encodes gene expression through 44 biologically grounded immune concepts representing immune cell states, tumormicroenvironment interaction and signaling pathways. Trained on 10,184 tumors across 33 cancer types, COMPASS achieves better average performance than 22 methods across 16 clinical cohorts spanning seven cancers and six ICIs, improving accuracy by 8.5% and area under the precision-recall curve by 15.7% on average across cohorts. COMPASS generalizes to cancer types and treatments not represented during fine-tuning and may inform indication selection and patient stratification. In survival analyses, patients classified by COMPASS as responders had longer overall survival (hazard ratio = 4.7, P < 0.0001).

Knorr, J.M., Patel, S., Abusafieh, H.T. et al. The beginning of the end for manual chart review: LLM-mediated database constructionnpj Digit. Med. (2026). 

[2026/8] Manual clinical data abstraction is the reference standard for research databases but is labor-intensive, costly, and susceptible to human error. We evaluated the accuracy of a locally deployed open-source large language model (LLM) framework for automated extraction of structured kidney cancer data from unstructured clinical documentation. In this retrospective study, 8366 patients undergoing nephrectomy for suspected kidney or upper tract urothelial cancer between 2009 and 2024 were identified from a large academic health system. The LLM processed 130,509 clinical notes to generate 136,425 data elements across 14 operative, pathology, and radiology variables. Overall agreement with a manually curated reference database was 97.5%, with pathologic variables exceeding 98% agreement and Cohen’s κ values > 0.90. Manual arbitration of disagreements (2.5%) favored the LLM for 10 of 14 variables. These findings demonstrate that locally deployed open-source LLMs can accurately and efficiently generate large, structured clinical databases while preserving institutional data governance.

Zhu, Y., Wang, Z., Qi, Y. et al. HealthFlow: automating electronic health record analysis via a strategically self-evolving multi-agent frameworknpj Digit. Med. 9, 660 (2026).

[2026/8] Electronic health records (EHRs) are a rich source of real-world clinical data, but turning them into valid analyses remains slow, brittle, and expert-intensive. Although recent AI agents can answer medical questions and use tools, automating full EHR workflows remains difficult because small planning or execution errors can invalidate otherwise plausible analyses. Here we present HealthFlow, a multi-agent framework that converts prior EHR analyses into structured, governed experience for planning under dataset-specific and methodological constraints. We also introduce EHRFlowBench, a rigorously curated benchmark of EHR analysis tasks derived from 51,280 peer-reviewed papers. Across EHRFlowBench and four established benchmarks (MedAgentBoard, MedAgentsBench, HLE, and CureBench), HealthFlow consistently outperforms strong baselines in generating valid clinical artifacts and completing complex EHR analysis pipelines. These results show that governed reuse of prior analytical experience improves robustness in health data science and opens a promising path to automating open-ended EHR analysis.

Li, A., Chen, Y., Long, W. et al. Toward federated large language models in medicine: a parameter-efficient framework for privacy-preserving, multi-institutional adaptationnpj Digit. Med. (2026). 

[2026/8] Large language models (LLMs) are increasingly adapted for medical applications, but most are trained using data from a single institution because privacy and governance constraints prevent multi-institutional data sharing. As a result, these models often generalize poorly across heterogeneous healthcare systems. We address this gap by introducing Fed-MedLoRA and Fed-MedLoRA + , a parameter-efficient federated framework for collaborative LLM adaptation across healthcare institutions. Fed-MedLoRA transmits only low-rank adapters rather than full model weights, reducing communication overhead. We also evaluate a privacy-preserving variant that applies Gaussian perturbation to transmitted adapter updates. Fed-MedLoRA+ further incorporates adaptive aggregation to better address cross-site heterogeneity in patient populations, annotation practices, and disease distributions. We evaluate the framework on clinical information extraction across five independent patient cohorts totaling 42,198 entities and 41,570 relations, and compare it with zero-shot and fine-tuned LLMs, domain-specific BERT models, and federated baselines. Across all settings, the proposed methods consistently improve extraction performance and generalize better to heterogeneous cohorts. In a real-world case study using clinical notes from the Yale New Haven Health System, the framework demonstrates strong performance under low-resource new-site deployment. These results suggest that federated, parameter-efficient LLM adaptation is feasible, scalable, and effective for multi-institutional clinical deployment.

 

~

 

 

In the Eyes of GPTs

(AJ in conversation with ChatGPT to understand how machines view human healthcare)

 

User  The debate between Dr. Zeke Emanuel and AMA CEO Dr. John Whyte is available here: https://www.youtube.com/watch?v=WTnZ6W7qsLQ. Please summarize their key arguments separately.

 

A white logo on a purple background

Description automatically generated ChatGPT

The debate is useful because Emanuel and AMA CEO John Whyte are not really debating whether AI is valuable—they both think it is. Their disagreement is primarily about what the existing evidence justifies and whether physicians should remain in the loop.

Ezekiel Emanuel: Autonomous AI will eventually outperform physician-led care

Emanuel’s central argument is that medicine is following a pattern already seen in chess and autonomous driving: humans initially outperform AI, then human-AI teams perform best, but as AI improves further, autonomous AI ultimately surpasses both. He argues that recent medical AI research is already approaching this transition for cognitive tasks such as eliciting patient information, generating differential diagnoses, ordering tests, delivering guideline-concordant care, and managing chronic disease.

His main arguments are:

  • Recent evidence increasingly favors AI. Emanuel says his group reviewed roughly 100 studies published since January 2024 and that the large majority showed AI outperforming physicians and physician-AI hybrids. He rejects older or “out-of-the-box” chatbot studies as poor representations of what appropriately configured modern AI can do.
  • AI should be compared with actual physician performance, not idealized medicine. Physicians also make errors, including misdiagnoses. Therefore, the relevant safety question is not whether AI makes mistakes, but whether AI makes fewer or less consequential mistakes than physicians.
  • Autonomous does not mean unregulated. Emanuel envisions specialized AI with training, reasoning capabilities, guardrails, defined scopes of practice, real-patient validation, licensing, periodic reassessment, and liability frameworks—not simply giving patients an off-the-shelf chatbot.
  • Physicians would not disappear, but their role would change. AI could take over substantial portions of primary care, triage, and chronic-disease management, freeing physicians to concentrate on patients and interactions requiring human expertise. This would also require medical education to place more emphasis on EQ and human relationships.
  • Autonomous AI could expand access. Emanuel argues that underserved, uninsured, poor, and rural populations may be among the first to benefit because autonomous AI could provide care where physicians are scarce.

In short, Emanuel sees physician oversight as potentially becoming a limitation rather than an inherent safety feature once AI itself becomes sufficiently capable.

John Whyte / AMA: AI should augment physicians, not replace them

Whyte’s central objection is that Emanuel draws conclusions that are broader than the evidence supports. He emphasizes that much of the literature cited involves simulated cases, benchmarks, or discrete cognitive tasks rather than actual clinical practice. Performing well on those tasks does not demonstrate that autonomous AI provides better medical care overall.

His main arguments are:

  • The evidence base is not sufficiently real-world. Many studies use simulations, standardized cases, or other artificial comparisons. Whyte argues that better performance on a board-style problem or diagnostic task cannot simply be extrapolated to better medical practice.
  • Medicine is more than cognitive task performance. Clinical care involves trust, judgment, recognizing when the apparent diagnosis does not fit, communicating uncertainty, accountability, and the physician-patient relationship. Reducing medicine to measurable cognitive tasks misses this “art and science” dimension.
  • Human oversight remains necessary. Whyte supports the AMA concept of “augmented intelligence”: AI should strengthen a physician-led care team rather than independently practice medicine. He would retain physician oversight even for seemingly routine activities such as medication refills.
  • AI can be extremely valuable without autonomy. He specifically highlights AI-assisted cancer detection, personalized treatment, and AI-supported second opinions as promising applications. The important objective is improving what clinicians can accomplish, not removing clinicians from the process.
  • Patient use of LLMs can be beneficial, but should connect back to clinicians. He accepts that patients already use AI to understand laboratory results, consider possible diagnoses, and prepare questions, but believes most patients still value dialogue and a relationship with physicians.
  • Safety evidence is mixed in sensitive areas. He points particularly to mental health, false negatives, and patients' difficulty interpreting probabilities as reasons not to assume that good benchmark performance translates into safe autonomous care.
  • He nevertheless agrees that governance must evolve. Whyte and Emanuel find common ground around the need for appropriate benchmarks, licensing, regulation, liability, and reimbursement frameworks.

The fundamental disagreement

I think the debate can be distilled into one sentence:

Emanuel asks: “If AI performs a medical task better than physicians, why should requiring physician involvement automatically make care safer?” Whyte asks: “How can we conclude AI provides better medical care when most studies have not actually tested autonomous AI in real-world medical care?”

That second question is especially relevant to the point you made in the newsletter. The disagreement is partly about different levels of evidence. Emanuel extrapolates from rapidly improving performance on individual cognitive functions toward a future model of medical care; Whyte argues that performance on those functions cannot yet establish superiority of autonomous AI at the system-of-care level.

Interestingly, both ultimately converge on something close to an LHS research agenda: deploy carefully, evaluate on real patients, define appropriate scopes of practice, monitor performance and safety, periodically reassess the technology, and learn as AI changes. Emanuel himself explicitly argues that test performance is “grossly insufficient” and that AI needs evaluation on hundreds of actual patient cases plus regular reassessment.

That convergence could be a particularly strong angle for your later newsletter discussion: before declaring a winner in “AI vs physician,” healthcare needs the infrastructure to continuously generate the real-world evidence capable of answering the question.

 

User  You are right—what is urgent now is to accelerate evidence generation across diseases and clinical settings. While this kind of debate is helpful, I believe that new evidence will ultimately guide us toward the future of medicine that best serves humanity.

In addition, do you think their debate missed a crucial element—preserving human values?

 

A white logo on a purple background

Description automatically generated ChatGPT

Yes. I think preserving human values is a crucial dimension that the debate touched on but did not fully address.

Whyte came closest when he emphasized trust, the physician-patient relationship, empathy, accountability, and the “art and science” of medicine. Emanuel responded that medicine has historically rewarded scientific competence more than empathy and argued that AI could allow physicians to concentrate on the human dimensions of care. But both largely framed the debate around who performs particular medical functions better and whether humans should remain in the loop.

There is a deeper question:

Even if autonomous AI eventually becomes better than physicians at certain cognitive tasks, how do we ensure that future medicine continues to reflect human values?

Medicine is not purely an optimization problem. Some decisions involve values that cannot be determined simply by maximizing diagnostic accuracy or predicted outcomes. What constitutes an acceptable tradeoff between longevity and quality of life? How much uncertainty or treatment burden is acceptable? When should aggressive treatment give way to comfort? What risks is a particular patient willing to accept? How should dignity, autonomy, fairness, compassion, cultural preferences, and family considerations influence a decision?

AI can help humans reason about these questions, but the values being optimized ultimately need to come from humans.

This also changes how I would frame the three positions. It isn't simply:

AI alone vs. physician alone vs. physician + AI.

There are actually two different questions:

1. What can AI do better?
This is an empirical question. We need much more real-world evidence across diseases, populations, clinical settings, and tasks. The answer will evolve rapidly as AI improves.

2. What kind of medicine do humans want?
This is a values question. Better predictive performance alone cannot answer it.

And I think your point about accelerating evidence generation connects these two beautifully. We shouldn't try to settle the future of medicine today through arguments based on an immature evidence base. We should build the capability to generate evidence much faster, across the full spectrum of medicine, and allow that evidence—together with human values—to guide how the roles of AI and clinicians evolve.

That could become a strong central thesis for your later commentary:

The urgent question is not whether AI or physicians will ultimately “win.” It is how quickly we can generate rigorous real-world evidence to discover the best roles for AI and humans across medicine—while ensuring that the healthcare system we build continues to serve human values.

There is also a deeper connection to LHS here. An AI-native Learning Health System should not merely optimize AI performance. It should continuously learn whether care is becoming more accurate, earlier, safer, more equitable, more accessible, and more aligned with what patients value.

In that sense, perhaps the destination isn't autonomous medicine or even human-AI medicine. It is human-centered medicine in which the respective roles of humans and AI are continuously determined by evidence and guided by human values.

I think that is a more durable framing than taking either side of the Emanuel–Whyte debate.

 

 

~

 

🔹 ELHS Institute🔹

Democratizing GenAI and LHS to Advance Global Health Equity

🤖 ELHS GenAI Platform

🌐 ELHS Institute Website

📰 ELHS Monthly Newsletters

▶️ ELHS Videos

👉 For Clinical AI technology support, contact us at support@elhsi.org 📩

 

~ the end ~