This is a preview of subscription content, access via your institution
Access options
Access Nature and 54 other Nature Portfolio journals
Get Nature+, our best-value online-access subscription
$32.99 / 30 days
cancel any time
Subscribe to this journal
Receive 12 print issues and online access
$259.00 per year
only $21.58 per issue
Buy this article
- Purchase on SpringerLink
- Instant access to the full article PDF.
USD 39.95
Prices may be subject to local taxes which are calculated during checkout
References
Beaulieu-Jones, B. & Nemati, S. Limited benchmarks constrain the conclusions of a general-purpose versus clinical AI comparison. Nat. Med. https://doi.org/10.1038/s41591-026-04638-6 (2026).
Vishwanath, K. et al. General-purpose large language models outperform specialized clinical AI tools on medical benchmarks. Nat. Med. 32, 2405–2409 https://doi.org/10.1038/s41591-026-04431-5 (2026).
Funding
E.K.O. is supported by the National Cancer Institute’s Early-Stage Surgeon Scientist Program (3P30CA016087-41S1) and the W.M. Keck Foundation. This work was supported by a grant from the Institute for Information & Communications Technology Planning and Evaluation (IITP) funded by the Ministry of Science and ICT (MSIT) of the Republic of Korea government (no. RS-2019-II190075 Artificial Intelligence Graduate School Program (KAIST); no. RS-2024-00509279, Global AI Frontier Lab). The funders had no role in study design, data collection and analysis, decision to publish or preparation of the manuscript.
Author information
Authors and Affiliations
Contributions
K.V., Y.A. and E.K.O., jointly supervised the study, conceptualized the design and wrote the initial draft. K.V. performed the statistical analyses and experiments. All authors reviewed and approved the final paper.
Corresponding authors
Ethics declarations
Competing interests
E.K.O. reports equity in MarchAI and Artisight, spousal employment by Eikon Therapeutics and consulting for Sofinnova Partners, Google and Alphatec Holdings. The remaining authors declare no competing interests.
Additional information
Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Rights and permissions
About this article
Cite this article
Vishwanath, K., Aphinyanaphongs, Y. & Oermann, E.K. Reply to: Limited benchmarks constrain the conclusions of a general-purpose versus clinical AI comparison. Nat Med (2026). https://doi.org/10.1038/s41591-026-04637-7
Received:
Accepted:
Published:
Version of record:
DOI: https://doi.org/10.1038/s41591-026-04637-7
Facts Only
* Limited benchmarks constrain the conclusions of a general-purpose versus clinical AI comparison.
* General-purpose large language models outperform specialized clinical AI tools on medical benchmarks.
* The study was published in Nature Medicine (2026).
* The research was supported by the National Cancer Institute’s Early-Stage Surgeon Scientist Program and the W.M. Keck Foundation.
* The work received funding from the Ministry of Science and ICT of the Republic of Korea government via the IITP.
* K.V. performed statistical analyses and experiments.
Executive Summary
Full Take
Sentinel — Human
This text appears to be the standard metadata and reference structure for a published academic paper, exhibiting characteristics consistent with legitimate scholarly output rather than synthetic content.