AIe2501008
IA & Tecnologia · Central de Estudos
🤖 IA & Tecnologia

AIe2501008

IA & Tecnologia 📄 Resumo de estudo 🏥 UTI / Emergência
§ 01

Visão geral

Figura
Fig. Figura

NEJM AI 2025;2(11)

EDITORIAL

Humanity’s Next Medical Exam: Preparing to Evaluate Superhuman Systems

Jack Gallifant

Figura
Fig. Figura

M.B.B.S.,1,2 and Danielle S. Bitterman

Figura
Fig. Figura

M.D.1,2

Received: September 4, 2025; Accepted: September 17, 2025; Published: October 23, 2025

Abstract

The rapid advances in health care AI necessitate a fundamental shift in how we evaluate these systems. Palepu et al. (2025) demonstrate that AI can outperform medical trainees in breast cancer management questions, illustrating that advances that would have been diffi­cult to foresee only a few years ago are imminent. However, current methods, often reliant on question and answer tasks, are inadequate for capturing the nuances of clinical practice, even as models begin to exceed human performance on these narrow metrics. The trajectory of AI development toward more generalized and autonomous systems introduces profound opportunities alongside substantial risks, making the limitations of our existing oversight frameworks an urgent problem. In this editorial, we propose Humanity’s Next Medical Exam, a novel approach designed to measure and promote the development of safe, human-aligned AI. This paradigm is built upon three foundational pillars: interactive interrogation to challenge models beyond rote knowledge, experiential learning in sandbox environments to assess decision-making under uncertainty, and real-world continuous learning to monitor and refine performance postdeployment. The maturation of these core components as part of a comprehensive evaluation framework is a critical step toward helping us prepare to take advantage and avoid risks of continually advancing AI technologies in the future. (Funded by the National Institutes of Health, the National Cancer Institute, and others.)

R and the architectural shift toward agentic self-reflection are genuine breakthroughs. gressed from simple fact recall to complex reasoning, collaboration, and tool use.1-4 The ability of these systems to reason through complex medical guidelines apid advances in health care AI warrant reflection. In a few years, AI has pro­

The case study by Palepu et al. (2025) highlights this progress.5 They evaluated the per­

formance of a proprietary, large language model (LLM)–based system in providing clinical The author affiliations are listed at

recommendations on 60 brief, synthetic breast cancer presentations. Their system outper­ the end of the article.

formed medical trainees with less than 1 year of oncology-specific training, but not attending Danielle S. Boncologists, in clinical management reasoning tasks. Rates of harm in the generated output were low. The system is an adaptation of Google’s Articulate Medical Intelligence Explorer, ent

or AMIE, based on Gemini 2.5 Pro, with access to Internet search and self-reflection. While of Radiation Oncology, Dana- Farber Cancer Institute/Brigham the authors acknowledge that the results fall short of immediate clinical applicability, they and Women’s Hospital, 75 Francis are illustrative of advances that would have been difficult to foresee only 3 years ago. Street, Boston, MA 02115.

NEJM AI is produced by NEJM Group, a division of the Massachusetts Medical Society.

Downloaded from ai.nejm.org by Fabio Ynoe de Moraes on November 27, 2025. For personal use only. No other uses without permission. Copyright © 2025 Massachusetts Medical Society. All rights reserved.

The clinical community has responded to this progress with measured caution, and rightly so. The promising reports have not led to a clinical consensus that such AI systems are true clinical reasoners. This caution stems from a deep understanding of clinical reality, not disbelief in study results. This is why passing an exam is insufficient to pro­duce a competent clinician and is not the sole requirement for licensure. Medical training builds knowledge that gen­eralizes to the complex reality of clinical care; its goal is to develop intuition and problem-solving skills, not rote mem­orization. The physician’s role is more than a sequence of questions and answers, and decisions often have no single “right” answer, creating a vast chasm between benchmarks and the bedside. Direct patient contact may also account for only 40% of a physician’s shift.6 The rest is spent navigating bureaucracy, reviewing disorganized data, coordinating care, speaking with families, and reacting to misaligned system incentives.

Evaluation and safety should evolve to keep pace with the exponential progress made in just the last 18 months. As we and others have argued, AI evaluations should test more realistic clinical tasks and meaningful outcomes.7,8 There are even more foundational concerns, however, with the nature of current evaluations. When scoped to simple question and answer tasks, systems today have the poten­tial to achieve superhuman performance, which refers to the progression of technology toward autonomous abil­ities exceeding human performance. Although the term superhuman has futuristic and provocative connotations, strictly defined superhuman performance on a controlled, measurable task is currently achievable. Soon, superhuman performance may extend to real-world settings where we anticipate and even desire experiential learning and devel­opment. Even with today’s narrow implementations, many hospitals are already grappling with the fact that there is no widely accepted or validated way to measure such implementations in the real world. As systems continue to progress, the risks of being unprepared for increasingly superhuman capabilities demand a fundamental transfor­mation in the conceptual and methodological frameworks underpinning evaluation.

These new strategies should prepare us for the “era of ­experience” in medical AI, where models evolve from experiential learning.9 For example, plotting attending physician performance as the 100% ceiling limits systems that may in the future exceed human knowledge recall.10-14 The challenge is shifting from whether machines can know as much as or even more than humans to how we measure and govern their ability to act wisely and in alignment with

human values. The opportunity is to leverage this expertise to revolutionize health care cost, quality, and accessibility.

We need a new paradigm: Humanity’s Next Medical Exam. This proposal is inspired by Humanity’s Last Exam, an extremely difficult general knowledge benchmark designed to avoid performance ceilings; the best LLMs score only 25%.15 This medical equivalent would realistically assess clinical skills, identify system gaps, and drive progress. Instead of a written test, we propose a multifaceted strategy for experiential learning and evaluation of real-world care delivery, built on three pillars.

The first pillar is interactive interrogation, resembling a rig­orous, oral board-style exam rather than a multiple-choice test. The question is whether a system can sustain a coher­ent clinical conversation, defend treatment plans, revise recommendations with new data, and reconcile contradic­tions with the fluency of a clinical peer. The focus must shift from knowledge recall to applied reasoning, adaptability, and accountability.

Second, interactive learning in sandbox environments. Early medical AI, trained on static data, was a form of sta­tistical imitation with inherent limits. The next frontier is learning from experience in high-fidelity simulated envi­ronments where AI can act, observe outcomes, and learn from measurable consequences. This approach enables a sustained cycle of advancement and evaluation that evolves in step with capabilities.9 Current reinforcement learning–based fine-tuning offers a robust framework to leverage these interactions, allowing a model to refine its policies beyond what is explicitly stated in a textbook. Such sandboxes could host virtual apprenticeships where models learn clinical trade-offs, such as balancing collecting addi­tional test results against diagnostic speed and the risk of a missed finding, by navigating realistic patient trajectories and goals of care. Though algorithms will evolve, the cen­tral challenge is creating a large, diverse set of high-quality interactive environments. This is the path toward systems that learn from consequences, not just information.

Third, real-world continuous learning. The learning health system, a decades-old promise that has largely failed to materialize, may finally be realized through AI.16 Here, the exam is continuous: Can the system improve with each patient interaction? Can it integrate the latest research for a specific patient or help titrate medication by synthesizing genomic data with patient feedback? This component will push the limits of human oversight toward superhuman performance — precisely the point of a learning health sys­tem. This would build and scale institutional knowledge in

NEJM AI 2

NEJM AI is produced by NEJM Group, a division of the Massachusetts Medical Society.

Downloaded from ai.nejm.org by Fabio Ynoe de Moraes on November 27, 2025. For personal use only. No other uses without permission. Copyright © 2025 Massachusetts Medical Society. All rights reserved.

near real time, making every encounter a systemwide learn­ing opportunity.

Building these pillars requires close collaboration between clinical and engineering communities to answer many open questions. Urgent investment is needed in benchmarking and evaluation to stress test current technologies and pre­pare for future, higher-performing systems.

This assessment requires an AI-ready, modernized health care system; technical advances are insufficient if incen­tives and infrastructure are outdated. This vision depends on secure, near-real-time research sandboxes that sup­port both high-fidelity simulation of historical trajecto­ries and generation of synthetic data for testing novel scenarios. Building these requires overcoming siloed data. Interoperable standards are not enough; we believe stron­ger government regulation is needed to mandate open, standardized application programming interfaces for elec­tronic health records and enforce information sharing. This would prevent vendor lock-in and foster a competitive ecosystem of innovators. Without this modernization, our most potent tools will remain unvetted for real-world use.

Physician skepticism reflects an understanding of systemic and clinical realities, not a fear of technology. Clinicians know their expertise extends beyond medical knowledge to the clinical judgment developed through practice, man­aging atypical patients with differing values, payer deni­als, competing demands, and providing comfort amid uncertainty.

The shift from static information to dynamic interaction marks a new era of experience for medical AI. Success will be contingent on evaluation that keeps pace with and accel­erates model capabilities, paired with a modernized data infrastructure with enforced standards to unlock siloed data. Medicine is practiced in the complex human space between facts, and it is here that AI must now be tested.

necessarily represent the views of PCORI, its Board of Governors, or its Methodology Committee.

Author Affiliations

1 Artificial Intelligence in Medicine (AIM) Program, Mass General Brigham, Harvard Medical School, Boston

2 Department of Radiation Oncology, Brigham and Women’s Hospital and Dana-Farber Cancer Institute, Boston

References

  • Wei J, Wang X, Schuurmans D, et al. Chain-of-thought prompt­ing elicits reasoning in large language models. January 28, 2022 (). Preprint.
  • Kojima T, Gu S, Reid M, Matsuo Y, Iwasawa Y. Large language models are zero-shot reasoners. May 24, 2022 (). Preprint.
  • Yang X, Chen A, PourNejatian N, et al. A large language model for electronic health records. NPJ Digit Med 2022;5:194. DOI: .
  • Singhal K, Azizi S, Tu T, et al. Large language models encode clin­ical knowledge. Nature 2023;620:172-180. DOI: .
  • Palepu A, Dhillon V, Niravath P, et al. Exploring large language models for specialist-level oncology care. NEJM AI 2025;2(11). DOI: .
  • American Medical Association. One driver of resident physi­cian burnout: too little time at bedside. June 28, 2024 ().
  • Kouzy R, Hong JC, Bitterman DS. One shot at trust: building credi­ble evidence for medical artificial intelligence. Lancet Digit Health 2025;7:100883. DOI: .
  • Raji ID, Daneshjou R, Alsentzer E. It’s time to bench the medical exam benchmark. NEJM AI 2025;2(2). DOI: .
  • Silver D, Sutton RS. Welcome to the era of experience. April 26, 2025 (
  • Disclosures ).

    Author disclosures are available at .

    We acknowledge financial support from the National Institutes of Health National Cancer Institute (grant numbers U54CA274516-01A1 [J.G. and D.S.B] and R01CA294033-01 [J.G. and D.S.B.]), the American Cancer Society and American Society for Radiation Oncology (ASTRO-CSDG-24-1244514-01-CTPS, grant DOI: [D.S.B.]), a Patient-Centered Outcomes Research Institute (PCORI) Project Program Award (grant number ME-2024C2-37484 [D.S.B.]), and the Woods Foundation [D.S.B.]. All statements in this report, including its findings and conclusions, are solely those of the authors and do not

  • Goh E, Gallo RJ, Strong E, et al. GPT-4 assistance for improvement of physician performance on patient care tasks: a randomized con­trolled trial. Nat Med 2025;31:1233-1238. DOI: .
  • Goh E, Gallo R, Hom J, et al. Large language model influence on diagnostic reasoning: a randomized clinical trial: a randomized clinical trial. JAMA Netw Open 2024;7:e2440969. DOI: .
  • Arora RK, Wei J, Hicks RS, et al. HealthBench: evaluating large language models towards improved human health. May 13, 2025 (). Preprint.
  • NEJM AI 3

    NEJM AI is produced by NEJM Group, a division of the Massachusetts Medical Society.

    Downloaded from ai.nejm.org by Fabio Ynoe de Moraes on November 27, 2025. For personal use only. No other uses without permission. Copyright © 2025 Massachusetts Medical Society. All rights reserved.

  • Korom R, Kiptinness S, Adan N, et al. AI-based clinical decision support for primary care: a real-world study. July 22, 2025 (). Preprint.
  • Chen S, Guevara M, Moningi S, et al. The effect of using a large lan­guage model to respond to patient messages. Lancet Digit Health 2024;6:e379-e381. DOI: .
  • Phan L, Gatti A, Han Z, et al. Humanity’s last exam. January 24, 2025 (). Preprint.
  • Institute of Medicine, Roundtable on Evidence-Based Medicine. The learning healthcare system: workshop summary. Olsen LA, Aisner D, McGinnis JM, eds. Washington, DC: National Academies Press, 2007. DOI: .
  • NEJM AI 4

    NEJM AI is produced by NEJM Group, a division of the Massachusetts Medical Society.

    Downloaded from ai.nejm.org by Fabio Ynoe de Moraes on November 27, 2025. For personal use only. No other uses without permission. Copyright © 2025 Massachusetts Medical Society. All rights reserved.

    Refs

    Referências

    Conteúdo migrado do Central de Estudos (Blog Dr. Jackson Fuck, Notion).