AIe2501008
Visão geral

NEJM AI 2025;2(11)
EDITORIAL
Humanity’s Next Medical Exam: Preparing to Evaluate Superhuman Systems
Jack Gallifant

M.B.B.S.,1,2 and Danielle S. Bitterman

M.D.1,2
Received: September 4, 2025; Accepted: September 17, 2025; Published: October 23, 2025
Abstract
The rapid advances in health care AI necessitate a fundamental shift in how we evaluate these systems. Palepu et al. (2025) demonstrate that AI can outperform medical trainees in breast cancer management questions, illustrating that advances that would have been difficult to foresee only a few years ago are imminent. However, current methods, often reliant on question and answer tasks, are inadequate for capturing the nuances of clinical practice, even as models begin to exceed human performance on these narrow metrics. The trajectory of AI development toward more generalized and autonomous systems introduces profound opportunities alongside substantial risks, making the limitations of our existing oversight frameworks an urgent problem. In this editorial, we propose Humanity’s Next Medical Exam, a novel approach designed to measure and promote the development of safe, human-aligned AI. This paradigm is built upon three foundational pillars: interactive interrogation to challenge models beyond rote knowledge, experiential learning in sandbox environments to assess decision-making under uncertainty, and real-world continuous learning to monitor and refine performance postdeployment. The maturation of these core components as part of a comprehensive evaluation framework is a critical step toward helping us prepare to take advantage and avoid risks of continually advancing AI technologies in the future. (Funded by the National Institutes of Health, the National Cancer Institute, and others.)
R and the architectural shift toward agentic self-reflection are genuine breakthroughs. gressed from simple fact recall to complex reasoning, collaboration, and tool use.1-4 The ability of these systems to reason through complex medical guidelines apid advances in health care AI warrant reflection. In a few years, AI has pro
The case study by Palepu et al. (2025) highlights this progress.5 They evaluated the per
formance of a proprietary, large language model (LLM)–based system in providing clinical The author affiliations are listed at
recommendations on 60 brief, synthetic breast cancer presentations. Their system outper the end of the article.
formed medical trainees with less than 1 year of oncology-specific training, but not attending Danielle S. Boncologists, in clinical management reasoning tasks. Rates of harm in the generated output were low. The system is an adaptation of Google’s Articulate Medical Intelligence Explorer, ent
or AMIE, based on Gemini 2.5 Pro, with access to Internet search and self-reflection. While of Radiation Oncology, Dana- Farber Cancer Institute/Brigham the authors acknowledge that the results fall short of immediate clinical applicability, they and Women’s Hospital, 75 Francis are illustrative of advances that would have been difficult to foresee only 3 years ago. Street, Boston, MA 02115.
NEJM AI is produced by NEJM Group, a division of the Massachusetts Medical Society.
Downloaded from ai.nejm.org by Fabio Ynoe de Moraes on November 27, 2025. For personal use only. No other uses without permission. Copyright © 2025 Massachusetts Medical Society. All rights reserved.
The clinical community has responded to this progress with measured caution, and rightly so. The promising reports have not led to a clinical consensus that such AI systems are true clinical reasoners. This caution stems from a deep understanding of clinical reality, not disbelief in study results. This is why passing an exam is insufficient to produce a competent clinician and is not the sole requirement for licensure. Medical training builds knowledge that generalizes to the complex reality of clinical care; its goal is to develop intuition and problem-solving skills, not rote memorization. The physician’s role is more than a sequence of questions and answers, and decisions often have no single “right” answer, creating a vast chasm between benchmarks and the bedside. Direct patient contact may also account for only 40% of a physician’s shift.6 The rest is spent navigating bureaucracy, reviewing disorganized data, coordinating care, speaking with families, and reacting to misaligned system incentives.
Evaluation and safety should evolve to keep pace with the exponential progress made in just the last 18 months. As we and others have argued, AI evaluations should test more realistic clinical tasks and meaningful outcomes.7,8 There are even more foundational concerns, however, with the nature of current evaluations. When scoped to simple question and answer tasks, systems today have the potential to achieve superhuman performance, which refers to the progression of technology toward autonomous abilities exceeding human performance. Although the term superhuman has futuristic and provocative connotations, strictly defined superhuman performance on a controlled, measurable task is currently achievable. Soon, superhuman performance may extend to real-world settings where we anticipate and even desire experiential learning and development. Even with today’s narrow implementations, many hospitals are already grappling with the fact that there is no widely accepted or validated way to measure such implementations in the real world. As systems continue to progress, the risks of being unprepared for increasingly superhuman capabilities demand a fundamental transformation in the conceptual and methodological frameworks underpinning evaluation.
These new strategies should prepare us for the “era of experience” in medical AI, where models evolve from experiential learning.9 For example, plotting attending physician performance as the 100% ceiling limits systems that may in the future exceed human knowledge recall.10-14 The challenge is shifting from whether machines can know as much as or even more than humans to how we measure and govern their ability to act wisely and in alignment with
human values. The opportunity is to leverage this expertise to revolutionize health care cost, quality, and accessibility.
We need a new paradigm: Humanity’s Next Medical Exam. This proposal is inspired by Humanity’s Last Exam, an extremely difficult general knowledge benchmark designed to avoid performance ceilings; the best LLMs score only 25%.15 This medical equivalent would realistically assess clinical skills, identify system gaps, and drive progress. Instead of a written test, we propose a multifaceted strategy for experiential learning and evaluation of real-world care delivery, built on three pillars.
The first pillar is interactive interrogation, resembling a rigorous, oral board-style exam rather than a multiple-choice test. The question is whether a system can sustain a coherent clinical conversation, defend treatment plans, revise recommendations with new data, and reconcile contradictions with the fluency of a clinical peer. The focus must shift from knowledge recall to applied reasoning, adaptability, and accountability.
Second, interactive learning in sandbox environments. Early medical AI, trained on static data, was a form of statistical imitation with inherent limits. The next frontier is learning from experience in high-fidelity simulated environments where AI can act, observe outcomes, and learn from measurable consequences. This approach enables a sustained cycle of advancement and evaluation that evolves in step with capabilities.9 Current reinforcement learning–based fine-tuning offers a robust framework to leverage these interactions, allowing a model to refine its policies beyond what is explicitly stated in a textbook. Such sandboxes could host virtual apprenticeships where models learn clinical trade-offs, such as balancing collecting additional test results against diagnostic speed and the risk of a missed finding, by navigating realistic patient trajectories and goals of care. Though algorithms will evolve, the central challenge is creating a large, diverse set of high-quality interactive environments. This is the path toward systems that learn from consequences, not just information.
Third, real-world continuous learning. The learning health system, a decades-old promise that has largely failed to materialize, may finally be realized through AI.16 Here, the exam is continuous: Can the system improve with each patient interaction? Can it integrate the latest research for a specific patient or help titrate medication by synthesizing genomic data with patient feedback? This component will push the limits of human oversight toward superhuman performance — precisely the point of a learning health system. This would build and scale institutional knowledge in
NEJM AI 2
NEJM AI is produced by NEJM Group, a division of the Massachusetts Medical Society.
Downloaded from ai.nejm.org by Fabio Ynoe de Moraes on November 27, 2025. For personal use only. No other uses without permission. Copyright © 2025 Massachusetts Medical Society. All rights reserved.
near real time, making every encounter a systemwide learning opportunity.
Building these pillars requires close collaboration between clinical and engineering communities to answer many open questions. Urgent investment is needed in benchmarking and evaluation to stress test current technologies and prepare for future, higher-performing systems.
This assessment requires an AI-ready, modernized health care system; technical advances are insufficient if incentives and infrastructure are outdated. This vision depends on secure, near-real-time research sandboxes that support both high-fidelity simulation of historical trajectories and generation of synthetic data for testing novel scenarios. Building these requires overcoming siloed data. Interoperable standards are not enough; we believe stronger government regulation is needed to mandate open, standardized application programming interfaces for electronic health records and enforce information sharing. This would prevent vendor lock-in and foster a competitive ecosystem of innovators. Without this modernization, our most potent tools will remain unvetted for real-world use.
Physician skepticism reflects an understanding of systemic and clinical realities, not a fear of technology. Clinicians know their expertise extends beyond medical knowledge to the clinical judgment developed through practice, managing atypical patients with differing values, payer denials, competing demands, and providing comfort amid uncertainty.
The shift from static information to dynamic interaction marks a new era of experience for medical AI. Success will be contingent on evaluation that keeps pace with and accelerates model capabilities, paired with a modernized data infrastructure with enforced standards to unlock siloed data. Medicine is practiced in the complex human space between facts, and it is here that AI must now be tested.
necessarily represent the views of PCORI, its Board of Governors, or its Methodology Committee.
Author Affiliations
1 Artificial Intelligence in Medicine (AIM) Program, Mass General Brigham, Harvard Medical School, Boston
2 Department of Radiation Oncology, Brigham and Women’s Hospital and Dana-Farber Cancer Institute, Boston
References
Disclosures ).
Author disclosures are available at .
We acknowledge financial support from the National Institutes of Health National Cancer Institute (grant numbers U54CA274516-01A1 [J.G. and D.S.B] and R01CA294033-01 [J.G. and D.S.B.]), the American Cancer Society and American Society for Radiation Oncology (ASTRO-CSDG-24-1244514-01-CTPS, grant DOI: [D.S.B.]), a Patient-Centered Outcomes Research Institute (PCORI) Project Program Award (grant number ME-2024C2-37484 [D.S.B.]), and the Woods Foundation [D.S.B.]. All statements in this report, including its findings and conclusions, are solely those of the authors and do not
NEJM AI 3
NEJM AI is produced by NEJM Group, a division of the Massachusetts Medical Society.
Downloaded from ai.nejm.org by Fabio Ynoe de Moraes on November 27, 2025. For personal use only. No other uses without permission. Copyright © 2025 Massachusetts Medical Society. All rights reserved.
NEJM AI 4
NEJM AI is produced by NEJM Group, a division of the Massachusetts Medical Society.
Downloaded from ai.nejm.org by Fabio Ynoe de Moraes on November 27, 2025. For personal use only. No other uses without permission. Copyright © 2025 Massachusetts Medical Society. All rights reserved.
Referências
Conteúdo migrado do Central de Estudos (Blog Dr. Jackson Fuck, Notion).