BOSTON — OpenAI's o1 reasoning model outperformed experienced physicians in diagnosing emergency department patients at Beth Israel Deaconess Medical Center in Boston, according to a study published April 30, 2026 in the journal Science. Researchers at Harvard Medical School and Beth Israel Deaconess Medical Center conducted the research, which tested the artificial intelligence system against doctors and the earlier GPT-4 model on diagnoses and care decisions.
The team ran experiments using actual cases from the hospital's emergency department, including a lupus patient who presented with a pulmonary embolism. The researchers graded the model's diagnostic accuracy at three stages from emergency room triage through hospital admission. The model outperformed two experienced physicians while using only electronic health records and the same limited information available to those clinicians. The study also drew on case reports from the New England Journal of Medicine and clinical vignettes to test the system against established diagnostic benchmarks.
In a trial of 76 emergency department patients, the model identified the exact or very close diagnosis in 67% of cases, compared with 50–55% accuracy for human doctors. When more detail was available, the model's diagnostic accuracy rose to 82%, compared with 70–79% for human doctors, a difference that was not statistically significant. In a comparison with 46 doctors on five clinical case studies, the model scored 89% on treatment planning tasks, compared with 34% for the doctors using conventional resources. The study found the model's advantage was particularly pronounced in triage situations requiring rapid decisions with minimal information.
"This is the big conclusion for me — it works with the messy real-world data of the emergency department. It works for making diagnoses in the real world." said Dr. Adam Rodman, a clinical researcher on the study.
Raj Manrai, an assistant professor of Biomedical Informatics, addressed the comparison with clinicians. "The model outperformed our very large physician baseline," he said.
The study tested only text-communicated patient data and did not include signals such as patient appearance or level of distress. The model relied on text alone, while clinicians also use inputs such as images, sounds and nonverbal cues when diagnosing and treating patients. Prior versions of large language models had difficulty managing uncertainty and generating lists of possible conditions known as differential diagnoses.
"You have something which is quite accurate, possibly ready for prime time," said Dr. David Reich, a chief clinical officer.
forum Comments (0)
No comments yet. Be the first to comment.