A recent study published in the journal Nature found that the autonomous medical AI agent MIRA achieved higher diagnostic accuracy compared to human physicians in a simulated environment. The study evaluated MIRA on a dataset of 574 real-world emergency department cases sourced from the Medical Information Mart for Intensive Care (MIMIC-IV) database.

MIRA, an artificial intelligence agent, operates within sandboxed electronic health record environments. It autonomously handles patient histories, orders diagnostic tests, andFormulates diagnoses and treatment plans within a controlled simulation. The cases included eight distinct diagnoses spanning surgery, internal medicine, and oncology, and MIRA navigated these with 11 specialized digital tools, making over 85,000 operational choices.

The agent achieved an 88.9% diagnostic accuracy across the entire 574-case dataset. In a comparison involving 311 matched cases, MIRA's diagnostic accuracy was 87.8%. This was compared to a cohort of four board-certified physicians, who achieved an average diagnostic accuracy of 78.1% for the same cases. A mixed-seniority medical team, comprising four residents and two board-certified doctors, averaged 71.1% diagnostic accuracy in the matched comparison.

MIRA also demonstrated specific diagnostic strengths, achieving 100% recall for laparoscopic appendectomies. Its diagnostic performance for pancreatic cancer was equivalent to that of board-certified physicians. For critical hospital admission decisions related to pneumonia and pulmonary embolism, MIRA achieved a perfect recall score of 1.00.

Regarding prescriptions, an independent, blinded medical review of 56 patient-level outputs showed zero high-severity drug-drug interactions caused by MIRA. An assessment of 468 prescriptions written by MIRA also found zero renal dosing incompatibilities and zero medication-allergy mismatches. While the authors stated MIRA did not achieve 100% perfection in all treatment choices, such as specific antibiotic selections, its route specification was 97% correct, representing its weakest prescription field.

MIRA requested a broader set of individual blood parameters than human doctors. However, its overall test selection remained below historical dataset baselines. The system specifically avoided the systemic over-ordering of high-cost radiological imaging, and the pulmonary embolism analysis suggested a tendency toward over-admission.