HANOVER — A new study from Dartmouth College, presented at the 2026 Annual Meeting of the Association for Computational Linguistics, indicates that artificial intelligence can mimic a doctor's tone but often does not parallel a doctor's thought process. The study analyzed 146,000 conversations between 10,105 patients and their primary care physicians at Dartmouth Health.

Researchers developed a tool to compare AI-generated responses with a dataset of actual responses created by health care professionals from Dartmouth Health. Sarah Preum, an assistant professor of computer science at Dartmouth College and co-corresponding author of the study, said, "We find that AI can sound like a doctor but not think like one."

AI-generated answers frequently misalign with what clinicians would write. These inconsistencies include responses that are too long, a failure to ask follow-up questions, and the use of irrelevant or inaccurate medical details. Preum stated, "We didn't just want to measure a platform's accuracy, but whether it actually helps with the workload, which in this case is measured by how much editing the physician is doing."

The research team evaluated physician responses drafted by Claude, Gemini, ChatGPT, Llama, Aloe, and Qwen. The study found that adapting AI to individual physician communication styles can improve accuracy by 33% and reduce editing by 26%. Tim Burdick, an associate professor of community and family medicine at Dartmouth's Geisel School of Medicine and a family medicine physician at Dartmouth Health, said, "I don't foresee a time when the portal can respond to a patient without a clinician editing it first."

Researchers created a technique called Thematic Agentic Direct Preference Optimization for Learning Enhancement (TADPOLE). This technique trains AI platforms using a hybrid model built from physician- and AI-generated responses. When TADPOLE was integrated into six commercial large language models, the drafted responses more closely matched physicians' standards for precision and information quality.

Burdick noted that physicians he works with report that AI-generated drafts save approximately 25% of their time on shorter messages. He added, "I would guess we need to get to where the physician is editing less than 30% of the content before it has substantial benefit." The researchers indicate that AI tends to be more empathetic and thorough than physicians who are constrained by time. The study also found that 65% of all portal messages analyzed came from individuals over 55, with patients over 65 generating 24% of all studied messages.