Key Takeaways
- Large language models (LLMs) are hypothetically capable of improving medical diagnosis by synthesizing disparate types of information more quickly and more systematically than humans.
- This randomized trial tested an EU-certified LLM called Prof. Valmed in diagnosing rheumatologic conditions.
- Use of Prof. Valmed did not improve diagnostic accuracy but reduced the time needed per case, suggesting it could make the process more efficient.
Prof. Valmed, a large language model (LLM) cleared by European regulators for medical diagnosis, outperformed human physicians for diagnosing rheumatologic conditions in terms of speed but not accuracy, a randomized trial showed.
Diagnostic accuracy was almost exactly the same in rheumatology case scenarios for physicians who used Prof. Valmed versus those reaching their diagnoses in their usual ways, according to Johannes Knitza, MD, PhD, MHBA, of Philipps-Universität Marburg in Germany, and colleagues.
But the human doctors relying on their own resources required an average of 206 seconds to make their diagnoses, compared with 94 seconds among those aided by Prof. Valmed (P<0.001), the group reported in a medRxiv preprint manuscript, which has not undergone peer review.
As well, use of Prof. Valmed increased physicians' confidence in their diagnoses -- perhaps too much, Knitza and colleagues suggested. "An interesting finding was that confidence exceeded observed accuracy across all evaluated settings and increased further after [Prof. Valmed] assistance in both groups, indicating persistent overconfidence," the investigators wrote.
"Exploratory analyses furthermore suggested substantial AI [artificial intelligence] over-reliance in the intervention group, whereas under-reliance was uncommon," they continued. "These findings suggest that diagnostic support may increase confidence more readily than correctness and highlight the need to evaluate calibration and behavioural reliance alongside accuracy."
AI system developers have long touted medicine as one of the most important potential applications for the emerging technology: disease diagnosis and management rely so strongly on synthesizing numerous, disparate forms of information that sufficiently capable machines ought to be able to do it better and faster than humans.
Whether the field has reached that point, however, remains uncertain. General AI systems have been tried with mixed results -- sometimes finding diagnoses that had eluded human professionals, but also prone to "hallucinations," decidedly false conclusions that appear to stem from the systems' emphasis on pleasing their users.
Prof. Valmed, founded by Vera Roedel, an attorney for Merck KGaA, and Heinz Wiendl, MD, a neuroimmunologist at University Hospital Freiburg, is designed specifically for medical applications with guardrails to limit hallucinations. It's billed as an "AI copilot" to assist but not replace physicians and received the European Union's CE mark in March 2025, allowing it to be sold for medical use. As an LLM, users can direct it through ordinary language.
For the trial, dubbed ALLIANCE, Knitza and colleagues sought to evaluate its performance in the rheumatology field. They recruited 82 physicians from seven institutions in Germany and Norway, randomizing them 1:1 to make diagnoses in three clinical scenarios either with or without Prof. Valmed's support. These scenarios involved Cogan syndrome, dermatomyositis, and familial Mediterranean fever, all drawn from published case reports. Participants were also told to quantify their confidence in each potential diagnosis.
For example, in the Cogan syndrome case, participants were asked to list up to three possible diagnoses, with probabilities, for a patient described as follows: "61-year-old man. Presented with tinnitus, progressive hearing loss, generalized joint pain, blurred vision, and redness in both eyes."
Participants were also asked to rate their satisfaction with the methods they used. For those assigned to the intervention group, the questions included several specific to Prof. Valmed about its ease of use and trustworthiness, and their interest in using it again.
Only about one-quarter of participants were rheumatology specialists. Many disciplines were represented, including general internal medicine, nephrology, endocrinology, and eight others.
The primary outcome was accuracy, defined as the percentage of instances in which a participant's most likely diagnosis matched the actual published one. This was achieved in 33.3% of intervention group cases versus 35.0% of those approached conventionally, a difference that did not come close to statistical significance. Prof. Valmed's performance looked somewhat better with a less stringent outcome -- having one of a participant's three possible diagnosis match the actual one -- with the system's use leading to a 49.2% success rate, compared with 39.2% in the control group, but the effect was not significant (P=0.291).
Knitza and colleagues also examined Prof. Valmed's "standalone" performance, without the human physician's input. Working by itself, the system was also 33.3% accurate for its top likelihood matching the true diagnosis, and 47.6% accurate in having one of its three candidates be correct.
Confidence ratings vastly exceeded accuracy at every step. In the intervention group, even before Prof. Valmed was called in, the mean confidence rating was 39%, compared with accuracy of 22%. Their 33% accuracy when using the AI system came with confidence ratings averaging 57%. The same pattern was seen in the control group.
Overall, participants liked Prof. Valmed, giving it high ratings for ease of use and a pleasing interface. About two-thirds said it was trustworthy, and over 80% said they would use it again. On the other hand, only 36% said it was easy to fix mistakes when using the system, suggesting the developers still have some work to do.
"LLM-based diagnostic support may be most valuable for improving efficiency, enhancing perceived support quality, and broadening differential diagnosis, while overconfidence and overreliance remain key safety concerns," Knitza and colleagues concluded. "Further real-world trials are needed to define the role of certified LLM-based decision support in routine care."
Facts Only
* An EU-certified LLM named Prof. Valmed was tested for diagnosing rheumatologic conditions in a randomized trial.
* The trial involved 82 physicians from Germany and Norway.
* Physicians used Prof. Valmed's support or conventional methods to make diagnoses in three clinical scenarios: Cogan syndrome, dermatomyositis, and familial Mediterranean fever.
* The intervention group (using Prof. Valmed) required an average of 94 seconds for diagnosis, compared to 206 seconds conventionally.
* Diagnostic accuracy did not show a statistically significant difference between the two groups across all evaluated settings.
* Accuracy in identifying one of three possible diagnoses matched the actual one was 49.2% for the intervention group versus 39.2% for the control group, with no statistical significance (P=0.291).
* The system's standalone performance showed 33.3% accuracy in matching the top likelihood diagnosis and 47.6% accuracy in having one of three candidates be correct.
* Confidence ratings were higher for the intervention group, averaging 57% when using Prof. Valmed, compared to 39% without assistance before intervention.
* Physicians rated the system highly regarding ease of use (over 80% would use it again) and trustworthiness (two-thirds).
Executive Summary
Full Take
Sentinel — Human
The text appears to be a well-structured summary of a scientific trial, presenting data and analysis that exhibits the complexity and specific citation style typical of human-written research reporting.
