UW Researchers Flag Errors in Popular AI Clinical Tool

Jessica Bergsbaken portrait in a white hallway
Jessica Bergsbaken, assistant teaching professor at the UW–Madison School of Pharmacy. | Photo by Sharon Vanorny

A School of Pharmacy study explores what OpenEvidence can get wrong about medications, including a dangerously high drug dose and misread research

By Susan Smith

AI tools are quickly integrating into many people’s daily workflows, including into the hands of clinicians looking for quick answers about medications.

OpenEvidence calls itself “the leading AI-powered medical search and decision-support platform.” The company claims that almost two-thirds of physicians in the U.S. actively use the tool for anything from drafting prior authorization requests and note-taking to supporting clinical decision-making.

But a new study from the University of Wisconsin–Madison School of Pharmacy’s Jessica Bergsbaken raises warning flags about its pharmacy advice.

In a paper published in a special issue of the Journal of the American College of Clinical Pharmacy devoted to practical applications of artificial intelligence in clinical pharmacy, Bergsbaken and colleagues found that the tool made mistakes in dosing recommendations for a heart medication and summarized research that did not apply to a question related to miscarriage management.

“AI tools should not be used in a silo, but in tandem with other resources.”
–Jessica Bergsbaken

Bergsbaken, assistant teaching professor and director of the pharmacotherapy skills laboratory at the School, says that pharmacy students, clinicians, and other users need to be aware of the limitations of OpenEvidence when using it for clinical decision-making. She points to an apt metaphor by another AI researcher, Kristian Hammond.

“If your car’s windows malfunctioned 1% of the time, it might be annoying, but you could live with it,” Hammond said. “However, if your brakes failed 1% of the time, that would be unacceptable — potentially fatal. The same applies to AI summarization: It all depends on the nature of that task and the potential impact of failure.”

Bergsbaken is co-chair of the School’s AI workgroup — tasked with exploring how the new tech fits into the curriculum — and is a practicing pediatric pharmacist at the American Family Children’s Hospital. Her research team included Paije Wilson, Susan Vandagriff, Leslie Christensen, and Lia Vellardita, all of the UW’s Ebling Library for the Health Sciences.

Because OpenEvidence is relatively new, there is sparse research on its efficacy, particularly in the realm of medication and dosing. That’s the gap that Bergsbaken set out to fill.

What is OpenEvidence?

OpenEvidence is a for-profit large language model (LLM) used by health care professionals to help answer clinical questions using sources from medical literature. It has partnerships with leading journals such as the JAMA network, the New England Journal of Medicine and the National Comprehensive Cancer Network, as well as other publishers. Currently, it is supported by advertising and free to any health care provider with a National Provider Identifier (NPI) or students who have proof of current enrollment in health sciences.

Jessica Bergsbaken works at her computer in the sunny atrium of Rennebohm Hall
Jessica Bergsbaken, assistant teaching professor at the UW–Madison School of Pharmacy. | Photo by Sharon Vanorny

There is a lack of transparency about how OpenEvidence processes queries and what sources it uses. It may lack access to relevant literature that is behind paywalls or use evidence that is not peer reviewed. Does it preferentially pull articles from the affiliated journals? We just don’t know the answer to that.

Other studies have found that the accuracy of OpenEvidence varies from being 100% accurate to less than 30%. It depends on how straightforward the question is; if it’s more nuanced, AI is not always able to discern the correct answer, which is illustrated by the examples in our paper. For both our examples, we created test questions and compared OpenEvidence’s answers to treatment guidelines, primary literature, and drug package insert information.

How did OpenEvidence do on prescribing advice?

It performed poorly in answering a medication question that was not clearly outlined in a drug package insert. We asked it to convert a dose of one common beta blocker, metoprolol tartrate, to an equivalent dose of carvedilol. It came up with a dose that was two to four times higher than an equivalent dose for switching between those two different beta blockers. This is scary, because it can have very significant consequences for patients. Unless you have knowledge and understanding about what an appropriate dose conversion should be, that would be easily missed.

And not only was the answer incorrect, but it also didn’t reference the two recommended dose conversions that are cited in literature.

How did OpenEvidence perform in summarizing studies?

We used a real-world clinical query encountered by one of our authors: In the case of an incomplete miscarriage, is misoprostol alone as effective as the combination of misoprostol and mifepristone? This was designed to test its abilities on more nuanced questions.

Again, the AI bot gave an incorrect answer, asserting that studies showed that combination was more effective. It cited two irrelevant studies and misrepresented the findings of a third, relevant study. Further, when we repeated the experiment six months later, two of the three sources were different, but it still gave an incorrect answer and referred to a “landmark study” that actually excluded patients with incomplete miscarriages.

What does this mean for AI-assisted clinical decision-making?

It’s important that AI tools are used in conjunction with other trusted resources. In pharmacy, that means looking at other drug information and primary literature that is out there. AI tools should not be used in a silo, but in tandem with other resources.

Keep Reading
Paula Voorheis portrait outside of a building with many large glass windows
AI Health Screening Meets User-Centered Design

I³ award brings Assistant Professor Paula Voorheis and an interdisciplinary team together to improve AI-based disease screening with user-centered design and AI-enabled evaluation.