Performance of AI models in pharmaceutical clinical reasoning: analysis of simulated questions regarding consistency, safety, and clarity

Authors

DOI:

https://doi.org/10.22491/2357-9730.155747

Keywords:

Generative Artificial Intelligence, Patient Safety, Clinical Reasoning

Abstract

Introduction: Widely known and freely accessible natural language models stand out for their simplified interfaces and their widespread use in daily practice by professionals and lay users alike, generating rapid responses based on large volumes of data. However, a limitation is that they also mine grey literature data. In this work, we evaluated the performance of three AI models (ChatGPT, Copilot, and Gemini) in solving simulated clinical questions, considering three domains: consistency, safety, and clarity of responses. Method: This exploratory, descriptive, and comparative study included 15 simulated clinical questions, categorized by level of difficulty (low, moderate, and high complexity). Each AI responded to the same questions, and scores were assigned by expert evaluators across the three domains of interest. Results: Overall, there were no statistically significant differences among the three models regarding consistency, safety, or clarity (p > 0.05), indicating equivalent general performance. All models performed better on low-complexity questions, with highlights for Copilot (87.5% consistency) and ChatGPT (100% safety). In moderate-complexity questions, Copilot achieved the best results (80.95% consistency, 88.10% safety, and 82.54% clarity). However, statistical analysis confirmed that question complexity significantly impacted the Safety domain (p = 0.0028). When the questions progressed from low to high difficulty, the ability of the AIs to provide safe responses plummeted significantly. In high-complexity questions, performance declined markedly, especially for ChatGPT (50% consistency and 54.17% safety). Gemini showed greater stability in this group, with 66.67% consistency, 70.83% safety, and 83.33% clarity. Copilot was the only AI to provide a completely correct answer in one high-complexity question. Conclusion: We corroborate that AI can be an ally to the clinical pharmacist, especially as technical support for standardized interventions. However, the statistically proven safety risks in highly complex scenarios reinforce that it does not replace the critical and interpretive role of the human professional. The use of these tools should be considered complementary to clinical reasoning.

Downloads

Download data is not yet available.

Published

2026-10-09

How to Cite

1.
Moura JGL, Hermes MGB, Mossmann JD da R, Sanches LF. Performance of AI models in pharmaceutical clinical reasoning: analysis of simulated questions regarding consistency, safety, and clarity. Clin Biomed Res [Internet]. 2026 Oct. 9 [cited 2026 Oct. 10];46. Available from: https://seer.ufrgs.br/index.php/hcpa/article/view/155747

Similar Articles

1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 > >> 

You may also start an advanced similarity search for this article.