Performance of AI models in pharmaceutical clinical reasoning: analysis of simulated questions regarding consistency, safety, and clarity
DOI:
https://doi.org/10.22491/2357-9730.155747Keywords:
Generative Artificial Intelligence, Patient Safety, Clinical ReasoningAbstract
Introduction: Widely known and freely accessible natural language models stand out for their simplified interfaces and their widespread use in daily practice by professionals and lay users alike, generating rapid responses based on large volumes of data. However, a limitation is that they also mine grey literature data. In this work, we evaluated the performance of three AI models (ChatGPT, Copilot, and Gemini) in solving simulated clinical questions, considering three domains: consistency, safety, and clarity of responses. Method: This exploratory, descriptive, and comparative study included 15 simulated clinical questions, categorized by level of difficulty (low, moderate, and high complexity). Each AI responded to the same questions, and scores were assigned by expert evaluators across the three domains of interest. Results: Overall, there were no statistically significant differences among the three models regarding consistency, safety, or clarity (p > 0.05), indicating equivalent general performance. All models performed better on low-complexity questions, with highlights for Copilot (87.5% consistency) and ChatGPT (100% safety). In moderate-complexity questions, Copilot achieved the best results (80.95% consistency, 88.10% safety, and 82.54% clarity). However, statistical analysis confirmed that question complexity significantly impacted the Safety domain (p = 0.0028). When the questions progressed from low to high difficulty, the ability of the AIs to provide safe responses plummeted significantly. In high-complexity questions, performance declined markedly, especially for ChatGPT (50% consistency and 54.17% safety). Gemini showed greater stability in this group, with 66.67% consistency, 70.83% safety, and 83.33% clarity. Copilot was the only AI to provide a completely correct answer in one high-complexity question. Conclusion: We corroborate that AI can be an ally to the clinical pharmacist, especially as technical support for standardized interventions. However, the statistically proven safety risks in highly complex scenarios reinforce that it does not replace the critical and interpretive role of the human professional. The use of these tools should be considered complementary to clinical reasoning.
Downloads
Published
How to Cite
Issue
Section
Categories
License
Copyright (c) 2026 Josué Guilherme Lisbôa Moura, Maria Gabriela Borges Hermes, Josiele Dias da Rosa Mossmann, Lauren Feiffer Sanches

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).