The performance of ChatGPT and other large language models on multiple-choice questions in biomedical disciplines: A meta-analysis

Authors

Document Type

Article

Publication Date

5-19-2026

Publication Title

Anatomical Sciences Education

Abstract

While large language models (LLMs) have shown promise as learning tools for medical education, their reported accuracy on multiple-choice questions (MCQs) varies widely across studies, necessitating synthesis. This meta-analysis synthesizes LLM accuracy on text-based MCQs from biomedical disciplines and USMLE Step 1-level content and explores study characteristics that moderate LLM performance. Studies published between January 1, 2022, and August 5, 2025, were identified using seven databases. Titles and abstracts were screened against eligibility criteria, which included testing LLM performance on text-based MCQs in English. Extracted data included accuracy rates (correct responses/total questions) and study characteristics. A random-effects proportional meta-analysis was used to calculate the pooled accuracy of LLM versions across studies, and a meta-regression tested the effects of study characteristics on accuracy. The final analysis included 41 articles comprising 207 proportions. Newer OpenAI models (GPT-4o through GPT o1 = 90%; GPT-4 = 82%) demonstrated significantly higher accuracy than earlier versions (GPT-3/3.5 = 59%; p < 0.001). Model version explained 56% of the total variance (p < 0.001). The pooled accuracy of newer Anthropic (Claude 3 through 3.7 = 86%) and DeepSeek (R1 = 86%) versions fell within the range of performance for the newer OpenAI versions, while newer Google (Gemini series = 76%) and Microsoft (Copilot = 69%) versions demonstrated lower performance. A discipline-level analysis demonstrated that OpenAI's GPT-4 through GPT-o1 models achieved pooled accuracy rates between 83 and 87% for most disciplines, including the anatomical sciences. These findings demonstrate that newer LLMs are achieving higher accuracy than older models on biomedical and Step 1-level MCQs, supporting their potential integration as supplementary study tools in foundational biomedical education.

First Page

1

Last Page

14

PubMed ID

42153764

Creative Commons License

Creative Commons Attribution 4.0 International License
This work is licensed under a Creative Commons Attribution 4.0 International License.

Share

COinS