Conference Coverage

GPT-5 Leads AI Models on Maxillofacial Trauma Questions

Edited by:

Key Highlights

  • GPT-5 achieved the highest mean accuracy score at 3.47 on a 4-point scale, followed by ScholarGPT at 3.41.
  • OpenEvidence scored significantly lower than GPT-5, ScholarGPT, and GPT-4o.
  • Ten of 11 responses with mean accuracy scores of 2 or lower stated a high level of certainty.
  • Question complexity affected the models differently rather than producing a consistent change in performance.

GPT-5, accessed through the free version of ChatGPT, had the highest mean accuracy score among 4 artificial intelligence (AI) models evaluated on maxillofacial trauma questions, according to research presented at the 2026 AAOMS Annual Meeting in Seattle, WA. However, nearly all low-scoring answers were delivered with high stated certainty.

Investigators developed 17 simple and 17 complex questions about maxillofacial trauma diagnosis and management. The same prompt structure, including a request for certainty, was submitted to OpenEvidence, ChatGPT-4o, and ScholarGPT on February 20, 2025, and to ChatGPT-5 on August 20, 2025. Three hospital-based oral and maxillofacial surgery attendings independently rated each response on a 4-point scale, from 1 (incorrect or misleading) to 4 (accurate).

Study Findings

Mean accuracy was 3.47 ± 0.74 for GPT-5, 3.41 ± 0.76 for ScholarGPT, 3.13 ± 0.84 for GPT-4o, and 2.61 ± 1.01 for OpenEvidence. Overall accuracy differed significantly across models (P < .001). OpenEvidence scored significantly lower than each of the other three platforms, and GPT-5 scored significantly higher than GPT-4o. GPT-5 and ScholarGPT did not differ significantly. Eleven platform-question responses received a mean reviewer score of 2 or lower. OpenEvidence accounted for 8 of these responses, GPT-4o for 2, ScholarGPT for 1, and GPT-5 for none; the distribution differed significantly by platform (P = .006). Ten of the 11 low-scoring responses stated high certainty.

The effect of question complexity differed by model (interaction P = .0054). GPT-4o and ScholarGPT scored higher on complex items than simple items, while GPT-5 scored slightly lower and OpenEvidence showed a larger decline on complex items.

Clinical Implications

According to the study authors, AI platforms still have clinically important accuracy limitations, and high expressed certainty does not ensure a correct response. They advised that clinicians critically appraise AI-generated answers and not use these platforms as the definitive basis for maxillofacial trauma diagnosis or management decisions.

Expert Commentary

"AI platforms require critical appraisal and should not be relied upon for definitive maxillofacial trauma diagnosis and management decisions," the researchers concluded.


Reference
Antonioni M, Martin HJ, Salama A, Kocak S. Artificial intelligence accuracy across platforms and time for maxillofacial trauma clinical questions. Poster presented at: 108th American Association of Oral and Maxillofacial Surgeons Annual Meeting, Scientific Sessions and Exhibition; September 30-October 3, 2026; Seattle, WA. Accessed September 14, 2026. https://aaoms.org/education-meetings/meetings/2026-annual-meeting/