ACCURACY OF CHATGPT AND GEMINI IN MANDIBULAR THIRD MOLAR TEETH IN PANORAMIC RADIOGRAPHS: COMPARISON OF UNCOMMANDED AND COMMANDED CASES


Aras S., Solak H., Demirezer K., Kıranşal M., Topaloğlu E. N., Sobi E., ...Daha Fazla

European Congress of DentoMaxilloFacial Radiology, Iasi, Romanya, 25 - 27 Haziran 2026, ss.96, (Özet Bildiri)

  • Yayın Türü: Bildiri / Özet Bildiri
  • Basıldığı Şehir: Iasi
  • Basıldığı Ülke: Romanya
  • Sayfa Sayıları: ss.96
  • İnönü Üniversitesi Adresli: Evet

Özet

Aim: To compare the accuracy of ChatGPT and Gemini for Winter angulation classification of

mandibular impacted third molars (#38, #48) on panoramic radiographs, and to evaluate whether

prompted querying (standardized instruction) improves performance versus unprompted

querying (no prior instruction), using an expert reference standard.

Materials and Methods: This retrospective study included 120 panoramic radiographs. Winter

angulation of #38 and #48 was determined by an oral and maxillofacial radiology specialist

(reference standard). ChatGPT and Gemini were queried under unprompted and prompted

conditions; in the prompted condition, a brief standardized description of Winter categories was

provided immediately before querying. Outputs were coded as correct (1) or incorrect (0). Each

radiograph yielded four evaluations per condition (ChatGPT-38, Gemini-38, ChatGPT-48,

Gemini-48), resulting in 480 paired evaluations per condition. Prompted vs unprompted

performance was tested with the exact McNemar test (two-sided α=0.05). To account for within-

case clustering, a sensitivity analysis used GEE logistic regression with robust standard errors

(condition, model, tooth, and condition×model).

Results: Accuracy increased from 52.5% (252/480) unprompted to 57.5% (276/480) prompted

(+5.0 percentage points), but was not significant (McNemar p=0.129; OR 1.23, 95% CI 0.95–

1.60). GEE showed no significant condition effect (p=0.269) and no condition×model

interaction. Gemini outperformed ChatGPT overall (GEE OR 1.40; 95% CI 1.06–1.84;

p=0.018).

Conclusion: Prompting produced a modest, non-significant improvement, whereas model

selection significantly affected accuracy, with Gemini performing better than ChatGPT.