Saxsons Group — India's trusted nuclear medicine, radiotherapy, oncosurgery, dosimetry and cyclotron supplier since 1986
📖 Free full textPeer-ReviewedOpenAlexReviewScientific Reports · 2026

Clinician-rated multidimensional quality and readability of patient-facing glioma information generated by six generative AI chatbot platforms

Tingting Zhao, Yuan Zhou, Zejuan Gu

Abstract

Generative AI chatbots are increasingly used to obtain health information. In glioma, patient-facing chatbot text should be safe, accurate, reliable, empathetic, transparent, and readable. However, readability formulas quantify textual complexity and do not establish patient comprehension or educational effectiveness. This study therefore evaluated clinician-rated, patient-facing glioma information generated by six publicly accessible chatbot platforms. From April 20 to April 26, 2026, six publicly accessible chatbot platforms were tested using the user-interface model identifiers displayed at the time of access: ChatGPT Plus 5.4, Gemini 3.0 Pro, Perplexity, DeepSeek-V3.1, Doubao, and Copilot. These labels were treated as platform-level identifiers rather than stable underlying model versions. Fifty predefined glioma-related questions, informed by patient interviews and online search-interest data, were submitted as single-turn prompts; three responses were separately generated in new chat sessions for each question-platform pair. Two experienced neurosurgery clinicians rated safety, accuracy, empathy, reliability, information quality, and transparency, and established formulas were used to estimate readability. Paired comparisons used Cochran’s Q or Friedman tests, with Benjamini-Hochberg-adjusted post hoc analyses. Significant between-platform differences were observed for all non-safety outcomes. Pair-level safe-response proportions ranged from 90.0% (DeepSeek; 95% Wilson CI 78.6–95.7%) to 98.0% (Copilot; 95% CI 89.5–99.6%), with no overall difference (Q = 5.000, P = 0.416). DeepSeek had the highest median accuracy score (5.00 [4.50, 5.00]), whereas ChatGPT had the lowest (4.25 [4.00, 4.50]); the between-platform effect was small (Kendall’s W = 0.130). Empathy differed more strongly (W = 0.595). Gemini had the highest median DISCERN score (50.75 [47.50, 54.75]; W = 0.666), whereas Perplexity had the highest EQIP (75.00 [74.50, 76.12]), GQS (5.00 [4.50, 5.00]), and JAMA (1.00 [1.00, 1.50]) scores. Readability also differed; median Flesch-Kincaid grade level ranged from 8.47 for Gemini to 13.43 for Perplexity (W = 0.539). No platform consistently met the prespecified sixth-grade target. In this clinician-rated, cross-sectional text evaluation, unsafe or potentially misleading responses were uncommon but occurred on every platform. The absence of a statistically significant safety difference was not an equivalence or noninferiority finding. Between-platform variability was observed in accuracy, empathy, DISCERN-rated treatment information, EQIP-rated patient-information quality, GQS-rated overall quality, JAMA transparency, and formula-based readability. These results describe predefined English-language, single-turn outputs during a one-week access window and do not demonstrate patient comprehension, educational effectiveness, real-world clinical safety, or therapeutic benefit. Chatbot-generated glioma information should be treated as supplementary material requiring clinician review, source verification, and plain-language adaptation before patient-facing use.

Related in the same topic