The Emerging Role of Large Language Models in Improving Prostate Cancer Literacy

Bioengineering (Basel) 2024 AI 6 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Page 1
The Health Literacy Gap in Cancer Care

When a patient is diagnosed with prostate cancer, their ability to understand their condition, treatment options, and what to expect has a direct impact on their outcomes. This concept -- health literacy -- encompasses how well patients can find, understand, and use health information to make decisions. Research consistently shows that higher health literacy correlates with better adherence to treatment, fewer complications, and improved survival.

Traditionally, the primary tools for improving patient health literacy have been printed materials such as patient guides, brochures, and pamphlets prepared by hospitals or patient advocacy organizations. These materials are carefully written, clinically vetted, and widely distributed, but they have real limitations: they cannot be updated quickly, they cover only general information, and they cannot respond to individual patient questions.

Since late 2022, large language models (LLMs) such as ChatGPT have entered widespread public use. These AI systems can engage in natural language conversation, synthesize complex information from their training data, and respond to specific questions in plain language. Patients are already using these tools to seek health information -- raising both opportunities and concerns for healthcare providers.

This study was designed to systematically evaluate how well three major publicly available LLMs -- ChatGPT 3.5, Microsoft CoPilot, and Google Gemini -- perform as sources of prostate cancer information, compared to the official printed Patient's Guide used in Romania. The evaluation was conducted by expert clinicians within a specific linguistic and cultural context, addressing a gap in research that has largely focused on English-language environments.

TL;DR: Patient health literacy is strongly linked to cancer outcomes, and AI chatbots are already widely used by patients seeking health information, motivating a rigorous comparison against official educational materials.
Pages 2-3
Study Design and Expert Evaluation Framework

The researchers developed 25 questions reflecting the most common information needs of newly diagnosed prostate cancer patients. Topics included symptoms, screening and diagnostic procedures, treatment options (surgery, radiation, hormone therapy), post-treatment care, and lifestyle recommendations. The question bank was developed and refined with input from clinical experts in oncology and urology.

On a single day in February 2024, all three LLMs and the official Patient's Guide were queried with the same standardized prompt: a patient with a new prostate cancer diagnosis asking for help understanding their condition. Queries were conducted in incognito browsing mode to eliminate personalized search biases. A single operator collected all responses to ensure consistency.

A panel of eight expert physicians specializing in prostate cancer, all affiliated with the largest prostate cancer treatment center in Bucharest, Romania, independently evaluated all responses. Critically, the experts were blinded to the source -- they did not know which response came from which LLM or from the Patient's Guide. Responses were randomized before presentation to prevent order effects.

Each response was scored on four criteria using a Likert scale from 1 to 5: accuracy (correctness of medical content), timeliness (currency of the information), comprehensiveness (completeness of the answer), and ease of use (clarity and accessibility for patients). Statistical analysis used linear mixed-effects models to account for variation among experts and to identify significant differences between information sources.

TL;DR: Eight blinded prostate cancer specialist physicians rated responses from three AI chatbots and an official patient guide on 25 expert-validated questions using a four-criteria scoring framework.
Pages 7-8
ChatGPT Leads, Gemini Lags

Across all five statistical models tested -- overall performance and each of the four individual criteria -- ChatGPT 3.5 consistently achieved the highest ratings. In the overall model, ChatGPT scored significantly higher than the Patient's Guide (estimate +55.00, p less than 0.001) and significantly higher than Gemini (estimate +54.75, p less than 0.001). These are statistically robust differences.

CoPilot also outperformed the Patient's Guide significantly in overall scoring (estimate +35.75, p = 0.005) and on specific criteria including accuracy and comprehensiveness. However, the difference between ChatGPT and CoPilot was not statistically significant, suggesting they perform at a comparable level. Both LLMs were seen as improvements over the Guide.

Gemini showed the weakest performance. In the overall model, its score was statistically indistinguishable from the Patient's Guide (estimate +0.25, p = 0.98). For timeliness, Gemini actually scored significantly lower than the Guide (p = 0.011). Gemini was rated significantly below both ChatGPT and CoPilot across multiple criteria, suggesting it is not a reliable source for prostate cancer information in this context.

The Patient's Guide itself received consistently high scores as a baseline -- its intercept estimate of 361.50 indicates it is a genuinely effective educational resource. However, ChatGPT and CoPilot's significantly higher scores demonstrate that AI tools can provide added value beyond what static print materials can offer, particularly in comprehensiveness and accuracy.

TL;DR: ChatGPT 3.5 significantly outperformed the official Patient's Guide and all other AI tools; CoPilot was comparable to ChatGPT; Gemini performed no better than the printed guide and was worse on timeliness.
Pages 9-10
What These Results Mean for Patient Education

The finding that ChatGPT and CoPilot outperformed the official Patient's Guide across multiple criteria suggests that AI tools have genuine potential to enhance prostate cancer patient education. The print guide's limitations -- space constraints, inability to be updated rapidly, and inability to answer follow-up questions -- are structural weaknesses that LLMs naturally overcome.

The fact that these results were obtained in the Romanian language is significant. Most evaluations of medical LLM performance have focused on English. This study demonstrates that at least ChatGPT and CoPilot maintain high quality across languages, which is important for ensuring equitable access to health information globally and reducing the disparities that can arise when only English-language health content is high quality.

The authors note that Gemini's poor performance relative to ChatGPT is notable, as this may reflect differences in training data, fine-tuning for medical contexts, or the models' built-in tendencies toward caution or brevity in medical responses. The divergence in performance suggests that not all LLMs are equal as health information tools, and specific evaluation before deployment is necessary.

The study authors and the broader literature emphasize that physician oversight remains essential. LLMs can produce plausible-sounding but incorrect medical information (a problem sometimes called hallucination), and patients may lack the medical background to identify errors. The ideal model is collaborative: clinicians help develop, verify, and update AI content, while patients benefit from more accessible and interactive information delivery.

TL;DR: AI chatbots can overcome the structural limitations of printed guides, and their performance quality in non-English languages makes them potentially valuable for global health equity, provided physician oversight prevents misinformation.
Pages 9-11
Ethical Dimensions and the EU AI Act

The widespread use of AI for medical information raises important ethical questions that the authors address directly. The accuracy of AI-generated health advice is not perfect, and when patients make decisions based on incorrect information, the consequences can be significant. This creates a responsibility for both AI developers and healthcare systems to ensure quality and accuracy.

The authors reference the EU AI Act, which was approved and will take effect from 2026 as the world's first comprehensive legal framework governing AI systems. Under this framework, AI systems used in healthcare contexts are classified as high-risk and subject to strict requirements for transparency, accuracy, human oversight, and safety evaluation. This regulatory development is directly relevant to the deployment of medical chatbots.

The concept of a human-LLM collaborative model is proposed as the appropriate framework for healthcare AI. Rather than patients relying on AI autonomously, a collaborative model would involve physicians actively participating in the development of AI health content -- contributing medical expertise, reviewing accuracy, updating information as guidelines change, and helping patients interpret AI-generated information within the context of their specific situation.

An important limitation acknowledged by the authors is that the expert panel consisted entirely of male physicians in a single institution, which may introduce perspective biases. The LLMs were evaluated only at a single point in time, and their responses change as models are updated -- a finding from today may not hold for the same model six months later, highlighting the need for ongoing evaluation rather than one-time certification.

TL;DR: Deploying AI for patient health education requires ongoing physician involvement for quality assurance, and emerging regulations like the EU AI Act will require high-risk medical AI systems to meet strict accuracy and oversight standards.
Page 11
A Path Forward for AI-Assisted Cancer Literacy

This study provides the first comparative evaluation of three major AI chatbots against an official patient guide for prostate cancer education within a specific non-English cultural context. The findings clearly show that ChatGPT 3.5 and CoPilot provide information of higher quality than the current Patient's Guide as assessed by expert clinicians -- a meaningful finding for improving patient education.

The practical implication is that healthcare systems and patient advocacy organizations should explore integrating LLMs -- particularly well-performing ones -- into patient education frameworks. This does not mean replacing clinical consultation or printed materials, but rather augmenting them with AI tools that can answer specific questions, explain complex concepts in accessible language, and provide more comprehensive information than static documents allow.

Future research directions should include evaluation of LLM performance on specific sub-topics within prostate cancer, investigation of how LLM-provided information affects actual patient behavior and decisions, and testing across more languages and cultural contexts. Studies examining whether AI health literacy tools improve patient outcomes -- not just information quality as judged by experts -- would provide the strongest evidence for clinical adoption.

The authors conclude with a call for continuous improvement, rigorous ongoing testing, and thoughtful clinical integration with appropriate ethical and regulatory oversight. AI tools like ChatGPT and CoPilot show real promise for democratizing cancer knowledge, but realizing this promise safely requires treating them as medical tools -- subject to the same standards of evidence, review, and accountability as any other medical intervention.

TL;DR: ChatGPT and CoPilot have demonstrated they can surpass official patient guides in prostate cancer education quality, supporting their cautious, physician-supervised integration into cancer literacy programs worldwide.
Citation: Open Access, . Available at: PMC11274300.