Patient- and clinician-based evaluation of large language models for patient education in prostate cancer radiotherapy

Strahlenther Onkol 2025 Treatment 7 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Pages 1-2
Why Patient Education Matters in Prostate Cancer Treatment

Prostate cancer is the most common cancer in men, affecting approximately 13 percent of men over their lifetime. Radiotherapy is one of the main treatment options, producing outcomes comparable to surgery for localized disease, but it involves complex technologies and procedures that patients often find confusing.

Patients facing a prostate cancer diagnosis actively seek information online. Good education helps patients understand their treatment, reduces anxiety, and encourages participation in their own care. However, the medical information available on the internet is often unreliable or difficult to understand.

Large language models (LLMs) like ChatGPT are AI systems that can generate human-like text responses to questions. Since their public release, these tools have attracted enormous interest as potential sources of personalized health information, allowing patients to ask specific questions and receive detailed, immediate answers.

Despite this promise, the use of LLMs specifically for educating prostate cancer patients about radiotherapy had not been rigorously studied from both the clinician and patient perspective. This study aimed to fill that gap.

TL;DR: Prostate cancer patients need reliable and accessible education about radiotherapy, and AI chatbots like ChatGPT have emerged as potential tools to help meet this need.
Pages 3-4
Putting Five AI Chatbots to the Test

Researchers developed six questions based on common inquiries from prostate cancer patients seen in clinical practice. The topics covered included what radiotherapy is, how it compares to surgery, its side effects, its impact on quality of life, preparations before treatment, and follow-up care after treatment.

These questions were posed to five different large language models: ChatGPT-4, ChatGPT-4o (both from OpenAI), Gemini (Google), Copilot (Microsoft), and Claude (Anthropic). Each was tested five times to check for consistency in responses.

The quality of responses was evaluated from two angles. First, five experienced radiation oncologists independently rated each response on a five-point scale for relevance, correctness, and completeness. Second, 35 prostate cancer patients who had recently completed radiotherapy evaluated ChatGPT-4's responses for comprehensibility, accuracy, relevance, trustworthiness, and overall usefulness.

Readability was measured using the Flesch Reading Ease Index, a formula that accounts for average sentence length and the number of syllables per word. Higher scores indicate easier reading, with scores above 60 considered appropriate for a general audience.

TL;DR: Five AI chatbots were tested using realistic patient questions about prostate radiotherapy, with both expert clinicians and actual patients independently rating the quality of the responses.
Pages 4-6
Clinicians Find Most Responses Relevant and Correct

Across all five AI systems, radiation oncologists rated the responses as generally relevant, with scores ranging from 4.2 to 4.7 out of 5. No response from any LLM was rated as irrelevant by any reviewer, indicating a baseline level of clinical appropriateness.

Correctness scores ranged from 3.8 to 4.5, with Claude AI and both ChatGPT versions scoring highest. ChatGPT-4o was the only model with no responses rated as incorrect by any reviewer. Errors that did occur included factual inaccuracies such as claiming that small tattoos are applied before radiotherapy (Claude), or that transrectal ultrasound is routinely performed before radiation (Gemini).

Completeness showed the most variation between models. ChatGPT-4, ChatGPT-4o, and Claude were rated as complete, with scores of 4.0, 3.9, and 4.2 respectively. Copilot and Gemini received neutral ratings of 3.2 and 2.8 respectively. Gemini's response about treatment preparation was the lowest rated, as it was the only one that failed to mention bladder and bowel preparation.

Statistical testing confirmed significant differences between models in relevance and completeness, though not in correctness. The only significant pairwise difference in completeness was between Claude and Gemini.

TL;DR: All five AI models produced generally relevant responses, but significant differences emerged in completeness, with ChatGPT and Claude outperforming Gemini and Copilot.
Pages 6-7
Patients Respond Positively Despite AI's Formal Writing Style

Thirty-five prostate cancer patients at a German university hospital reviewed ChatGPT-4's responses after completing their radiotherapy treatment. The median patient age was 73 years, and patients had received various forms of hypofractionated or standard radiotherapy.

The patient ratings were strikingly positive. 94 percent found the information easy to understand, and 86 percent said it did not contain medical terms that were too difficult. This is notable because the Flesch Reading Ease Index identified the same text as objectively difficult to read.

Most patients (89 percent) found the information accurate and relevant to their experience with prostate radiotherapy, and 91 percent said it matched their personal treatment experience. 76 percent expressed confidence in the information from ChatGPT, and 80 percent said it would have helped them feel better informed before or during treatment.

77 percent of patients said they would use ChatGPT for future medical questions. However, 26 percent were neutral or disagreed with this statement, suggesting that not all patients are ready to rely on AI for health information.

TL;DR: Patients who had just completed prostate radiotherapy rated ChatGPT responses as easy to understand, accurate, and helpful, with most saying they would use AI for future health questions.
Pages 7-8
The Gap Between Readability Scores and Patient Experience

An intriguing finding was that patients rated ChatGPT's text as easy to understand, even though the Flesch Reading Ease Index scores indicated it was difficult, ranging from 24 to 39 for the various models. A score of 60 to 70 is considered appropriate for a general audience.

A likely explanation is that patients were surveyed after their treatment, meaning they were already familiar with the topic and the terminology. Prior exposure to the same information through a standardized information sheet may have made the AI-generated text feel more familiar and comprehensible than it would to a newly diagnosed patient.

Researchers found that prompting ChatGPT-4o to use simpler language improved its Flesch score from 24 to 44, demonstrating that prompt engineering, adjusting the way questions are phrased, can meaningfully improve the accessibility of AI-generated health information.

The study also highlights the risk of AI hallucinations, where the model presents incorrect information as factual. Errors in this study were mostly imprecise rather than completely invented, but even minor inaccuracies can confuse patients making important medical decisions.

TL;DR: Patients familiar with radiotherapy found AI responses understandable despite low readability scores, and the accuracy of AI output can be improved by adjusting how questions are posed.
Page 8
LLMs as a Supplement, Not a Replacement for Doctors

An important framing in this study is that AI chatbots are positioned as supplementary information tools rather than replacements for physician consultations. ChatGPT itself encouraged patients to discuss questions with their oncologist, and the study design reflected this role by presenting AI information alongside standard care.

Interestingly, a previous study found that ChatGPT responses to patient questions were rated as more empathetic than physician responses in an online forum, countering the common criticism that AI lacks human warmth. LLMs may actually be better suited than formal written materials for engaging patients emotionally.

Ethical concerns remain, including patient privacy, the risk of harmful misinformation, the lack of standardized evaluation criteria for AI medical content, and the fact that LLMs are trained on general internet text rather than specialized medical literature. These limitations warrant caution when deploying AI for health education at scale.

All five tested models consistently reminded patients to consult their treatment team, suggesting that even without explicit programming for safety, these systems include appropriate caveats. This behavior should be maintained and reinforced as LLMs are adopted in healthcare settings.

TL;DR: AI chatbots are best positioned as supplements to clinical care, helping patients feel more informed and engaged while always directing them back to their medical team.
Pages 8-9
Promise with Room to Grow

Large language models demonstrated genuine promise as patient education tools in the context of prostate cancer radiotherapy. Both clinicians and patients responded positively to the information these systems generated.

However, significant gaps remain. Readability needs to improve so that patients with less prior medical knowledge can fully benefit. Accuracy must be made more consistent, particularly for nuanced questions about treatment preparation and follow-up care where errors can cause confusion or poor preparation.

The field is advancing rapidly, and results from studies like this reflect only a snapshot in time. As new model versions are released and as researchers learn more about how to craft effective prompts, the performance of LLMs in medical education is expected to improve substantially.

Future research should expand patient evaluation to include newly diagnosed patients who have not yet received treatment, test performance across multiple languages, and develop validated scoring criteria specifically designed for assessing AI-generated medical content.

TL;DR: LLMs show real promise for prostate cancer patient education but need improvements in readability and factual accuracy before they can be reliably deployed in clinical settings.
Citation: Open Access, . Available at: PMC11839798.