Performance of large language models (LLMs) in providing prostate cancer information

BMC Urol 2024 AI 8 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Pages 1-2
Why Prostate Cancer Patients Turn to AI Chatbots

Prostate cancer (PCa) is the second most common cancer in men worldwide, and its diagnosis and management involve considerable complexity. Patients frequently seek additional information beyond what they receive during clinical visits, turning to online resources and increasingly to AI-powered chatbots.

Large language models (LLMs) are a class of artificial intelligence trained on vast amounts of text data to generate human-like responses. The best-known examples include ChatGPT (developed by OpenAI) and Google Bard, both of which became widely accessible to the public in 2022 and 2023 respectively.

The complexity of prostate cancer -- spanning topics like PSA testing, digital rectal examination (DRE), biopsy, active surveillance, radiation therapy, radical prostatectomy, and androgen deprivation therapy -- makes it a meaningful test case for evaluating whether AI chatbots can reliably support patient education.

This study was the first to directly compare the performance of ChatGPT-3.5, ChatGPT-4, and Google Bard specifically for prostate cancer patient education, evaluating them on accuracy, comprehensiveness, readability, and consistency of responses.

TL;DR: Prostate cancer's diagnostic and treatment complexity drives patients toward AI chatbots, making evaluation of these tools' accuracy and usefulness a pressing question.
Page 2
How the Study Was Designed and Conducted

Researchers compiled 52 questions about prostate cancer from reputable public health and oncology organizations, including the American Society of Clinical Oncology (ASCO), the CDC, Prostate Cancer UK, and the Prostate Cancer Foundation. Questions spanned general knowledge, diagnosis, treatment, and screening or prevention.

Each question was submitted to all three AI systems on a fixed date (July 31, 2023), and responses were independently reviewed by two board-certified urologists. A third urologist resolved any disagreements. Reviews were benchmarked against guidelines from the NCCN, AUA, and European Association of Urology.

Responses were rated on a 3-point accuracy scale (correct, mixed, or completely incorrect) and a 5-point comprehensiveness scale. Readability was assessed using two validated formulas: the Flesch Reading Ease (FRE) score and the Flesch-Kincaid Grade Level (FKGL), which estimate how easy a text is to read.

A stability analysis was also performed by submitting 30 selected questions three times each to assess whether each AI produced consistent answers across repeated queries.

TL;DR: Fifty-two common prostate cancer questions were submitted to three AI chatbots and independently assessed by urologists for accuracy, completeness, readability, and consistency.
Pages 3-4
Accuracy: How Often Did Each AI Get It Right?

Overall, ChatGPT-3.5 answered correctly in 82.7% of cases, ChatGPT-4 in 78.8%, and Google Bard in 63.5%. While the difference in overall accuracy was not statistically significant (p = 0.100), meaningful gaps emerged in specific categories.

For general knowledge questions, the gap was striking: ChatGPT-3.5 scored 88.9% correct, ChatGPT-4 scored 77.8%, and Google Bard scored only 22.2% -- a statistically significant difference (p = 0.018). This suggests Bard struggled considerably with foundational cancer knowledge.

For diagnosis-related questions, both ChatGPT-3.5 and Google Bard achieved 100% accuracy while ChatGPT-4 scored 80%. Treatment questions showed no significant differences (p = 0.496), with ChatGPT-4 leading at 85.2%.

In the screening and prevention category, all three models performed similarly with no statistically significant differences (p = 0.884), ranging from 63.6% to 81.8% accuracy.

TL;DR: All three AI models answered most prostate cancer questions correctly, though ChatGPT models outperformed Google Bard -- especially on general knowledge questions.
Pages 4-5
Comprehensiveness and Readability: Depth vs. Clarity

ChatGPT-4 produced the most comprehensive responses, rated comprehensive or very comprehensive in 71.1% of cases, compared to 40.4% for ChatGPT-3.5 and 50% for Google Bard. This difference was statistically significant (p = 0.028), suggesting ChatGPT-4 provides more thorough explanations.

When it came to readability, the results were reversed. Google Bard generated the most accessible text, achieving the highest Flesch Reading Ease score of 54.7 (compared to 40.3 for ChatGPT-4 and 34.8 for ChatGPT-3.5) and the lowest Flesch-Kincaid Grade Level of 10.2, meaning a high school sophomore could read it comfortably.

By contrast, ChatGPT-3.5 responses required roughly a college graduate reading level (FKGL of 14.0), and ChatGPT-4 responses fell in between (FKGL of 12.3). This means ChatGPT's more detailed responses come at the cost of being harder for average patients to understand.

The pattern reflects a genuine tension in health communication: more comprehensive information is often harder to read. Google Bard achieved better accessibility but delivered less thorough content, while ChatGPT-4 was the most complete but required a higher reading level.

TL;DR: ChatGPT-4 provided the most thorough answers while Google Bard was easiest to read, revealing a trade-off between depth and accessibility in AI health communication.
Page 5
Consistency: Do AI Answers Stay the Same When Asked Again?

A key practical concern about AI chatbots is whether they give consistent answers to the same question asked multiple times. The study tested stability by submitting 10 questions to each model three times and comparing responses.

All three models were highly consistent: ChatGPT-4 and Google Bard had 100% consistent responses across all repeated questions. ChatGPT-3.5 showed one inconsistency -- an answer to a screening question that was less accurate on the first attempt than on the second and third -- giving it a 90% consistency rate.

This high level of response stability is an encouraging sign for patient education applications, suggesting that different patients asking the same question at different times are likely to receive similar answers rather than contradictory guidance.

TL;DR: All three AI models demonstrated near-perfect consistency when the same prostate cancer questions were asked multiple times, strengthening confidence in their reliability.
Pages 6-8
What These Results Mean for Patient Education

The study's findings suggest that AI chatbots can serve as a useful supplement for patient education, particularly for the large volume of routine informational inquiries that patients often direct at healthcare providers. Managing this steady stream of patient messages has been identified as a significant contributor to physician burnout.

Each model has distinct strengths. ChatGPT-4 is best for thorough, detailed information needed by patients who want to fully understand their situation. Google Bard is more accessible and may be better suited for patients with lower health literacy or those seeking quick, straightforward summaries.

One important insight from the discussion is the concept of qualia -- the philosophical idea that subjective, lived experience cannot be captured in objective knowledge. AI may be able to learn every fact about prostate cancer, but it cannot replicate the empathy, clinical judgment, and tactile understanding that comes from treating patients. This means AI is better framed as an assistant to physicians rather than a replacement.

The authors also note that Google Bard showed relatively poor accuracy on general knowledge questions, scoring only 22.2% correct versus over 80% for both ChatGPT versions. This may reflect limitations in Bard's inferential and clinical reasoning abilities, as identified in prior studies comparing its diagnostic skills to those of physicians.

TL;DR: AI chatbots can meaningfully support patient education but work best as physician assistants, with each model offering different trade-offs between depth and accessibility.
Pages 8-9
Limitations and Practical Considerations for Clinical Use

A key limitation is that ChatGPT's training data cuts off at September 2021, meaning its knowledge may not reflect the latest treatment guidelines or clinical trial results. Prostate cancer management evolves rapidly, and outdated information could mislead patients about current standard-of-care options.

The study tested only 52 questions, which, while covering the most commonly asked topics, does not represent the full range of inquiries a patient might have. More nuanced or personalized questions -- such as those about specific drug interactions or individualized treatment decisions -- were not evaluated.

All responses were generated on a single date, meaning the evaluation captured only a snapshot of each AI's capabilities. AI models are regularly updated, so findings may not reflect current performance. The study also notes that Google Bard refused to answer one question entirely, which introduces another potential limitation in real-world use.

Despite these limitations, the authors emphasize that LLMs have the potential to democratize access to medical knowledge -- particularly for patients in underserved areas or those facing long wait times for specialist care, where AI chatbots could help bridge critical information gaps.

TL;DR: Outdated training data, limited question scope, and the inability to handle personalized queries are key limitations that must be considered when deploying AI for patient education.
Page 9
Conclusions and Future Directions

The study concluded that all three AI chatbots -- ChatGPT-3.5, ChatGPT-4, and Google Bard -- are capable of generating accurate, reasonably comprehensive, and readable information about prostate cancer. None produced primarily incorrect answers, and all were highly consistent across repeated queries.

While AI cannot replace healthcare professionals, this research supports a role for LLMs as patient education and pre-consultation support tools. They can answer common questions efficiently, allowing clinical staff to focus on more complex interactions that require human judgment and empathy.

Future research should evaluate more personalized and context-specific questions, test whether providing clinical context in a prompt improves response quality, and study how different patient populations interpret and act on AI-generated health information.

TL;DR: ChatGPT and Google Bard can reliably support prostate cancer patient education, though future studies should explore personalized queries and the real-world impact on patient understanding.
Citation: Open Access, . Available at: PMC11342655.