Prostate cancer (PCa) is the second most common cancer in men worldwide, and its diagnosis and management involve considerable complexity. Patients frequently seek additional information beyond what they receive during clinical visits, turning to online resources and increasingly to AI-powered chatbots.
Large language models (LLMs) are a class of artificial intelligence trained on vast amounts of text data to generate human-like responses. The best-known examples include ChatGPT (developed by OpenAI) and Google Bard, both of which became widely accessible to the public in 2022 and 2023 respectively.
The complexity of prostate cancer -- spanning topics like PSA testing, digital rectal examination (DRE), biopsy, active surveillance, radiation therapy, radical prostatectomy, and androgen deprivation therapy -- makes it a meaningful test case for evaluating whether AI chatbots can reliably support patient education.
This study was the first to directly compare the performance of ChatGPT-3.5, ChatGPT-4, and Google Bard specifically for prostate cancer patient education, evaluating them on accuracy, comprehensiveness, readability, and consistency of responses.
Researchers compiled 52 questions about prostate cancer from reputable public health and oncology organizations, including the American Society of Clinical Oncology (ASCO), the CDC, Prostate Cancer UK, and the Prostate Cancer Foundation. Questions spanned general knowledge, diagnosis, treatment, and screening or prevention.
Each question was submitted to all three AI systems on a fixed date (July 31, 2023), and responses were independently reviewed by two board-certified urologists. A third urologist resolved any disagreements. Reviews were benchmarked against guidelines from the NCCN, AUA, and European Association of Urology.
Responses were rated on a 3-point accuracy scale (correct, mixed, or completely incorrect) and a 5-point comprehensiveness scale. Readability was assessed using two validated formulas: the Flesch Reading Ease (FRE) score and the Flesch-Kincaid Grade Level (FKGL), which estimate how easy a text is to read.
A stability analysis was also performed by submitting 30 selected questions three times each to assess whether each AI produced consistent answers across repeated queries.
Overall, ChatGPT-3.5 answered correctly in 82.7% of cases, ChatGPT-4 in 78.8%, and Google Bard in 63.5%. While the difference in overall accuracy was not statistically significant (p = 0.100), meaningful gaps emerged in specific categories.
For general knowledge questions, the gap was striking: ChatGPT-3.5 scored 88.9% correct, ChatGPT-4 scored 77.8%, and Google Bard scored only 22.2% -- a statistically significant difference (p = 0.018). This suggests Bard struggled considerably with foundational cancer knowledge.
For diagnosis-related questions, both ChatGPT-3.5 and Google Bard achieved 100% accuracy while ChatGPT-4 scored 80%. Treatment questions showed no significant differences (p = 0.496), with ChatGPT-4 leading at 85.2%.
In the screening and prevention category, all three models performed similarly with no statistically significant differences (p = 0.884), ranging from 63.6% to 81.8% accuracy.
ChatGPT-4 produced the most comprehensive responses, rated comprehensive or very comprehensive in 71.1% of cases, compared to 40.4% for ChatGPT-3.5 and 50% for Google Bard. This difference was statistically significant (p = 0.028), suggesting ChatGPT-4 provides more thorough explanations.
When it came to readability, the results were reversed. Google Bard generated the most accessible text, achieving the highest Flesch Reading Ease score of 54.7 (compared to 40.3 for ChatGPT-4 and 34.8 for ChatGPT-3.5) and the lowest Flesch-Kincaid Grade Level of 10.2, meaning a high school sophomore could read it comfortably.
By contrast, ChatGPT-3.5 responses required roughly a college graduate reading level (FKGL of 14.0), and ChatGPT-4 responses fell in between (FKGL of 12.3). This means ChatGPT's more detailed responses come at the cost of being harder for average patients to understand.
The pattern reflects a genuine tension in health communication: more comprehensive information is often harder to read. Google Bard achieved better accessibility but delivered less thorough content, while ChatGPT-4 was the most complete but required a higher reading level.
A key practical concern about AI chatbots is whether they give consistent answers to the same question asked multiple times. The study tested stability by submitting 10 questions to each model three times and comparing responses.
All three models were highly consistent: ChatGPT-4 and Google Bard had 100% consistent responses across all repeated questions. ChatGPT-3.5 showed one inconsistency -- an answer to a screening question that was less accurate on the first attempt than on the second and third -- giving it a 90% consistency rate.
This high level of response stability is an encouraging sign for patient education applications, suggesting that different patients asking the same question at different times are likely to receive similar answers rather than contradictory guidance.
The study's findings suggest that AI chatbots can serve as a useful supplement for patient education, particularly for the large volume of routine informational inquiries that patients often direct at healthcare providers. Managing this steady stream of patient messages has been identified as a significant contributor to physician burnout.
Each model has distinct strengths. ChatGPT-4 is best for thorough, detailed information needed by patients who want to fully understand their situation. Google Bard is more accessible and may be better suited for patients with lower health literacy or those seeking quick, straightforward summaries.
One important insight from the discussion is the concept of qualia -- the philosophical idea that subjective, lived experience cannot be captured in objective knowledge. AI may be able to learn every fact about prostate cancer, but it cannot replicate the empathy, clinical judgment, and tactile understanding that comes from treating patients. This means AI is better framed as an assistant to physicians rather than a replacement.
The authors also note that Google Bard showed relatively poor accuracy on general knowledge questions, scoring only 22.2% correct versus over 80% for both ChatGPT versions. This may reflect limitations in Bard's inferential and clinical reasoning abilities, as identified in prior studies comparing its diagnostic skills to those of physicians.
A key limitation is that ChatGPT's training data cuts off at September 2021, meaning its knowledge may not reflect the latest treatment guidelines or clinical trial results. Prostate cancer management evolves rapidly, and outdated information could mislead patients about current standard-of-care options.
The study tested only 52 questions, which, while covering the most commonly asked topics, does not represent the full range of inquiries a patient might have. More nuanced or personalized questions -- such as those about specific drug interactions or individualized treatment decisions -- were not evaluated.
All responses were generated on a single date, meaning the evaluation captured only a snapshot of each AI's capabilities. AI models are regularly updated, so findings may not reflect current performance. The study also notes that Google Bard refused to answer one question entirely, which introduces another potential limitation in real-world use.
Despite these limitations, the authors emphasize that LLMs have the potential to democratize access to medical knowledge -- particularly for patients in underserved areas or those facing long wait times for specialist care, where AI chatbots could help bridge critical information gaps.
The study concluded that all three AI chatbots -- ChatGPT-3.5, ChatGPT-4, and Google Bard -- are capable of generating accurate, reasonably comprehensive, and readable information about prostate cancer. None produced primarily incorrect answers, and all were highly consistent across repeated queries.
While AI cannot replace healthcare professionals, this research supports a role for LLMs as patient education and pre-consultation support tools. They can answer common questions efficiently, allowing clinical staff to focus on more complex interactions that require human judgment and empathy.
Future research should evaluate more personalized and context-specific questions, test whether providing clinical context in a prompt improves response quality, and study how different patient populations interpret and act on AI-generated health information.