After a prostate cancer diagnosis, patients frequently turn to the internet for information -- especially when anxiety makes it difficult to fully absorb what their doctor has told them. Unfortunately, much of the health content on popular platforms like YouTube, TikTok, and Instagram is of poor quality, inaccurate, or difficult to act upon.
ChatGPT, a freely available AI language model developed by OpenAI, offers a different kind of resource. Unlike search engines that return a list of links, ChatGPT generates conversational, plain-language responses to direct questions. Its answers have been shown to be difficult for patients to distinguish from those written by actual healthcare providers, suggesting it may carry meaningful trust for health advice.
However, ChatGPT does not retrieve information in real time -- its responses are based on patterns learned from text published before September 2021. This creates a risk of outdated or fabricated information (sometimes called hallucinations), where the model generates plausible-sounding but incorrect statements. Assessing when ChatGPT is reliable and when it falls short is a critical research need.
This brief study asked: does ChatGPT generate the kinds of questions newly diagnosed prostate cancer patients actually ask? And are its answers to those questions medically sound, understandable, and actionable enough to be genuinely helpful?
The researchers used ChatGPT 3.5 to generate the ten most commonly asked questions about prostate cancer, then compared that list to the top 25 Google search trends for prostate cancer in 2022. This allowed them to check whether AI-generated questions genuinely reflect what real patients search for.
ChatGPT was then given each of its own generated questions and asked to respond. A board-certified urologist reviewed every response and rated it as either clinically appropriate or inappropriate for a newly diagnosed patient.
Content quality was assessed using two validated tools: the PEMAT (Patient Education Materials Assessment Tool), which measures both understandability (can patients process and explain the information?) and actionability (can patients identify what to do?), scored 0-100% with 75% as the high-quality threshold; and the QUEST (Quality Evaluation Scoring Tool), which evaluates authorship transparency, referencing, currentness, and tone, scored out of 28.
Readability was also measured using the Flesch-Kincaid test, which estimates the grade level of reading required to understand a piece of text -- an important consideration for ensuring health information reaches patients across different educational backgrounds.
Eight out of ten questions that ChatGPT generated independently matched real patient Google searches, confirming that the AI correctly anticipates what newly diagnosed patients want to know. All ten responses from ChatGPT were rated by the urologist as clinically appropriate -- a 100% accuracy rate for correctly guiding patient understanding of their condition.
On understandability, seven of the ten responses scored above the high-quality threshold of 75%, with a mean PEMAT understandability score of 91.7%. Responses used everyday language, clear purpose statements, and organized content. Most included bullet points, and medical terms were generally explained in accessible language -- though one response mentioned BRCA gene mutations without explanation.
For actionability -- whether patients could identify what to do based on the information -- the mean PEMAT score was 76%. All responses suggested at least one action, and most referenced the patient-doctor relationship. However, three responses lacked organizational features and three would have benefited from visual aids. Some responses gave general guidance without spelling out explicit steps, which can limit patient ability to act independently.
The biggest weakness emerged with the QUEST information quality scores: the mean was only 40.4% out of a possible 100%, well below accepted standards. ChatGPT responses lacked clear authorship, citations to expert sources, and timestamps indicating when information was current -- all features that reputable health websites typically provide. Notably, all responses used cautious vocabulary acknowledging medical uncertainty, and 8 out of 10 explicitly encouraged patients to consult their physicians.
The finding that ChatGPT-generated questions matched real patient search behavior is practically significant: it suggests AI could serve as a question-generating companion for patients who are overwhelmed and unsure what to ask their doctor. This is especially valuable in early post-diagnosis consultations when anxiety is high and patients may not think clearly.
While ChatGPT outperformed social media platforms in accuracy and clinical appropriateness, its transparency limitations are real. Unlike a reputable medical website, ChatGPT does not tell you who wrote its content, when it was last updated, or what evidence it is based on. This makes it harder for patients to critically evaluate the advice they receive -- a concern especially when information may be outdated or subtly inaccurate.
The college-level reading grade of ChatGPT's responses is a meaningful equity concern. Not all patients -- particularly older adults or those with lower health literacy -- can comfortably engage with college-level text. Additionally, ChatGPT 3.5 lacks the ability to follow up dynamically as patients develop follow-on questions, limiting its usefulness for the natural back-and-forth of real patient-provider dialogue.
A critical limitation cited in the study: only 26% of ChatGPT responses to clinical prostate cancer questions based on European Association of Urology (EAU) guidelines were accurate in other studies, suggesting that while basic educational questions fare well, nuanced clinical questions carry a higher risk of misinformation. Professional oversight remains essential for any clinical AI application.
The study's authors envision ChatGPT and similar AI tools as supplementary educational companions rather than replacements for clinical consultation. Used appropriately, they could help patients arrive at appointments better prepared, more informed, and with specific questions ready -- improving the quality of patient-physician dialogue and overall satisfaction with care.
Customization represents a particularly promising opportunity: rather than relying on a general-purpose chatbot, healthcare organizations could develop disease-specific AI tools trained on validated prostate cancer patient education content, integrated with current clinical guidelines, and connected to real-time medical literature. This would address transparency and accuracy limitations while preserving the conversational accessibility that makes ChatGPT appealing.
The study also highlights an important accessibility consideration. Older adults -- who represent the majority of prostate cancer patients -- have been shown to find chatbot interfaces usable with low cognitive burden, suggesting that age should not be assumed to be a barrier. However, ensuring access for patients with limited internet literacy or those from lower-resource settings remains an open challenge for any digital health tool.
As AI continues to evolve rapidly, the authors emphasize that evaluation studies like this one need to keep pace with new model versions. ChatGPT 4.0 offers more current knowledge and customization options than version 3.5, but requires a paid subscription, raising questions about health equity in access. Ongoing benchmarking of AI tools against clinical standards will be essential to ensure patients receive safe and accurate guidance.
For a patient newly diagnosed with prostate cancer, ChatGPT can serve as a useful starting point -- a tool that anticipates common questions and provides clinically sound, readable explanations in plain language. It consistently nudges patients toward their doctors, uses cautious phrasing about medical uncertainty, and avoids the sensationalism and misinformation common on social media platforms.
The practical value is clear: patients who don't know what they don't know can use ChatGPT to generate a list of relevant questions before their first specialist appointment. The fact that 80% of AI-generated questions matched real Google search trends suggests the tool has a genuine understanding of what newly diagnosed patients care about most -- including treatment options, prognosis, side effects, and the role of genetics.
At the same time, patients and clinicians should be aware that ChatGPT's responses are not authored, referenced, or regularly updated. It cannot account for individual patient circumstances, and its accuracy drops for more detailed clinical questions. Using it wisely means treating its output as a starting conversation, not a definitive medical resource.
Looking ahead, the integration of AI into prostate cancer patient education is likely to deepen. The challenge for healthcare systems will be to harness the accessibility and conversational strengths of large-language models while addressing their known weaknesses through guideline integration, version oversight, and meaningful collaboration between AI developers and clinical specialists.