The internet has become a primary source of health information for patients, particularly for older adults managing diagnoses like prostate cancer. By early 2023, ChatGPT had over one billion monthly users worldwide -- a rapid adoption that suggests many patients are already turning to AI chatbots rather than traditional search engines for medical answers.
Unlike Google, which returns a list of links for users to sift through, ChatGPT provides direct, conversational responses to questions. This creates a more efficient and accessible format for patients -- but also a riskier one, because the quality of the answer is entirely dependent on the AI's training data and cannot be verified by simply checking a source URL.
AI hallucinations -- cases where AI models generate plausible-sounding but factually incorrect or fabricated information, including false citations -- represent a known hazard in medical AI applications. OpenAI itself states that ChatGPT is not fine-tuned to provide medical information and should not be used to deliver diagnostic or treatment guidance for serious conditions.
Despite this, patients will use these tools regardless of official guidance. The important research question is: how accurate is ChatGPT when patients ask about prostate cancer, and can the technology be prompted to deliver that information in a more readable, understandable way for the general public?
Nine prostate cancer questions were identified via Google Trends, covering diagnosis, treatment, and postoperative follow-up. Each question was entered into ChatGPT 3.5, and the response was recorded. Then, each response was re-entered with a new instruction: generate a simplified version understandable at or below a sixth-grade reading level, while remaining accurate and comprehensive.
Medical quality was assessed by 53 urological experts (36 urologists and 17 urology residents) recruited through social media and prior research consent. They rated each original ChatGPT response for accuracy, completeness, and clarity using a five-point scale. A response was considered to pass the "correctness trifecta" when it scored 4 or 5 (agree/strongly agree) on all three dimensions.
Readability was measured using six validated tools -- Flesch-Kincaid Grade Level, Gunning Fog Score, SMOG Index, Coleman Liau Index, Automated Readability Index, and Flesch Reading Ease -- applied to both the original and simplified responses. The standard recommendation for patient health education materials is a sixth-grade reading level or below.
To assess how the general public perceived the simplified summaries, 514 crowdsourced workers from Amazon Mechanical Turk (MTurk) rated each layperson summary for clarity on a five-point scale. They then answered a multiple-choice comprehension question about the summary's main message -- a direct test of whether the simplified version was genuinely understood, not just rated as readable.
The urological experts rated ChatGPT's original responses fairly highly across the nine prostate cancer scenarios. Accuracy ratings ranged from 71.7% to 96.2% depending on the scenario. The best performance was on a postoperative follow-up question (Scenario 9), where 94.3% of experts rated the answer accurate. The worst performance was on a diagnostic question (Scenario 3), where only 71.7% agreed the answer was accurate.
The correctness trifecta -- accuracy, completeness, and clarity all rated 4 or 5 -- was achieved in 62.3% to 90.6% of expert responses, depending on the scenario. Treatment-related questions and follow-up questions tended to score better than diagnosis questions, suggesting ChatGPT handles established clinical information more reliably than nuanced diagnostic decision-making.
Two independent physician reviewers assessed the simplified layperson summaries: they agreed that 8 out of 9 summaries were accurate, and that 8 out of 9 provided enough information for a patient to make an informed decision. This 88.9% rate of accuracy and decision-support sufficiency suggests that the simplification prompt preserved the medical substance of the answers while making them more accessible.
Inter-rater agreement between the two reviewing physicians was strong, ranging from 88.9% to 100% across evaluated categories -- suggesting the quality assessments were reliable and not just idiosyncratic opinions of individual reviewers.
The original ChatGPT responses were written at a post-secondary reading level: the Flesch-Kincaid Grade Level averaged 12.8, the Gunning Fog Score averaged 15.8, and the Flesch Reading Ease score averaged only 36.5 (higher is better, and scores below 50 are considered difficult). These are well above the recommended sixth-grade standard for patient health materials.
After prompting ChatGPT to simplify the outputs, readability improved dramatically across all six metrics (all p less than 0.001). The mean Flesch-Kincaid Grade Level dropped to 7.4, the Gunning Fog Score dropped to 9.5, and the Flesch Reading Ease improved to 70.2 -- significantly more accessible while maintaining medical substance. This demonstrates that a simple prompt modification can yield dramatically more patient-appropriate content.
Among the 514 MTurk respondents assessing the simplified summaries, clarity ratings were high: 89.5% to 95.7% of participants rated each scenario as clear. However, when tested on actual comprehension via multiple-choice questions, performance was more variable: the easiest scenario had 87.4% of participants answering correctly, while the hardest scenario had only 63.0% -- a meaningful gap between perceived clarity and actual understanding.
This divergence between clarity ratings and comprehension scores is clinically important: patients may feel they understand AI-generated health information better than they actually do. Even simplified AI content can contain subtleties that lead to misunderstanding, and the gap between feeling informed and being informed is a known challenge in health communication research.
The study's most actionable finding is that the way ChatGPT is prompted matters enormously. Simply instructing the model to respond at a sixth-grade level produced statistically significant improvements in all readability metrics while maintaining clinical accuracy. This means that medical organizations building AI-powered patient education tools have a concrete, easy-to-implement optimization available to them right now.
Despite this, the researchers emphasize that ChatGPT was not designed for medical education and should not be used as a standalone health resource. Its accuracy is imperfect -- never reaching 100% in any scenario -- and it cannot account for individual patient circumstances, local clinical guidelines, or nuances that a specialist urologist would naturally incorporate. AI hallucinations (fabricated but plausible content) remain an unsolved risk for medical LLM applications.
The study points toward domain-specific medical AI models as the more promising direction. Tools like Google's MedPalm2, trained specifically on medical knowledge, or custom GPT models fine-tuned on current clinical guidelines and prostate cancer literature, could provide higher accuracy and better calibration to clinical standards than a general-purpose chatbot.
The authors call on AI developers building medical chatbots to integrate up-to-date clinical guidelines, include citations and source transparency, collaborate with medical experts during development, and subject outputs to ongoing validation studies. As AI tools evolve, the research community must keep pace with new versions -- what holds for ChatGPT 3.5 may not hold for later models, requiring continuous re-evaluation.
For prostate cancer patients, this study offers a nuanced message: ChatGPT can provide a reasonable starting point for understanding their diagnosis, treatment options, and follow-up expectations -- but the information should always be verified with a healthcare provider. Accuracy ranges from 71% to 96% depending on the question, meaning roughly one in four to one in twenty answers may contain inaccurate or incomplete information.
For healthcare providers and patient advocacy organizations, the study highlights an opportunity: if AI tools must exist for patient education (and patients are going to use them regardless), it is worth investing in customized, guideline-integrated, transparency-labeled chatbot interfaces rather than leaving patients to use general-purpose commercial tools with unknown accuracy profiles.
The readability findings have equity implications. At a 12th-grade reading level, ChatGPT's default outputs are inaccessible to a large proportion of the population -- including many older adult prostate cancer patients. The simple act of requesting sixth-grade-level summaries is a practical improvement, but healthcare systems should not rely on patients knowing to ask for this.
This study is one piece of a growing body of research evaluating AI in medicine. Its limitations include reliance on ChatGPT 3.5 (an older version), lack of demographic data from MTurk respondents, and focus on a limited set of nine questions. As the technology evolves, periodic re-evaluation will be essential to ensure that AI tools used in prostate cancer education remain safe, accurate, and accessible for all patients.