Can the ChatGPT and other large language models solve the questions and concerns of patient with prostate cancer

J Transl Med 2023 AI 6 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Page 1
Testing AI Chatbots on Prostate Cancer Questions

Large language models (LLMs) like ChatGPT have attracted enormous attention in medicine as potential tools for patient education and information access. But how accurately and helpfully do they actually respond to the questions real prostate cancer patients ask? This letter-style study from Renji Hospital, Shanghai, set out to find a systematic answer.

The researchers designed 22 questions based on patient education guidelines from the CDC and the clinical reference platform UpToDate, supplemented by their own clinical experience. The questions covered screening, prevention, treatment options, and postoperative complications, and ranged from basic factual queries to complex clinical scenario questions requiring analysis and interpretation.

Five state-of-the-art LLMs were tested: ChatGPT (both free and paid Plus versions), YouChat, NeevaAI, Perplexity (both concise and detailed modes), and Chatsonic. Three experienced urologists jointly evaluated each response across five dimensions: accuracy, comprehensiveness, patient readability, provision of humanistic care, and response stability.

An important distinction among these models: ChatGPT was trained on data up to September 2021 and cannot access the internet, while YouChat, NeevaAI, Perplexity, and Chatsonic are internet-connected and can retrieve current information. A key question was whether real-time web access would give internet-connected models a performance advantage over the offline-trained ChatGPT.

TL;DR: Three urologists evaluated five AI chatbots on 22 prostate cancer questions covering screening, treatment, and postoperative care to assess whether LLMs can accurately serve as patient information tools.
Pages 1-2
Accuracy Results: Above 90% for Most Models

Most LLMs achieved accuracy above 90% across all 22 questions -- a higher baseline than many might expect for AI systems answering complex medical questions. ChatGPT achieved the highest overall accuracy rate, with the free version performing slightly better than the paid Plus version. NeevaAI and Chatsonic were the exceptions, falling below the 90% threshold.

For basic factual questions -- such as 'What is prostate cancer?' or 'What is PSA?' -- most LLMs performed well. Accuracy was highest for questions with clear, definitive answers that appear frequently in standard medical education materials.

Performance dropped significantly on harder questions requiring contextual analysis, nuance, or synthesis of multiple clinical factors. Examples of difficult questions included: 'My prostate was totally removed by surgery -- why is my PSA still high?' and 'Which is better for prostate cancer -- apalutamide or enzalutamide?' These questions require interpreting a patient's specific situation rather than recalling standard facts.

Importantly, internet-connected models did not outperform the offline ChatGPT despite their ability to access current literature. This suggests that model training quality matters more than real-time data access for answering well-defined patient questions -- the underlying language model architecture and training data determine whether a model can reason through clinical scenarios, not merely whether it can retrieve the latest publications.

TL;DR: Most LLMs exceeded 90% accuracy on basic prostate cancer questions, with ChatGPT performing best overall, though all models struggled with complex scenario-specific questions requiring clinical reasoning.
Pages 1-2
Comprehensiveness, Readability, and Humanistic Care

In assessing comprehensiveness, the LLMs generally performed well on most questions. They could articulate the significance of different PSA levels, acknowledge that PSA is a screening marker and not a definitive diagnosis, compare treatment options with their respective advantages and disadvantages, and consistently recommend that patients discuss findings with their own doctors.

Comprehensiveness failures were typically errors of omission rather than commission. Perplexity, for example, failed to mention screening as an important strategy in prostate cancer prevention. Some models gave imprecise guidance on PSA testing frequency -- recommending a case-by-case approach without specifying age-appropriate screening intervals that guidelines explicitly define.

Readability was satisfactory for most models. Responses were generally written at a level accessible to non-specialist patients. The exception was NeevaAI, which -- reflecting its search-engine-based architecture -- tended to reproduce excerpts from medical literature without summarizing or explaining the content in patient-friendly language.

All LLMs demonstrated humanistic care for the question about life expectancy, consistently reassuring patients about prostate cancer's relatively favorable survival compared to other cancers. However, they showed limited empathy or emotional sensitivity in other responses. The models cannot ask follow-up questions, cannot adapt to a patient's emotional state, and cannot provide the contextual comfort that a human clinician would in a difficult conversation.

TL;DR: LLMs were broadly comprehensive and readable for most questions, but struggled with detailed protocol specifics, showed limited empathy beyond standard reassurance, and failed on search-engine architectures to explain rather than cite.
Page 2
Where LLMs Made Errors and Why

Analysis of incorrect or incomplete answers revealed two main error types. The most common was the inclusion of outdated or incorrect information. In one notable example, a model claimed that open surgery is more commonly performed than robot-assisted surgery for radical prostatectomy -- a factually incorrect statement that reflects the training data's historical composition rather than current clinical practice, where robotic approaches now predominate in many countries.

A second example of inaccuracy involved comparing apalutamide and enzalutamide -- two androgen-blocking drugs used in advanced prostate cancer. A model provided incorrect information about the approved indications for each drug, a medically significant error because these drugs are prescribed for different disease states and the distinction matters clinically.

A subtler but important error type was context misapplication. Several models, having been trained that 'PSA is not the final diagnostic test for prostate cancer,' applied this statement incorrectly when answering a question about why PSA remains elevated after complete prostate removal (prostatectomy). In that context, PSA is being used to detect cancer recurrence, not to diagnose initial cancer -- a clinically important distinction the models failed to recognize.

These errors highlight a fundamental limitation of current LLMs: they are pattern-matchers trained on text, not reasoners trained on clinical logic. They can retrieve and combine information well, but struggle to recognize when standard information does not apply in a specific patient context -- exactly the kind of situational judgment that clinical training emphasizes.

TL;DR: LLM errors fell into three categories: outdated information reflecting stale training data, factual inaccuracies on specific drug indications, and context misapplication where correct general knowledge was applied to the wrong clinical situation.
Pages 2-3
Potential for Democratizing Medical Knowledge

Despite their limitations, the study concludes that LLMs have genuine potential as patient education tools. For the majority of basic questions that prostate cancer patients commonly ask -- about symptoms, risk factors, screening, diagnosis, and treatment options -- most tested models provided accurate, readable, and reasonably comprehensive answers.

The most compelling potential application is in medical deserts -- communities with limited access to specialist urology care, longer waiting times for appointments, or geographic or financial barriers to medical consultation. For patients in these situations, an LLM that can accurately answer 90% of basic questions and consistently advise consulting a doctor for the remainder represents a meaningful improvement in information access over having no guidance at all.

LLMs could also support shared decision-making by helping patients understand their treatment options before appointments, generate informed questions for their physicians, and process the substantial amount of information that cancer diagnoses require patients to absorb in a short time. A well-informed patient is better positioned to participate meaningfully in discussions with their care team.

The authors emphasize that LLMs at their current capability cannot replace physicians. They cannot gather additional history by asking follow-up questions, cannot examine a patient, cannot access medical records, and cannot provide the individualized reasoning a specialist brings to complex cases. They are best understood as a supplement to medical care -- providing information accessibility -- rather than a substitute for clinical expertise.

TL;DR: LLMs show real promise for expanding basic cancer information access, particularly for underserved populations without easy access to specialists, while remaining clearly unsuited to replace clinical judgment for complex individual cases.
Page 3
Summary and Implications for AI in Patient Communication

This study provides a first systematic evaluation of multiple LLMs on a standardized set of prostate cancer patient questions. The central finding is that most leading LLMs can answer the majority of basic prostate cancer questions accurately and in patient-accessible language -- a capability that, while imperfect, is genuinely useful for information dissemination at scale.

The finding that offline ChatGPT outperformed internet-connected models is counterintuitive but important. It suggests that the quality of the underlying language model -- its ability to understand and reason about medical questions -- matters more than the recency of information available to it. For well-established clinical topics like prostate cancer screening and treatment, the core information has not changed so rapidly that real-time access provides a decisive advantage.

Key limitations of the current evaluation include its single time-point design (all responses generated on one day in February 2023), the relatively small number of questions, and the fact that LLM capabilities are evolving rapidly. Models that perform below 90% today may improve substantially with future training iterations.

The broader implication is that LLMs represent a new category of patient-facing health information tool -- more conversational and adaptive than static websites, but not yet reliable enough for clinical decision support without physician oversight. As these systems improve, establishing quality standards for AI-generated patient education content -- similar to standards already applied to patient brochures and websites -- will be an important priority for oncology and medical education communities.

TL;DR: LLMs can answer most basic prostate cancer patient questions accurately and readably, with model training quality mattering more than internet access, but current limitations in contextual reasoning mean physician oversight remains essential.
Citation: Open Access, . Available at: PMC10115367.