The Need Patients with cirrhosis and hepatocellular carcinoma (HCC) face complex, life-threatening conditions but often have low health literacy and unmet information needs. Accessible, reliable health information could improve adherence and outcomes.
The Study Researchers from Cedars-Sinai Medical Center systematically evaluated ChatGPT's accuracy and reproducibility in answering 164 patient-oriented questions about cirrhosis and HCC, reviewed by transplant hepatologists.
The Questions Questions came from two sources: frequently asked questions from professional societies and institutions, and real questions posted by patients and caregivers in Facebook liver disease support groups.
The Finding ChatGPT correctly answered approximately 79% of cirrhosis questions and 74% of HCC questions, but only 47% and 41% respectively were rated as 'comprehensive' - suggesting adequate but not expert-level responses.
Question Collection A total of 91 cirrhosis questions and 73 HCC questions were selected after excluding duplicate, vague, individual-specific, and non-medical questions from a pool identified from professional societies and patient Facebook groups.
ChatGPT Version Responses were generated using the ChatGPT Dec 15 version (GPT-3.5 LLM). Each question was entered twice independently using the 'New Chat' function to assess reproducibility.
Grading System Two independent board-certified transplant hepatologists graded each response on a 4-point scale: (1) Comprehensive, (2) Correct but inadequate, (3) Mixed with correct and incorrect/outdated data, (4) Completely incorrect. A blinded senior hepatologist resolved disagreements.
Additional Tests ChatGPT was also tested against published physician/trainee questionnaires on HCC surveillance, 26 AASLD-recommended quality measures for cirrhosis management, and emotional support scenarios.
Overall Accuracy ChatGPT achieved correct or comprehensive responses to 75% or more of questions in the 'basic knowledge', 'treatment', 'lifestyle', and 'others' domains. No responses were rated completely incorrect.
Weaker Domains Performance dropped in 'diagnosis' (66.7% correct or comprehensive) and 'preventive medicine' (50% correct or comprehensive), with 33.3% of diagnosis questions and 50% of preventive medicine questions containing mixed correct and incorrect/outdated information.
Quality Measures Performance On 26 AASLD-recommended cirrhosis quality measures, ChatGPT correctly answered 20 (76.9%). Strong areas included diagnostic paracentesis algorithms, spontaneous bacterial peritonitis management, and hepatic encephalopathy protocols.
Reproducibility 90.48% of all cirrhosis questions produced two similar responses from ChatGPT - demonstrating high consistency.
Overall HCC Performance ChatGPT provided correct or comprehensive responses to 74% of 73 HCC questions, with strong performance in basic knowledge (75%+ correct) and treatment categories.
Diagnosis Failures In the diagnosis category, 50% of questions contained mixed correct and incorrect information, and 33.3% were completely incorrect - a concerning pattern for clinical decision-making.
Staging System Errors ChatGPT used TNM (tumor-node-metastasis) staging in 6.7% of treatment questions instead of the BCLC (Barcelona Clinic Liver Cancer) staging system recommended by international guidelines for HCC treatment decisions.
Surveillance Knowledge Gaps When tested against published physician questionnaires on HCC screening, ChatGPT correctly answered only 4 of 8 questions in one study and just 1 of 7 in another - failing to specify correct age cutoffs for screening or appropriate imaging modalities for patients with ascites.
Patient Emotional Support When asked about coping with an HCC diagnosis, ChatGPT acknowledged the patient's emotional distress, provided clear actionable starting points for newly diagnosed patients, encouraged proactive treatment engagement, and emphasized both physical and mental well-being.
Caregiver Support For caregiver queries, ChatGPT provided organized, multifaceted practical and psychological recommendations including supporting patients with treatment adherence, finding support groups, and crucially - encouraging caregivers to also maintain their own health.
Quality of Emotional Responses Two physician reviewers assessed the emotional support responses as organized, empathetic, and offering actionable recommendations appropriate for patients and caregivers facing serious liver disease diagnoses.
Limits of Emotional AI While the responses were practical and empathetic, the reviewers acknowledged that ChatGPT's emotional support algorithms may not fully comprehend the complexity of human emotional responses in serious illness contexts.
Addressing Health Literacy Gaps Prior studies showed that only 15.7% of chronic liver disease patients knew safe acetaminophen doses. ChatGPT's conversational style and accurate basic knowledge could help bridge this dangerous knowledge gap.
Efficiency for Providers ChatGPT could draft framework responses to patient questions that physicians review and refine, rather than physicians drafting from scratch - potentially saving significant time and reducing healthcare costs.
Accessibility and Equity Free access to ChatGPT could reduce health disparities by providing reliable medical information to financially constrained patients who lack access to specialist care, potentially improving surveillance rates and treatment compliance.
Patient-Centered Care Patients empowered with better information about their liver disease can participate more effectively in shared decision-making with their medical teams, which studies show improves treatment adherence and outcomes.
Knowledge Cutoff ChatGPT's training data ended in 2021, meaning it may provide outdated treatment recommendations as guidelines evolve. Liver cancer management has seen rapid changes with new immunotherapy approvals.
Regional Guideline Blindness Without specifying geographic region, ChatGPT cannot distinguish between AASLD (American), EASL (European), or Asian regional guidelines - which differ significantly for HCC surveillance criteria, creating confusion for international patients.
Cannot Replace Expert Care Given that only 47% of cirrhosis and 41% of HCC responses were rated 'comprehensive', ChatGPT should be considered an adjunct tool that requires expert review, not a replacement for hepatologist consultation.
Future Optimization Future ChatGPT versions should be programmed to ask clarifying questions about geographic location, clinical context, and guideline preferences to generate more accurate, tailored recommendations for liver disease patients.