Computer Algorithms Show Potential for Improving Dermatologists' Accuracy to Diagnose Cutaneous Melanoma: Results of ISIC 2017

J Am Acad Dermatol 2020 AI 6 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Pages 1-2
AI vs. Dermatologists: The ISIC 2017 Reader Study at Memorial Sloan Kettering

The Question This Study Asks Can computer algorithms from an international melanoma detection challenge improve dermatologist diagnostic accuracy? Most prior AI-vs-dermatologist studies focused on raw performance comparisons in isolated test settings. This study from Memorial Sloan Kettering went further, using a controlled reader study to test not only head-to-head performance, but whether substituting AI decisions for human decisions in low-confidence cases could improve overall accuracy.

Study Design The study used 150 dermoscopy images (50 melanomas, 50 nevi, 50 seborrheic keratoses) randomly selected from the ISIC 2017 challenge test dataset. Eight specialist dermatologists and nine dermatology residents evaluated all images online, classifying each lesion and reporting diagnostic confidence on a 0-6 Likert scale. Their performance was then compared to the top-ranked algorithm from 23 participating teams.

Novel Imputation Analysis The study's key innovation was an imputation analysis: for cases where clinicians reported low diagnostic confidence (Likert score 0-3), the algorithm's classification was substituted for the human's. This tested a practical scenario where AI serves not as a replacement, but as a targeted second opinion for uncertain cases - the most realistic model of clinical deployment.

TL;DR: This reader study compared the ISIC 2017 top-ranked algorithm against 8 dermatologists and 9 residents on 150 dermoscopy images, and also tested whether substituting AI decisions for human decisions in low-confidence cases improved diagnostic accuracy.
Pages 2-3
ISIC 2017 Challenge Design and Reader Study Protocol

Dataset and Challenge The ISIC 2017 challenge used 2,750 dermoscopy images from the ISIC Archive: 521 melanomas (19%), 1,843 nevi (67%), and 386 seborrheic keratoses (14%). Images were split into training (n=2,000), validation (n=150), and test (n=600) datasets. Twenty-three algorithm teams competed, all using neural networks. Algorithms were ranked by ROC area under the curve for melanoma classification.

Reader Study Participants Eight dermatologists specializing in skin cancer and nine residents from four countries participated. The dermatologists had a mean of 14 years of post-residency experience and 14.5 years using dermoscopy. Readers classified each lesion as melanoma, nevus, or seborrheic keratosis, indicated a management decision (biopsy or observe), and rated their confidence. Readers were blinded to clinical metadata, histopathology, and prior evaluations.

Imputation Threshold Confidence scores of 0-3 (out of 6) were designated as 'low confidence' and triggered algorithm imputation. This constituted 51% of resident evaluations and 26.6% of dermatologist evaluations, reflecting the inherent uncertainty both groups experienced. The algorithm was pre-dichotomized at a 90% sensitivity threshold before imputation to prioritize melanoma detection.

TL;DR: The reader study enrolled 8 dermatologists and 9 residents evaluating 150 dermoscopy images with confidence ratings; algorithm classifications were substituted for human classifications in 51% of resident and 26.6% of dermatologist evaluations where confidence was low.
Page 4
Algorithm Outperforms Both Dermatologists and Residents in ROC Analysis

Head-to-Head Performance The top-ranked algorithm achieved a ROC area of 0.87 for melanoma classification. Dermatologists achieved ROC area 0.74 and residents 0.66. All comparisons were statistically significant (p less than 0.001). This means the algorithm had substantially better overall discriminatory ability than all human readers - whether measured by the area under the ROC curve or at specific sensitivity/specificity operating points.

Specificity at Matched Sensitivity At the dermatologists' operating sensitivity of 76.0%, the algorithm achieved a specificity of 85.0% versus dermatologists' 72.6% (p=0.001). For management decisions (biopsy vs. observe), at the dermatologists' sensitivity of 89.0%, the algorithm specificity was 61% versus dermatologists' 51.1% (p=0.02). In both scenarios, the algorithm demonstrated statistically significantly higher specificity at the same level of melanoma detection.

Improvement Over 2016 Challenge Comparing to the same eight dermatologists who participated in the 2016 ISIC challenge, the performance gap between the top algorithm and dermatologists had widened. This suggests algorithms are improving year over year, likely driven by larger and more diverse training datasets and advances in network architectures.

TL;DR: The top algorithm achieved ROC area 0.87 versus 0.74 for dermatologists and 0.66 for residents - at matched sensitivity, the algorithm's specificity was 12.4 percentage points higher than dermatologists, and the performance gap had grown since 2016.
Pages 4-5
Imputing AI for Low-Confidence Human Evaluations Improves Accuracy

Resident Improvement After algorithm imputation for low-confidence resident evaluations (51% of all resident responses), resident sensitivity increased from 56.0% to 72.9%, and overall correct classification increased from 69.4% to 72.6%. The dramatic improvement in sensitivity (+16.9 percentage points) was particularly important for melanoma detection, where missing a case has severe consequences. There was a small decrease in specificity (76.3% to 72.6%).

Dermatologist Improvement After imputation for dermatologist low-confidence evaluations (26.6% of responses), sensitivity increased from 76.0% to 80.8% and specificity held steady at 72.6% to 72.8%. Overall correct classification rose from 73.8% to 75.4%. The improvement was more modest than for residents, reflecting both the smaller proportion of imputed evaluations and the higher baseline confidence of experienced specialists.

The Clinical Interpretation The authors argue that the most realistic model of AI use is not replacement but targeted assistance: a clinician uncertain about a diagnosis would seek a second opinion, and AI could provide that second opinion more efficiently than a colleague. The data supports this model - when AI is applied precisely where human certainty is lowest, it consistently improves outcomes.

TL;DR: Substituting AI decisions for low-confidence human evaluations increased resident sensitivity from 56% to 73% and dermatologist sensitivity from 76% to 81%, supporting AI as a targeted second opinion rather than a replacement.
Page 5
Artificial Study Setting and Challenges to Real-World Generalization

Curated Dataset Bias The 150-image test set did not include banal (routine, obviously benign) lesions or uncommon melanoma presentations. In real clinical practice, the vast majority of skin lesions seen by dermatologists are ordinary - the dataset's 33% melanoma prevalence is far higher than real-world rates, which inflates the apparent clinical impact of the algorithm.

Absence of Clinical Context Readers evaluated images without patient age, personal or family history of melanoma, lesion symptoms, or other contextual information that strongly influences real diagnostic decisions. Algorithms also lack this context, which may partially explain why both were evaluated in the same artificial setting. Adding clinical metadata could change both human and algorithm performance unpredictably.

Lack of External Validation The algorithm was tested only on ISIC 2017 data - the same distribution it was trained on. External validation on images from different dermatoscopes, institutions, and patient populations is needed before claiming generalizability. The authors note the precedent of MelaFind, an FDA-approved device that demonstrated performance in reader studies but was eventually discontinued - a reminder that positive reader study results do not guarantee real-world success.

TL;DR: The study's artificially high melanoma prevalence, lack of clinical metadata, and absence of external validation limit direct translation to clinical practice - despite the algorithm's clear performance advantage in this controlled setting.
Page 5
What Is Needed Before AI Enters the Dermatology Clinic

Real-World Prospective Trials The authors call for future studies demonstrating clinical utility in real-world settings. This means prospective studies using consecutive unselected patients, including all skin lesion types presented for evaluation, evaluated with full clinical context, and measuring patient outcomes rather than just image classification accuracy.

Optimal Algorithm Thresholds The study pre-set the algorithm at 90% sensitivity for imputation. Different clinical settings require different thresholds - a community screening clinic prioritizes sensitivity to avoid missed cancers, while a specialized melanoma center might tolerate lower sensitivity to reduce unnecessary biopsies. Future research should determine the optimal operating thresholds for specific clinical scenarios.

Expanding the ISIC Archive The ISIC initiative is building the world's largest publicly available skin image archive. As this archive grows with more images, more diverse patient populations, and clinically relevant metadata, future ISIC challenges can address current limitations. Continuous public challenges allow the research community to track algorithm progress against a consistent benchmark and identify where AI improvements matter most.

TL;DR: Clinical utility of AI in dermatology requires prospective real-world trials with unselected patients and full clinical context, and optimizing algorithm decision thresholds for specific clinical settings - elements the controlled ISIC challenge setting cannot provide.
Citation: Open Access, 2020. Available at: PMC7006718.