The Question This Study Asks Can computer algorithms from an international melanoma detection challenge improve dermatologist diagnostic accuracy? Most prior AI-vs-dermatologist studies focused on raw performance comparisons in isolated test settings. This study from Memorial Sloan Kettering went further, using a controlled reader study to test not only head-to-head performance, but whether substituting AI decisions for human decisions in low-confidence cases could improve overall accuracy.
Study Design The study used 150 dermoscopy images (50 melanomas, 50 nevi, 50 seborrheic keratoses) randomly selected from the ISIC 2017 challenge test dataset. Eight specialist dermatologists and nine dermatology residents evaluated all images online, classifying each lesion and reporting diagnostic confidence on a 0-6 Likert scale. Their performance was then compared to the top-ranked algorithm from 23 participating teams.
Novel Imputation Analysis The study's key innovation was an imputation analysis: for cases where clinicians reported low diagnostic confidence (Likert score 0-3), the algorithm's classification was substituted for the human's. This tested a practical scenario where AI serves not as a replacement, but as a targeted second opinion for uncertain cases - the most realistic model of clinical deployment.
Dataset and Challenge The ISIC 2017 challenge used 2,750 dermoscopy images from the ISIC Archive: 521 melanomas (19%), 1,843 nevi (67%), and 386 seborrheic keratoses (14%). Images were split into training (n=2,000), validation (n=150), and test (n=600) datasets. Twenty-three algorithm teams competed, all using neural networks. Algorithms were ranked by ROC area under the curve for melanoma classification.
Reader Study Participants Eight dermatologists specializing in skin cancer and nine residents from four countries participated. The dermatologists had a mean of 14 years of post-residency experience and 14.5 years using dermoscopy. Readers classified each lesion as melanoma, nevus, or seborrheic keratosis, indicated a management decision (biopsy or observe), and rated their confidence. Readers were blinded to clinical metadata, histopathology, and prior evaluations.
Imputation Threshold Confidence scores of 0-3 (out of 6) were designated as 'low confidence' and triggered algorithm imputation. This constituted 51% of resident evaluations and 26.6% of dermatologist evaluations, reflecting the inherent uncertainty both groups experienced. The algorithm was pre-dichotomized at a 90% sensitivity threshold before imputation to prioritize melanoma detection.
Head-to-Head Performance The top-ranked algorithm achieved a ROC area of 0.87 for melanoma classification. Dermatologists achieved ROC area 0.74 and residents 0.66. All comparisons were statistically significant (p less than 0.001). This means the algorithm had substantially better overall discriminatory ability than all human readers - whether measured by the area under the ROC curve or at specific sensitivity/specificity operating points.
Specificity at Matched Sensitivity At the dermatologists' operating sensitivity of 76.0%, the algorithm achieved a specificity of 85.0% versus dermatologists' 72.6% (p=0.001). For management decisions (biopsy vs. observe), at the dermatologists' sensitivity of 89.0%, the algorithm specificity was 61% versus dermatologists' 51.1% (p=0.02). In both scenarios, the algorithm demonstrated statistically significantly higher specificity at the same level of melanoma detection.
Improvement Over 2016 Challenge Comparing to the same eight dermatologists who participated in the 2016 ISIC challenge, the performance gap between the top algorithm and dermatologists had widened. This suggests algorithms are improving year over year, likely driven by larger and more diverse training datasets and advances in network architectures.
Resident Improvement After algorithm imputation for low-confidence resident evaluations (51% of all resident responses), resident sensitivity increased from 56.0% to 72.9%, and overall correct classification increased from 69.4% to 72.6%. The dramatic improvement in sensitivity (+16.9 percentage points) was particularly important for melanoma detection, where missing a case has severe consequences. There was a small decrease in specificity (76.3% to 72.6%).
Dermatologist Improvement After imputation for dermatologist low-confidence evaluations (26.6% of responses), sensitivity increased from 76.0% to 80.8% and specificity held steady at 72.6% to 72.8%. Overall correct classification rose from 73.8% to 75.4%. The improvement was more modest than for residents, reflecting both the smaller proportion of imputed evaluations and the higher baseline confidence of experienced specialists.
The Clinical Interpretation The authors argue that the most realistic model of AI use is not replacement but targeted assistance: a clinician uncertain about a diagnosis would seek a second opinion, and AI could provide that second opinion more efficiently than a colleague. The data supports this model - when AI is applied precisely where human certainty is lowest, it consistently improves outcomes.
Curated Dataset Bias The 150-image test set did not include banal (routine, obviously benign) lesions or uncommon melanoma presentations. In real clinical practice, the vast majority of skin lesions seen by dermatologists are ordinary - the dataset's 33% melanoma prevalence is far higher than real-world rates, which inflates the apparent clinical impact of the algorithm.
Absence of Clinical Context Readers evaluated images without patient age, personal or family history of melanoma, lesion symptoms, or other contextual information that strongly influences real diagnostic decisions. Algorithms also lack this context, which may partially explain why both were evaluated in the same artificial setting. Adding clinical metadata could change both human and algorithm performance unpredictably.
Lack of External Validation The algorithm was tested only on ISIC 2017 data - the same distribution it was trained on. External validation on images from different dermatoscopes, institutions, and patient populations is needed before claiming generalizability. The authors note the precedent of MelaFind, an FDA-approved device that demonstrated performance in reader studies but was eventually discontinued - a reminder that positive reader study results do not guarantee real-world success.
Real-World Prospective Trials The authors call for future studies demonstrating clinical utility in real-world settings. This means prospective studies using consecutive unselected patients, including all skin lesion types presented for evaluation, evaluated with full clinical context, and measuring patient outcomes rather than just image classification accuracy.
Optimal Algorithm Thresholds The study pre-set the algorithm at 90% sensitivity for imputation. Different clinical settings require different thresholds - a community screening clinic prioritizes sensitivity to avoid missed cancers, while a specialized melanoma center might tolerate lower sensitivity to reduce unnecessary biopsies. Future research should determine the optimal operating thresholds for specific clinical scenarios.
Expanding the ISIC Archive The ISIC initiative is building the world's largest publicly available skin image archive. As this archive grows with more images, more diverse patient populations, and clinically relevant metadata, future ISIC challenges can address current limitations. Continuous public challenges allow the research community to track algorithm progress against a consistent benchmark and identify where AI improvements matter most.