Results of the 2016 International Skin Imaging Collaboration ISBI Challenge: Comparison of the Accuracy of Computer Algorithms to Dermatologists for the Diagnosis of Melanoma from Dermoscopic Images

J Am Acad Dermatol 2018 AI 5 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Pages 1-2
AI vs. Dermatologists: The 2016 ISIC Challenge

Study purpose: The 2016 International Skin Imaging Collaboration (ISIC) challenge at ISBI was the first large-scale, public benchmark comparing AI computer algorithms to board-certified dermatologists for melanoma detection from dermoscopic images.

Competition design: Twenty-five algorithm teams submitted 79 automated entries. Eight board-certified dermatologists independently evaluated the same 100 test images, which included 55 non-melanoma and 45 melanoma cases drawn from the ISIC archive.

Primary metric: Performance was measured by the area under the receiver-operating characteristic curve (AUC), allowing fair comparison across different classification thresholds between algorithms and human readers.

Significance: This challenge established the first structured methodology for head-to-head comparison of AI and human clinical experts in skin lesion diagnosis, creating a reproducible benchmark referenced by subsequent years of ISIC challenges.

TL;DR: The 2016 ISIC challenge was the first rigorous public benchmark pitting 25 teams of AI algorithms against 8 dermatologists on 100 dermoscopy images to measure melanoma detection accuracy.
Pages 2-3
Challenge Dataset and Evaluation Protocol

Test set composition: The 100-image test set was balanced at 45% melanoma prevalence - far higher than the 5-10% prevalence in real clinical dermoscopy practice - which inflated all AUC scores relative to real-world deployment.

Dermatologist evaluation: Eight board-certified dermatologists with varying subspecialty expertise (general dermatology, dermoscopy, melanoma surgery) reviewed images independently without clinical history, providing binary melanoma vs. non-melanoma decisions.

Algorithm evaluation: Each of the 25 participating teams could submit up to 10 algorithmic runs. Algorithms were ranked primarily by AUC, with secondary metrics including sensitivity, specificity, and average precision.

Fusion approach: A post-hoc fusion algorithm was applied to the top-performing algorithm outputs by averaging probability scores, testing whether combined ensemble predictions could outperform individual models.

TL;DR: The challenge used a 100-image test set with 45% melanoma prevalence, AUC as the primary metric, and allowed each of 25 teams up to 10 algorithm submissions to be compared against 8 independent dermatologists.
Pages 3-4
Key Performance Findings: AI vs. Human Experts

Top algorithm performance: The highest-performing single algorithm achieved an AUC of 0.807, while the best dermatologist achieved an AUC of 0.798, indicating the leading AI was essentially equivalent to the best-performing human reader.

Average dermatologist performance: Dermatologists achieved a mean AUC of 0.712 (range 0.631-0.798), while the top 10 algorithms all scored above the human average, with the best AI outperforming 7 of 8 dermatologists.

Fusion algorithm advantage: Combining the outputs of the top-scoring algorithms through a fusion approach yielded an AUC of 0.860, substantially exceeding both any single algorithm (0.807) and any individual dermatologist (0.798), suggesting ensemble methods are a key strategy for AI improvement.

Sensitivity-specificity tradeoff: The best algorithm matched dermatologist sensitivity at approximately 82% while maintaining higher specificity, indicating AI could potentially maintain clinical safety while reducing unnecessary biopsies.

TL;DR: The top AI algorithm (AUC 0.807) outperformed 7 of 8 dermatologists (mean AUC 0.712), and a fusion of top algorithms reached AUC 0.860, significantly exceeding any individual reader.
Pages 4-5
What These Results Mean for Dermatology Practice

AI as a screening aid: The results suggest AI algorithms could serve as a triage or screening tool, flagging suspicious lesions for dermatologist review, especially in settings where specialist access is limited.

Human-AI collaboration: Rather than replacing dermatologists, the data suggest AI performs best as a complementary tool - dermatologist review of AI-flagged cases could combine human contextual reasoning with algorithmic pattern recognition.

Performance gaps in real practice: The artificially high 45% melanoma prevalence in the test set inflates AUC compared to the 3-10% prevalence in clinical practice; real-world performance would likely differ substantially from challenge metrics.

Benchmark value: Annual standardized benchmarks like ISIC allow transparent tracking of AI progress over time, letting clinicians and regulators assess whether AI tools are improving toward clinical deployment thresholds.

TL;DR: AI results from the 2016 challenge suggest algorithms could serve as screening or second-opinion tools, but real-world deployment requires accounting for much lower clinical melanoma prevalence than used in the challenge.
Page 5
Caveats and the Path Forward

Dataset limitations: The 100-image test set is small and curated from high-quality ISIC archive images; real dermoscopy practice includes variable image quality, non-standard lighting, and rare lesion subtypes not well represented.

Challenge design biases: The 45% melanoma prevalence is 5-10x higher than typical clinical settings, preventing direct extrapolation of challenge AUC scores to expected clinical performance.

Dermatologist context: Human dermatologists normally have access to clinical context - patient age, skin tone, lesion history, personal/family history - none of which were provided in this image-only challenge, potentially disadvantaging human readers.

Future goals: Subsequent ISIC challenges (2017, 2018, 2019, 2020) expanded to larger datasets, multi-class diagnosis, and patient-level metadata to address these limitations and move AI evaluation closer to real-world conditions.

TL;DR: The 2016 challenge used a small, high-prevalence test set without clinical context, limiting real-world extrapolation; subsequent ISIC challenges have progressively addressed these design constraints.
Citation: Open Access, 2018. Available at: .