Study purpose: The 2016 International Skin Imaging Collaboration (ISIC) challenge at ISBI was the first large-scale, public benchmark comparing AI computer algorithms to board-certified dermatologists for melanoma detection from dermoscopic images.
Competition design: Twenty-five algorithm teams submitted 79 automated entries. Eight board-certified dermatologists independently evaluated the same 100 test images, which included 55 non-melanoma and 45 melanoma cases drawn from the ISIC archive.
Primary metric: Performance was measured by the area under the receiver-operating characteristic curve (AUC), allowing fair comparison across different classification thresholds between algorithms and human readers.
Significance: This challenge established the first structured methodology for head-to-head comparison of AI and human clinical experts in skin lesion diagnosis, creating a reproducible benchmark referenced by subsequent years of ISIC challenges.
Test set composition: The 100-image test set was balanced at 45% melanoma prevalence - far higher than the 5-10% prevalence in real clinical dermoscopy practice - which inflated all AUC scores relative to real-world deployment.
Dermatologist evaluation: Eight board-certified dermatologists with varying subspecialty expertise (general dermatology, dermoscopy, melanoma surgery) reviewed images independently without clinical history, providing binary melanoma vs. non-melanoma decisions.
Algorithm evaluation: Each of the 25 participating teams could submit up to 10 algorithmic runs. Algorithms were ranked primarily by AUC, with secondary metrics including sensitivity, specificity, and average precision.
Fusion approach: A post-hoc fusion algorithm was applied to the top-performing algorithm outputs by averaging probability scores, testing whether combined ensemble predictions could outperform individual models.
Top algorithm performance: The highest-performing single algorithm achieved an AUC of 0.807, while the best dermatologist achieved an AUC of 0.798, indicating the leading AI was essentially equivalent to the best-performing human reader.
Average dermatologist performance: Dermatologists achieved a mean AUC of 0.712 (range 0.631-0.798), while the top 10 algorithms all scored above the human average, with the best AI outperforming 7 of 8 dermatologists.
Fusion algorithm advantage: Combining the outputs of the top-scoring algorithms through a fusion approach yielded an AUC of 0.860, substantially exceeding both any single algorithm (0.807) and any individual dermatologist (0.798), suggesting ensemble methods are a key strategy for AI improvement.
Sensitivity-specificity tradeoff: The best algorithm matched dermatologist sensitivity at approximately 82% while maintaining higher specificity, indicating AI could potentially maintain clinical safety while reducing unnecessary biopsies.
AI as a screening aid: The results suggest AI algorithms could serve as a triage or screening tool, flagging suspicious lesions for dermatologist review, especially in settings where specialist access is limited.
Human-AI collaboration: Rather than replacing dermatologists, the data suggest AI performs best as a complementary tool - dermatologist review of AI-flagged cases could combine human contextual reasoning with algorithmic pattern recognition.
Performance gaps in real practice: The artificially high 45% melanoma prevalence in the test set inflates AUC compared to the 3-10% prevalence in clinical practice; real-world performance would likely differ substantially from challenge metrics.
Benchmark value: Annual standardized benchmarks like ISIC allow transparent tracking of AI progress over time, letting clinicians and regulators assess whether AI tools are improving toward clinical deployment thresholds.
Dataset limitations: The 100-image test set is small and curated from high-quality ISIC archive images; real dermoscopy practice includes variable image quality, non-standard lighting, and rare lesion subtypes not well represented.
Challenge design biases: The 45% melanoma prevalence is 5-10x higher than typical clinical settings, preventing direct extrapolation of challenge AUC scores to expected clinical performance.
Dermatologist context: Human dermatologists normally have access to clinical context - patient age, skin tone, lesion history, personal/family history - none of which were provided in this image-only challenge, potentially disadvantaging human readers.
Future goals: Subsequent ISIC challenges (2017, 2018, 2019, 2020) expanded to larger datasets, multi-class diagnosis, and patient-level metadata to address these limitations and move AI evaluation closer to real-world conditions.