Distribution-shift detection in skin cancer AI: SAGE improves accuracy but excludes most dark-skin patients
A team at Oregon Health & Science University (OHSU) built SAGE, a detector that combines three uncertainty signals to flag, before any diagnosis, skin lesion photos that differ too much from the ones a model was trained on. Tested across five datasets spanning seven countries, SAGE almost perfectly separates completely unrelated images and correctly catches 92% of shifts that combine a change of imaging device with new diseases, and filtering images by its score improves a downstream malignancy-prediction model's accuracy. But that filter, calibrated to keep 90% of training-like cases, retains only 16.4% of dark-skin patients' photos — the very group the tool was meant to protect better.
The context
Deep learning models for skin cancer diagnosis post impressive numbers on benchmark datasets like HAM10000 — tens of thousands of dermoscopic photos, an imaging technique that uses a dermatoscope (a lit magnifier placed on the skin) to better visualize a lesion's structure. The well-documented problem is that this performance collapses as soon as the model meets images unlike its training data: a different camera, a smartphone instead of a dermatoscope, a patient population with more varied skin tones, or simply a disease absent from the original training set. This phenomenon has a name in machine learning: distribution shift, or out-of-distribution (OOD) data — data that departs from what the model has learned.
Until now, most tools for detecting this shift relied on a single signal — the model's confidence in its own prediction, or an image's distance from its training set. The OHSU team starts from a simple observation: a single signal poorly captures very different kinds of shift (a new camera type has nothing to do with a new disease), and proposes combining several. The other blind spot it documents is demographic: most public dermatology datasets overrepresent lighter skin (Fitzpatrick types I to IV), exposing darker-skinned patients to a double risk — underrepresentation in training and, potentially, under-detection of problematic cases.
The method
SAGE (Supervised Autoencoders for Generalization Estimates) is built on a ResNet-50, a convolutional neural network architecture standard in computer vision, pretrained on ImageNet and reused here as an encoder: it compresses each image into a 256-number vector summarizing its visual content. That vector feeds three modules. A decoder attempts to reconstruct the original image from the compressed vector — the worse the reconstruction, the more likely the image is unusual. A classifier predicts the diagnostic category and outputs a confidence score. And a search over the 20 nearest neighbors in the compressed space measures how much the image resembles examples already seen during training. The three signals — reconstruction error, classifier confidence, distance to neighbors — are converted into probabilities and combined via a geometric mean into a single score between 0 and 1: the higher it is, the further the image departs from the training distribution.
The model was trained on HAM10000 (10,015 dermoscopic images, Australia and Austria), then tested on four independent datasets representing increasing shift: HIBA (Argentina, 1,616 images, a mix of dermoscopy and smartphone photos), UFES (Brazil, 2,255 images, clinical smartphone photos only), DDI (Stanford, USA, 656 images, including 36 diagnoses absent from HAM10000, some of them rare cancers), and MILK10K (the 2025 ISIC challenge dataset, 10,480 images, five countries). A dataset entirely unrelated to medicine, Caltech-101 (photos of everyday objects), serves as an extreme control. Downstream, the authors measured the effect of SAGE-score filtering on a separate malignancy-prediction model — a pretrained Inception v3 that classifies each lesion as benign or malignant (basal cell carcinoma, squamous cell carcinoma, or melanoma).
The results
For shift detection, measured by AUROC (the area under the curve indicating how well the score separates a normal image from an out-of-distribution one), SAGE reaches 1.00 against Caltech-101 (everyday objects obviously have nothing to do with a skin lesion), 0.92 on mixed shifts combining device change and new diseases, 0.88 on device change alone (dermatoscope to smartphone), and 0.81 on images that are still fairly close to the training distribution. The median score logically climbs with distance from training: 0.63 on the HAM10000 test set, 0.74 for HIBA, 0.79 for UFES, 0.80 for MILK10K, up to 0.96 for DDI — the most distant dataset, with the rarest diagnoses.
Clinical translation: for malignancy cases entirely absent from training (cutaneous T-cell lymphoma, Kaposi sarcoma, metastatic carcinoma), 81% of the diagnostic model's false negatives — cancers it would have wrongly classified as benign — are correctly flagged as suspect by SAGE and thus withheld from an automated decision. Out of 100 such cases a diagnostic model working alone would miss, roughly 81 would be caught by the filter. But filtering has a direct cost on the downstream malignancy model's overall accuracy: for dark-skin patients, AUROC rises from 0.68 unfiltered to 0.78 after filtering — a genuine improvement — but only 16.4% of that group's images pass the filter calibrated to keep 90% of training-like cases. For light-skin patients, AUROC rises from 0.75 to 0.77, without the corresponding coverage rate being reported in the paper.
The paper also documents what raises the SAGE score independently of any pathology: presence of a measuring ruler (+7.2%), camera flash (+5.8%), non-skin background (+7.4 to +10.5% depending on density), hair (+5.1%), and the combination of hair and background (+11.5%). The ruler effect is stronger on dark skin (+9.8%) than on light skin (+6.1%). After filtering out low-quality images, the bias gap between skin tones (measured by Cliff's delta, an indicator of the gap between two distributions) shrinks from −0.434 to −0.206, without disappearing.
What is good
Three signals instead of one, validated on a controlled gradient of shifts. Many OOD-detection papers test a single type of change. Here the authors explicitly distinguish modality shift (dermatoscope vs. smartphone), class shift (new diseases), and their combination, and show SAGE stays competitive across all four settings — not just on the easiest case (non-medical objects).
Manual annotation that hunts for causes, not just symptoms. 4,527 images were manually annotated across twelve criteria (hair density, contrast, flash, ruler, background, blur) by an observer blinded to diagnosis, malignancy, and skin type. That work is what makes it possible to identify precise photographic artifacts as the cause of a high score, rather than settling for an aggregate number.
Code, weights, and data published under an open license. The full SAGE code and trained models are available on GitHub (github.com/pdxgx/sage) under a CC BY 4.0 license, and all images used are public. Independent replication is possible today, which remains rare in this field.
What is less good
A population bias the filter reduces without solving — and which paradoxically worsens in terms of coverage. The SAGE score itself is structurally higher for darker skin because of correlated artifacts (ruler, background) more common in the corresponding datasets. Direct consequence: at the threshold chosen to preserve 90% of training-like cases, 83.6% of dark-skin patients' images are excluded from automated diagnosis. The AUROC gain (0.68 to 0.78) is real, but it applies to a very small residual fraction of the group — a textbook case of a misleading metric: presenting a performance improvement without giving equal weight to the coverage rate that comes with it hides the fact that, in practice, the tool excludes the majority of the patients it was meant to protect better from automated decision-making.
Shortcut learning documented but not fixed at the source. The model reacts to elements that have nothing to do with pathology — a ruler, a flash, a room background — rather than to the lesion alone. The authors name and quantify it precisely, which is to their credit, but the proposed fix (filtering after the fact) treats the symptom: it removes risky images rather than preventing the model from learning these spurious correlations upstream.
A single diagnostic architecture tested, on a training set that remained modest. The downstream malignancy model is a single Inception v3 published in 2023 — a comparator no longer state of the art in 2026. Nothing indicates the same gains and the same coverage costs would reproduce with a more recent architecture or a larger training set than the 7,207 images used here. The authors themselves flag a risk of data leakage between training and test sets — different lesions from the same patient possibly ending up on both sides — without being able to fully rule it out.
What this changes
For the research community, SAGE sets a reproducible method for testing a shift detector across a controlled gradient of severities, and its main strength — catching 81% of false negatives on rare cancers never seen in training — deserves to be reproduced with other downstream diagnostic models. The next necessary step is already flagged by the authors: fixing demographic bias at training time rather than compensating for it with a filter that disproportionately excludes dark-skin patients.
For clinicians, nothing changes in immediate practice: SAGE is a research tool, not a validated clinical device. It does illustrate a principle transferable to any dermatology diagnostic AI: a high confidence score can reflect photo quality (ruler, flash, background) rather than diagnostic clarity, and a poorly calibrated image-quality filter can deprive certain patients of an automated opinion without that showing up in aggregate metrics.
For patients and the general public, the study concretely shows why a skin-diagnosis AI trained mostly on light skin remains less reliable, or less available, for other skin tones — a gap that is not solved simply by adding a safety filter at the output, but by diversifying the data at the source, from training onward.
Further reading
The paper: Multi-criterion uncertainty estimation improves skin cancer distribution shift detection and malignancy prediction, W. Max Schreyer, Ravi Samatham, Elizabeth Berry, Reid F. Thompson et al. (Oregon Health & Science University; VA Portland Healthcare System; Washington University in St. Louis), npj Digital Medicine, September 24, 2026, DOI 10.1038/s41746-026-03293-y. Code and models: github.com/pdxgx/sage.
On subgroup bias in dermatology imaging, see also our decryption of selective prediction with MC-Dropout on HAM10000. On automated subtyping of skin cancers, see our decryption of biopsy-free basal cell carcinoma subtyping.