Subtyping basal cell carcinoma without a biopsy: what an AUC of 0.784 in dermatoscopy is actually worth
A team at the Aristotle University of Thessaloniki trained a pre-trained vision transformer to separate aggressive basal cell carcinomas from less serious forms using a single dermatoscopic image, on 1,271 images collected in one Greek department between 2018 and 2022. The model reaches 0.784 AUC against 0.755 for a ResNet-50, and 70.0% sensitivity at 70% specificity, where a published study reported 56.4% for trained dermatologists. But the human comparator comes from a different image set, the operating point was fixed to make the comparison possible, no external validation was run, and the paper never states how many patients lie behind its 1,271 images.
The context
Basal cell carcinoma is by far the most common cancer in humans: the authors note it accounts for nearly 80% of skin cancer diagnoses. It almost never metastasises, but it destroys locally, and above all it is not a homogeneous entity. Pathology distinguishes several histological subtypes, and that distinction drives treatment. The so-called aggressive forms — infiltrative, morpheaform, micronodular, metatypical — spread irregularly beneath the skin, with extensions the eye cannot see; they call for Mohs micrographic surgery, a layer-by-layer excision with intraoperative margin control, slow and costly. Superficial forms, by contrast, can often be treated with cryotherapy, imiquimod, photodynamic therapy or simple excision.
The practical problem is that this triage currently requires a diagnostic biopsy first: an invasive, painful procedure that costs money and time, and delays treatment. Hence the question this preprint poses: does dermatoscopy — examining the lesion through a lit magnifier that suppresses surface reflection from the skin — carry enough information to predict the subtype without going through the laboratory?
Clinically the question is not new. The dermatoscopic literature has described criteria associated with aggressive subtypes, and a study cited by the authors measured what human experts extract from them: thirty readers, before and after dedicated training, classifying 235 basal cell carcinomas. Their sensitivity for spotting aggressive forms rose from 42.3% to 56.4% after training, at a constant specificity of around 71%. In other words, even trained, an expert dermatologist misses more than four aggressive tumours in ten by eye. That is the low bar this paper sets out to beat — and according to the authors, no prior work had combined dermatoscopy and deep learning on this specific subtyping task.
The method
A binary task, not subtyping. Despite the title, the model does not predict the histological subtype. The authors group their lesions into three families — superficial (275 images), non-superficial non-aggressive (726, mostly nodular) and aggressive (270) — then merge the first two into a single "low risk" class of 1,001 images. The model therefore answers one question: aggressive or not? That is clinically defensible, since this is the boundary that triggers Mohs surgery, but it changes what the word "subtyping" means.
The model. The authors start from BEiT-Large, a vision transformer (ViT): an architecture that cuts the image into small squares — here 16 × 16 pixels on a 512 × 512 input, a 32 × 32 token grid — and treats them like the words of a sentence, each square able to "look at" every other. Roughly 300 million parameters, 24 layers, 16 attention heads. The starting weights come from pre-training by masked image modelling (parts of the image are hidden and the model learns to reconstruct them) on natural images — the exact corpus is never named in the paper, an awkward omission for reproducibility.
Two adaptation stages. First domain adaptation: the model is fully retrained for 100 epochs on 53,826 images from the public ISIC archive across nine skin lesion classes (melanoma, nevus, keratosis, squamous cell carcinoma and others). Then a second full retraining on the 1,271 Thessaloniki images: 10 epochs, learning rate 10⁻⁵ with cosine decay, batches of 16, minority-class oversampling to balance each batch, standard augmentations (flips, rotations, shifts, brightness and contrast variation). The comparators are a ResNet-18 and a ResNet-50 trained "with the same setup and computational budget".
The data. A private set, assembled at the First Department of Dermatology of the Aristotle University between January 2018 and December 2022, on Caucasian patients, with a Dermlite Foto dermatoscope at 10-fold magnification. The biopsy-confirmed histological subtype is the reference standard; mixed subtypes and incomplete excisions are excluded. Fitzpatrick phototype is not reported, patient background is described only by the word "Caucasian", and — a point we return to below — the number of patients is never stated.
The evaluation. Stratified five-fold cross-validation, repeated three times, giving fifteen splits; 20% of each fold's training set serves as validation, and the checkpoint is selected on validation AUC. No external centre. No prospective evaluation. All data is retrospective.
The results
In machine learning metrics. AUC (area under the ROC curve: the probability that a randomly drawn aggressive tumour scores higher than a randomly drawn non-aggressive one) comes to 0.784 ± 0.026 for BEiT, against 0.755 ± 0.018 for ResNet-50 and 0.735 ± 0.023 for ResNet-18. The ± figures are standard deviations across the fifteen splits, not confidence intervals. At an operating point fixed arbitrarily at 70% specificity, sensitivity reaches 70.0% ± 6.6 for BEiT, 67.4% ± 4.3 for ResNet-50 and 61.6% ± 5.8 for ResNet-18. These are the only published metrics: no accuracy, no F1, no positive predictive value, no confusion matrix, no per-subtype results, no ROC curves, no statistical test.
In clinical translation. This is where the number becomes meaningful. In the dataset composition, 21.2% of lesions are aggressive. Take 1,000 basal cell carcinomas at that same proportion: 212 aggressive, 788 not. At the chosen operating point, the model flags roughly 384 as suspicious. Of those, 148 are genuinely aggressive and 236 are not: those 236 patients would be routed to Mohs surgery they do not need. And 64 aggressive tumours would fall below threshold, hence towards inappropriate conservative treatment, with the recurrence risk that entails. One in three flagged decisions is correct; roughly one aggressive tumour in three is missed.
Against humans. On the same 1,000 lesions, the trained dermatologists of the cited study (56.4% sensitivity at 70.7% specificity) would catch about 120 aggressive tumours against 148 for the model, at a comparable false positive count. The fourteen-point sensitivity gap therefore represents, in this arithmetic, about 28 more aggressive tumours correctly routed per 1,000 lesions. That is not nothing. But, as we are about to see, it is not a comparison either.
What is good
The clinical question is the right one, and it is well framed. Much computational dermatology keeps hammering at detection — benign versus malignant — even though that step is largely handled and the bottleneck has moved on. Here, the choice of an "aggressive versus not" boundary maps exactly onto a real, binary treatment decision: Mohs or not Mohs. The potential gain is not a score, it is a biopsy avoided and a delay shortened. Few skin imaging papers identify this precisely the point in the care pathway where their model would slot in.
The algorithmic comparator is honest and the authors do not oversell themselves. ResNet-18 and ResNet-50 are trained with the same setup and the same computational budget, which makes the 0.03 AUC gap interpretable. More importantly, the authors do not dress 0.784 up as a success: the abstract calls it a "preliminary investigation", the conclusion an "initial proof-of-concept exploration", and the limitations — single centre, Caucasian population, uncontrolled human comparison — are stated plainly rather than buried. In a field where AUCs of 0.95 are regularly presented as clinic-ready, that restraint deserves noting.
The data constraint is handled with a fitting method. With 270 positive cases, training a 300-million-parameter model ought to fail. The intermediate domain adaptation on ISIC's 53,826 public images is the right methodological answer: bring the model into the world of skin lesions first, then fine-tune on the rare task. Choosing a masked-pretrained transformer over a CNN trained from scratch runs in the same direction, and the fact that the gap over the ResNets is modest rather than spectacular is consistent with the dataset size.
What is less good
Data leakage is not excluded, and cannot be. The paper never says whether cross-validation separates folds by patient or only by image. The terms "patient-level", "lesion-level", "group" and "GroupKFold" appear nowhere, the stratification variable is unnamed, and above all the number of patients is not reported — only 1,271 images. Yet in dermatoscopy it is common for the same lesion to be photographed several times, and for one patient to carry several basal cell carcinomas, often of the same subtype. If two images of the same tumour land on either side of a fold, the model is not generalising, it is recognising. This is the failure mode that most systematically inflates published AUCs in medical imaging, and nothing in the text rules it out.
Missing metadata betrays a heterogeneous assembly. The demographics table contains an anomaly the authors do not comment on: in the aggressive group, 158 of 270 lesions (58.5%) have neither sex nor anatomical site recorded, against 32% in the superficial group and 25.8% in the non-superficial group. The exact same count is missing for both variables, suggesting a block of 158 aggressive cases drawn from a data source distinct from the rest. If that source also differs in device, settings or period — and the paper contradicts itself here, announcing a single dermatoscope in section II then "different dermatoscopic devices" in section III — the model may learn a provenance signature rather than tumour morphology. The distribution of anatomical sites points the same way: 37.1% of superficial lesions are on the trunk against 9.3% of aggressive ones, 51.4% of non-superficial ones on the head and neck against 25.1% of superficial ones. A model that guessed the body region from the image would already score respectably without ever looking at the tumour. No audit — domain classifier, saliency map, error breakdown by source — is run. We described this mechanism in the context of dataset fusion in screening mammography, and the effect of missing metadata on hidden bias in subgroup auditing in medical imaging.
The human comparison is not one, and the paper half-admits it. The dermatologists' 56.4% sensitivity comes from a separate study, run on 235 different basal cell carcinomas, with thirty readers and a different protocol. The 70% specificity operating point was not chosen for a clinical reason: it was chosen, as the authors explicitly write in the table caption, to allow comparison with that study. So the threshold that aligns the x-axis is selected, then the gap on the y-axis is announced. The authors acknowledge that "different dataset compositions preclude definitive conclusions" — but the abstract sentence asserts performance "superior to previously-reported human reader performance". Add two details: the introduction gives 71.7% specificity for readers before training where the table gives 71.1%, and the clinician co-author is himself lead author of one of the dermatoscopic references the paper relies on, with that overlap declared nowhere.
Nothing is verifiable, and the ethics are undocumented. No code repository, no published weights, no data availability statement: the Thessaloniki set is private and nothing indicates it will be released. There is no conflict of interest declaration, and above all no mention of ethics committee approval or consent, for retrospective clinical images linked to age, sex and anatomical site. The only declared funding is the Special Account for Research Funds of the Aristotle University (reference 90169/2024). Finally, this preprint has not been peer reviewed, and the complete absence of ablations — notably on the ISIC adaptation stage, whose contribution is therefore unknown — leaves open where the 0.03 gap actually comes from.
What it changes
For the research community. The useful result is not 0.784, it is the mapping of a ceiling. A 300-million-parameter transformer, properly domain-adapted, tops out around 0.78 AUC on 1,271 images from a single centre — three hundredths above a ResNet-50, roughly what one expects from an architecture change when data is the limiting factor. The message, stated by the authors themselves, is that the bottleneck is the size and diversity of the annotated corpus, not model capacity. The logical next step is therefore not a bigger model but a multicentre, multi-ethnic dataset with a declared patient-level split, and a human reading protocol run on the same images. Until that last point is done, the claim "better than dermatologists" remains unverifiable.
For clinicians. Nothing changes today. The false negative rate — roughly one aggressive tumour in three missed — is incompatible with a decision to skip biopsy, because the error is not symmetric: an infiltrative carcinoma treated with cryotherapy recurs, sometimes years later, and the salvage operation is more mutilating. The only cautious reading of the result is the opposite of what the title suggests: not "replace the biopsy", but possibly prioritise which lesions to biopsy or send straight to micrographic surgery when theatre access is constrained. That prioritisation requires calibrated probabilities, which the paper does not provide, and validation on the local population, which the paper does not allow. Cross-centre generalisation in dermatoscopy was examined more finely in other external validation work on skin cancer.
For patients and the public. The intuition to keep is that Black, brown and mixed skin is absent from this work: the cohort is described as exclusively Caucasian and phototype is not even reported. A subtyping tool trained this way carries no guarantee of working elsewhere, and basal cell carcinomas on darker phototypes often present in pigmented, hence visually different, form. More broadly, resist the reading the title invites: "biopsy-free" describes an ambition, not a result. Today the biopsy remains the only way to know which type of basal cell carcinoma one has, and a score of 0.78 does not come close.
Further reading
The preprint: Deep Learning for Biopsy-Free Subtyping of Basal Cell Carcinoma from Dermatoscopic Images, arXiv:2609.07180 [cs.CV], DOI 10.48550/arXiv.2609.07180, submitted 7 September 2026, under CC BY 4.0. Authors: Alexandros Papadopoulos, Chrysa Episkopou, Ioannis Sarafis and Anastasios Delopoulos (School of Electrical and Computer Engineering, Aristotle University of Thessaloniki) and Aimilios Lallas (First Department of Dermatology, Aristotle University of Thessaloniki). Funding: Special Account for Research Funds of the Aristotle University of Thessaloniki, reference 90169/2024. No conflict of interest declaration and no ethics approval statement appear in the manuscript. The reused architecture is BEiT (Bao et al., arXiv:2106.08254), and the domain adaptation corpus comes from the ISIC archive. The human comparator is the study by Camela and colleagues published in 2024 in the Journal of the American Academy of Dermatology, covering thirty readers and 235 basal cell carcinomas; the dermatoscopic criteria relied on trace back to the 2014 work of Lallas and colleagues. On the place of micrographic surgery in managing aggressive subtypes, the European EADO guidelines and the NCCN guidelines on non-melanoma skin cancers are the baseline references.