Skip to content
Learning

Melanoma classification from dermoscopic images

Classify dermoscopic images as benign or melanoma

A compact convolutional network built and trained in Neural Designer 8.0.1 scores all 900 images of the ISIC 2016 lesion-classification training release. On 180 held-out testing images it reaches ROC AUC 0.739. At a threshold fixed on the validation images, it flags 26 of 37 melanomas and leaves 98 of 143 benign lesions unflagged. It is an educational research example, not a diagnostic tool.

900Dermoscopic images (CC0)
0.739Testing ROC AUC
70.3%Testing sensitivity at threshold 0.11
68.5%Testing specificity at threshold 0.11

1. Clinical question and intended use

Dermoscopy magnifies the pigment structures of a skin lesion, and melanoma is the skin cancer that dermoscopic triage most needs to catch. This example asks a narrow research question: how well can a small convolutional network, trained from scratch on 540 labeled dermoscopic images, rank new images toward the melanoma label?

The example shows the complete Neural Designer workflow for image classification: importing an image folder, training a network, choosing an operating threshold on validation data and reporting every error on testing data. It does not establish clinical performance.

Licensed, traceable images

Every image comes from the CC0 ISIC 2016 release, with source and file hashes in the manifest.

Threshold fixed in advance

The operating point is chosen on validation images before the testing images are evaluated.

Complete error accounting

Report missed melanomas and false alarms, not only accuracy.

Dermatology researchBiomedical imagingModel validationMedical ML education
Intended-use boundary. Educational research example only. The model is not clinically validated and must not be used to diagnose skin lesions. A qualified clinician must assess any suspicious lesion.

2. Images, labels and endpoint

The data are the 900 dermoscopic images of the ISIC 2016 Challenge, Part 3 training release, published by the International Skin Imaging Collaboration under CC0 1.0. Each image has one label: 727 are benign and 173 are malignant (melanoma), a melanoma share of 19.2%. The endpoint is that label; the positive class is malignant.

The original JPEG images range from 576 × 542 to 4288 × 2848 pixels (median 1024 × 768). Each one is converted to RGB and resized to 300 × 300 pixels with Lanczos resampling, without cropping, so non-square lesions are slightly distorted. This is the same preparation as the 102-image OpenNN example, whose images are bit-identical to the corresponding files here.

Eight dermoscopic training images: four benign lesions in the top row and four melanomas in the bottom row
Training images after preparation. Top row: benign labels (ISIC_0000000, 0000001, 0000006, 0000008). Bottom row: malignant labels (ISIC_0000026, 0000030, 0000031, 0000049).

Neural Designer imports the folder with one subfolder per class and assigns images at random to training (60%), validation (20%) and testing (20%) subsets. The project stores this assignment.

SubsetImagesBenignMalignantMelanoma sharePurpose
Training54043810218.9%Fit the network parameters
Validation1801463418.9%Compare candidates and fix the threshold
Testing1801433720.6%Final evaluation, used once
Total90072717319.2%ISIC 2016 Part 3 training release
Provenance. The split is by image. The ISIC 2016 release does not include patient identifiers, so lesions from the same patient may fall in different subsets. The official ISIC 2016 test release (379 images) is not used here. The download includes the manifest with the SHA-256 of every source JPEG and prepared BMP.

3. Model

The network is the default image-classification architecture that Neural Designer 8.0.1 creates for 300 × 300 RGB images: one convolutional layer with eight 3 × 3 filters, max pooling and a single sigmoid output. It has 180,225 trainable parameters.

LayerInput shapeOutput shapeConfigurationParameters
Scaling300 × 300 × 3300 × 300 × 3ImageMinMax pixel scaling0
Convolutional300 × 300 × 3300 × 300 × 88 filters, 3 × 3, stride 1, same padding, ReLU224
Pooling300 × 300 × 8150 × 150 × 8Max pooling, 2 × 2, stride 20
Flatten150 × 150 × 8180,0000
Dense (classification)180,0001Sigmoid180,001

Output contract. The output is a score between 0 and 1 for the malignant class; the benign score is its complement. The scores are not calibrated probabilities. The image is flagged as suspicious when the malignant score reaches the decision threshold described in Section 5.

Network architecture: scaling, convolutional layer with eight filters, max pooling, flatten layer with 180,000 values and one sigmoid output named benign_malignant
Network architecture. Open the image to inspect the layer dimensions.

4. Training strategy

Training minimizes the cross-entropy error with the adaptive moment estimation (Adam) optimizer: learning rate 0.0001, mini-batches of 32 images, L2 regularization with weight 0.01 and 100 epochs. Training ran on an NVIDIA GeForce RTX 5070 Ti GPU.

The training and validation errors below are the mean binary cross-entropy of the malignant score over the images, without the regularization term. For reference, a constant score of 0.5 gives 0.693.

Adaptive moment estimation resultsValue
Epochs number100
Elapsed time00:00:29
Stopping criterionMaximum epochs number
Training error0.261
Validation error0.451

Training and validation cross-entropy error by epoch. Enable JavaScript to explore the chart.

Training and validation error at each epoch, as reported by Neural Designer.

The training error falls from 0.753 to 0.261, while the validation error stays between about 0.40 and 0.50 after the first epochs. The network keeps fitting the training images without improving on unseen ones: with 540 training images, the data limit what this architecture can learn. The saved model is the one at epoch 100.

5. Model selection and threshold

Five configurations were trained on the same split and compared by their final validation error. The testing images were not used for this choice.

CandidateConfigurationFinal validation error
No regularizationDefault network, Adam, learning rate 0.0001, 100 epochs0.482
L2 regularization 0.01 (selected)Same, with L2 weight 0.010.451
Higher learning rateNo regularization, learning rate 0.0010.476
Shorter trainingNo regularization, 50 epochs0.464
Hidden dense layer128 ReLU neurons after flatten (23.0 million parameters)2.989 (training failed)

The first four candidates differ by at most 0.031, which is within the epoch-to-epoch fluctuation of the validation error (up to 0.08), so the choice among them is weak; the selected model keeps the regularization. The larger network with a hidden layer of 128 neurons did not train: its error stayed near 3 and its malignant score saturated at 0 for every image checked.

Decision threshold

Melanomas are 19% of the training images, so the model’s scores are low: the median malignant score is 0.075 for benign and 0.171 for malignant testing images. The default threshold of 0.5 therefore flags very few images. The threshold was set on the validation images, as the point of the validation ROC curve closest to perfect classification: 0.11, with validation ROC AUC 0.750, sensitivity 76.5% and specificity 65.8%.

Baseline. Predicting “benign” for every image is correct for 143 of the 180 testing images (79.4% accuracy) but detects no melanoma. Accuracy alone is therefore not a useful measure here.

6. Clinical validation

The final model is evaluated once on the 180 testing images: 37 malignant (20.6%) and 143 benign. Neural Designer reports a testing ROC AUC of 0.739. Calculated from the same scores, the 95% confidence interval is about 0.64–0.84 (Hanley–McNeil).

ROC curve of the testing images with the chosen threshold marked. Enable JavaScript to explore the chart.

ROC curve on the testing images, from Neural Designer’s sensitivity and specificity table. The orange point is the threshold fixed on the validation images.

Operating point at threshold 0.11

Actual / predicted at 0.11Flagged (malignant)Not flagged (benign)Total
Malignant label261137
Benign label4598143
Total71109180
Testing metricThreshold 0.11Count-based readingThreshold 0.5
Sensitivity70.3%26 of 37 melanomas flagged10.8% (4 of 37)
Specificity68.5%98 of 143 benign lesions not flagged97.2% (139 of 143)
Precision / observed PPV36.6%26 of 71 flagged images are melanomas50.0% (4 of 8)
Observed NPV89.9%98 of 109 unflagged images are benign80.8% (139 of 172)
Accuracy68.9%124 of 180 images classified correctly79.4% (143 of 180)
F1 score0.481Harmonic mean of precision and sensitivity0.178

At 0.5 the model behaves almost like the all-benign baseline: same accuracy, and only 4 of 37 melanomas detected. At 0.11 it detects 26 of 37 melanomas, at the cost of 45 false alarms among 143 benign lesions. On the testing images, the ROC point closest to perfect classification lies at a threshold of 0.13. That value is descriptive only, because it was found on the testing data.

Clinical interpretation. The model learns a real but modest signal from the images. Eleven of the 37 testing melanomas score below the threshold, and only 36.6% of flagged images are melanomas at this 20.6% prevalence. On the training images the ROC AUC is 0.955, so performance on new images is much lower than on the images it was fitted to.

7. Workflow and reproducibility

A responsible real-world analogue would begin after a lesion has entered an approved clinical pathway. It would check image quality and format, calculate a versioned score, and route flagged and unflagged lesions alike to clinician review, with histopathology as the reference where indicated.

Calculating outputs for new images

Neural Designer’s Calculate outputs task scores one image at a time. The two testing images below show one melanoma above the threshold and one benign lesion below it.

Example 1

Dermoscopic image ISIC_0000004 (malignant label, testing subset)
Input image: malignant/ISIC_0000004.bmp (testing subset, malignant label)

Malignant score: 0.8151
Flagged: the score is above the threshold of 0.11.

ClassOutput score
benign0.1849
malignant0.8151

Example 2

Dermoscopic image ISIC_0000028 (benign label, testing subset)
Input image: benign/ISIC_0000028.bmp (testing subset, benign label)

Malignant score: 0.0209
Not flagged: the score is below the threshold of 0.11.

ClassOutput score
benign0.9791
malignant0.0209

These are illustrative outputs for images already counted in the testing metrics. The scores are uncalibrated.

Reproduce the example

The ZIP contains the 900 prepared images, the manifest, the trained Neural Designer 8.0.1 model (model/melanoma.ndm and its parameter file), the Neural Designer inference script and scores.csv with the malignant score of every image and its subset. All the metrics above can be recalculated from scores.csv. To score a new image prepared in the same way, run:

model\melanoma_predict.bat lesion.bmp --table

The script reports its predicted class at 0.5; apply the 0.11 threshold to the malignant score instead. To retrain, create an image-classification project in Neural Designer from the melanoma_dataset_bmp folder, keep the default network and use the training settings in Section 4. Neural Designer draws a new random split and initialization, so the results will vary; the saved split is listed in model/melanoma.ndm and scores.csv.

8. Safety, generalizability and governance

  • Internal random split only. Training, validation and testing images come from the same release. No external hospital, device, time period or prospective cohort is evaluated.
  • Patient grouping cannot be verified. The release has no patient identifiers, so related lesions may appear in more than one subset, which can make internal results optimistic.
  • Few melanomas in the test. With 37 testing melanomas, each missed case changes sensitivity by 2.7 percentage points, and the confidence interval of the AUC is wide.
  • Challenge selection, not a clinical population. The 19% melanoma share reflects how the challenge images were collected. Predictive values change with prevalence.
  • Image preparation. Downsampling to 300 × 300 discards fine dermoscopic structure, the resize distorts non-square images, and no color or artifact normalization is applied.
  • Uncalibrated scores. A threshold of 0.11 shows that the scores are not probabilities of melanoma.
  • No subgroup evidence. Skin type, age, sex, anatomical site and acquisition device are not available in the release, so performance across them is unknown.
  • No clinical-utility evaluation. The example does not compare with dermatologists or measure the effect on patient outcomes.
Decision boundary. Use this model for education and method research only. Any clinical use would require larger and external data, calibration, a prespecified threshold, regulatory review and evaluation within the clinical pathway.

References