Skip to content
Learning

Erythemato-squamous disease classification with machine learning

Audit a compact six-class dermatology benchmark without unnecessary architecture search

This reproducible Neural Designer example classifies six erythemato-squamous disease labels from 11 clinical, 22 histopathological and one age variable in the 366-record UCI Dermatology data set. The fixed 34-to-6 softmax model classifies 70 of 73 internal test records correctly. All three errors are pityriasis rosea records assigned to seborrheic dermatitis, and no external or prospective validation is available. The artifact is for research education and model audit, not patient diagnosis.

366historical public records
73internally held-out test records
95.0%macro sensitivity across six labels
70.0%pityriasis rosea sensitivity (7/10)

1. Clinical question and intended use

The defensible purpose of this artifact is retrospective machine-learning education, software verification and critical appraisal of a classic dermatology benchmark. It reproduces six historical source labels from already collected clinical assessments and biopsy-derived microscopy scores.

The intended users are dermatology researchers, dermatopathology teams, clinical data scientists, biostatisticians and model-governance professionals. The supported action is to inspect data provenance, preprocessing, model parsimony and class-level internal test behaviour before designing a clinically valid study.

Prefer parsimony

Audit the direct 34-to-6 classifier before adding hidden neurons to an already separable benchmark.

Localize errors

Move beyond aggregate accuracy and identify the pityriasis rosea–seborrheic dermatitis confusion.

Define the evidence gap

Separate internal random-split performance from calibration, external validity and clinical utility.

Dermatology researchDermatopathologyClinical data scienceBiostatisticsModel governance
Intended-use boundary. Do not use this model to diagnose, exclude, triage or treat a skin disorder. It requires biopsy-derived histopathological inputs and has not been validated as a medical device or evaluated in a clinical workflow.

2. Cohort, measurements and endpoint

The analysis uses the UCI Dermatology data set (DOI 10.24432/C5FK5P), accessed 13 August 2026 and distributed under CC BY 4.0. UCI reports 366 records and 34 features. Patients were first assessed clinically and skin samples were subsequently evaluated microscopically for histopathological features.

The processed CSV used here contains 11 clinical variables, 22 histopathological variables, age and the six-class endpoint. Except for binary family history and linear age, UCI encodes feature degree from 0 (absent) to 3 (largest amount). These are ordinal assessments, not continuous physical measurements.

Source classMachine labelRecordsShare
Chronic dermatitiscronic_dermatitis5214.2%
Lichen planuslichen_planus7219.7%
Pityriasis rubra pilarispitiriasis_rubra_pilaris205.5%
Pityriasis roseapityriasis_rosea4913.4%
Psoriasispsoriasis11230.6%
Seborrheic dermatitisseboreic_dermatitis6116.7%

The machine-readable labels preserve spelling in the supplied project for exact reproducibility; standard clinical spelling is used in the narrative.

Distribution of 366 UCI Dermatology records across six source labels

Sample composition is not disease prevalence. Psoriasis is the largest source class (112 records); pityriasis rubra pilaris is the smallest (20).
SubsetRowsChronicLichenPRPPityriasis roseaPsoriasisSeborrheic
Training220313911336937
Selection7371966278
Testing7314143101616

The stored project uses a random 60/20/20 row split. The audit found no duplicate complete rows or duplicate input vectors. Eight records have missing age; the project applies mean replacement.

Preprocessing leakage. The exported age scaler is centred at 36.2961 years, exactly the mean of all 358 observed ages, rather than the 35.5185-year mean in the training subset. The test partition therefore influenced preprocessing and is not a fully untouched estimate. A corrected study should fit imputation and scaling on training data only.

Ten largest univariate associations between dermatology inputs and the six-class source label

These label-coding- and method-dependent univariate associations are not feature importance, causal effects or evidence that a finding is independently diagnostic.
Endpoint provenance. The accompanying paper describes records with known diagnoses, but the public repository does not provide the individual reference-standard procedure, reader count, blinding, adjudication or recruitment setting for each record. The six labels must therefore be treated as historical benchmark annotations, not a fully characterized contemporary clinical ground truth.

3. Model

All 34 inputs are scaled and connected directly to six softmax outputs. There is no hidden layer, so the model is a multinomial logistic classifier with 210 trainable parameters: 204 input weights and six biases.

This compact structure is appropriate to audit first: the data set is small, the inputs are expert-scored and the classes are already strongly separable. Additional hidden neurons would increase flexibility and selection burden without evidence that the base model needs it.

Output contract. The six exported softmax values sum to one but were not calibrated. They are model scores for the source labels, not patient-level diagnostic probabilities or measures of clinical certainty.
Direct 34-input to six-output dermatology classifier without a hidden layer
Initial and final architecture: 34 scaled inputs connect directly to six softmax scores. The diagram summarizes the six scores under the endpoint node and omits middle inputs for readability.

4. Training strategy

The stored project minimizes multiclass cross-entropy with L2 regularization weight 0.01 using the quasi-Newton optimizer. Training ends on minimum loss decrease after 35 stored epochs (0–34).

Training cross-entropy falls from 2.0043 to 0.0544. Selection cross-entropy falls from 0.2776 to its minimum of 0.0743 at epoch 7 and ends at 0.0784. The modest post-minimum increase is an overfitting signal; the exported final state is not documented as a restored best-selection checkpoint.

Training and selection cross-entropy over 35 quasi-Newton epochs

The selection curve reaches its minimum at epoch 7 while training continues. This supports checkpointing the best selection state in a future leakage-free experiment.

5. Model selection and baseline

No input selection, neuron selection or architecture search was performed. The project contains default configuration panels for growing inputs and growing neurons, but no corresponding selection task or result. The direct 34-to-6 classifier is both the initial and final model.

Testing referenceAccuracyMacro sensitivityModel complexity
Always predict psoriasis21.9% (16/73)16.7%No fitted parameters
Fixed direct softmax model95.9% (70/73)95.0%210 parameters

The base model substantially exceeds the transparent majority-class baseline. Because the result is already strong on this internal split, the next useful experiment is not a wider network: it is a leakage-free, repeated or nested evaluation followed by independent external validation.

6. Clinical validation

The exact Python export was recomputed against the 73 records marked as testing in the project after applying the project-wide age mean. This reproduces the supplied confusion table: 70 correct and three incorrect classifications.

Calibration, multiclass ROC AUC and PR AUC were not evaluated. Class counts range from only three pityriasis rubra pilaris records to 16 psoriasis and 16 seborrheic dermatitis records, so apparently perfect class sensitivities remain uncertain.

Decision rule and confusion matrix

No binary or class-specific threshold was selected. The reported operating point assigns the label with the largest of the six softmax scores (argmax); no threshold was optimized on the test subset.

Actual / predictedChronicLichenPRPPityriasis roseaPsoriasisSeborrheicTotal
Chronic dermatitis140000014
Lichen planus014000014
Pityriasis rubra pilaris0030003
Pityriasis rosea00070310
Psoriasis000016016
Seborrheic dermatitis000001616
Total141437161973
One-vs-rest labelTest prevalenceTP / FN / FP / TNSensitivity (95% Wilson CI)SpecificityObserved precision / PPV
Chronic dermatitis14/73 (19.2%)14 / 0 / 0 / 59100% (78.5–100)100%100%
Lichen planus14/73 (19.2%)14 / 0 / 0 / 59100% (78.5–100)100%100%
Pityriasis rubra pilaris3/73 (4.1%)3 / 0 / 0 / 70100% (43.9–100)100%100%
Pityriasis rosea10/73 (13.7%)7 / 3 / 0 / 6370.0% (39.7–89.2)100%100%
Psoriasis16/73 (21.9%)16 / 0 / 0 / 57100% (80.6–100)100%100%
Seborrheic dermatitis16/73 (21.9%)16 / 0 / 3 / 54100% (80.6–100)94.7%84.2%
Overall test summaryValue
Accuracy95.9% (70/73)
Macro sensitivity / balanced multiclass recall95.0%
Macro precision97.4%
Macro F10.956
Majority-class baseline accuracy21.9% (16/73)
Interpretation. On this split, errors are not diffuse: three of ten pityriasis rosea records are assigned to seborrheic dermatitis. The 100% observed sensitivities for other classes are based on 3–16 records each and have wide confidence intervals. Predictive values reflect this benchmark composition and must not be transferred to a clinical population.

7. Workflow and reproducibility

A responsible workflow for this artifact is a controlled research reproduction:

Licensed de-identified record
Schema and ordinal-scale check
Training-only preprocessing
Six uncalibrated scores
Error and uncertainty audit
Researcher review

Use batch inference only with the supplied schema. A valid record requires all 34 ordered inputs, including 22 biopsy-derived histopathological assessments. Reject unknown encodings, out-of-range ordinal values and missing fields outside the documented age-imputation reproduction path.

No interactive patient calculator is provided: manual entry of 34 expert-scored findings would invite transcription errors and imply a clinical use that this evidence does not support.

Verified reference calculation

The exact export was checked with this ordered 34-value vector:

[2, 1, 3, 3, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 2, 0, 2, 0, 2, 0, 2, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 1, 0, 36]

Source labelModel score
Chronic dermatitis0.967870
Lichen planus0.002199
Pityriasis rubra pilaris0.001761
Pityriasis rosea0.003867
Psoriasis0.006463
Seborrheic dermatitis0.017840

The Python package contains the exact export, ordered input schema, reference vector and expected scores. The Neural Designer package contains the project and exact processed CSV, preserving the stored split, base model and regenerated analyses.

8. Safety, generalizability and governance

  • Internal evidence only: there is no independent-site, temporal or prospective validation and no workflow-impact study.
  • Preprocessing leakage: the scaler uses complete-data statistics, so the random test subset is not an untouched estimate.
  • Historical provenance: the records were donated in 1997; the public metadata does not establish contemporary diagnostic criteria, recruitment, geography, reader agreement or adjudication for every record.
  • Small class samples: the test set contains only three pityriasis rubra pilaris and ten pityriasis rosea records. Perfect observed results do not imply negligible error rates.
  • Calibration and thresholds: calibration, multiclass ROC AUC, PR AUC, decision-curve utility and clinically prespecified operating points are absent.
  • Expert-dependent inputs: 22 variables require microscopy of skin samples and many fields use subjective ordinal grading. Inter-reader and inter-site reproducibility were not evaluated.
  • Missingness: eight ages are missing and reproduced using a complete-data mean. Deployment missingness mechanisms may differ.
  • Subgroups and spectrum: performance by age, sex, skin tone, disease stage, treatment status and coexisting conditions is not available.
  • Human oversight: any future study requires specialist review by dermatology and dermatopathology professionals, locked preprocessing, audit logging, an abstention pathway and confirmatory clinical procedures.

Evidence required before any clinical research transition

  1. Reconstruct the reference standard, inclusion criteria, acquisition sites and reader process from primary records.
  2. Repeat development with stratification and training-only imputation/scaling, using nested or repeated resampling for model comparison.
  3. Prespecify clinically meaningful operating points and evaluate calibration, uncertainty and abstention.
  4. Validate externally across sites, time periods, acquisition workflows and relevant demographic and disease subgroups.
  5. Compare against dermatologist and dermatopathologist assessment and evaluate prospective workflow effects before considering regulated use.
Decision boundary. This article demonstrates reproducibility, parsimony and critical evaluation of a historical benchmark. It does not establish clinical validity, clinical benefit or permission to use the model with patient data.

References

  1. İlter N, Güvenir H. Dermatology. UCI Machine Learning Repository; 1998. DOI: 10.24432/C5FK5P.
  2. Güvenir HA, Demiröz G, İlter N. Learning differential diagnosis of erythemato-squamous diseases using voting feature intervals. Artificial Intelligence in Medicine. 1998;13(3):147–165. DOI: 10.1016/S0933-3657(98)00028-1.