Audit a compact six-class dermatology benchmark without unnecessary architecture search
This reproducible Neural Designer example classifies six erythemato-squamous disease labels from 11 clinical, 22 histopathological and one age variable in the 366-record UCI Dermatology data set. The fixed 34-to-6 softmax model classifies 70 of 73 internal test records correctly. All three errors are pityriasis rosea records assigned to seborrheic dermatitis, and no external or prospective validation is available. The artifact is for research education and model audit, not patient diagnosis.
1. Clinical question and intended use
The defensible purpose of this artifact is retrospective machine-learning education, software verification and critical appraisal of a classic dermatology benchmark. It reproduces six historical source labels from already collected clinical assessments and biopsy-derived microscopy scores.
The intended users are dermatology researchers, dermatopathology teams, clinical data scientists, biostatisticians and model-governance professionals. The supported action is to inspect data provenance, preprocessing, model parsimony and class-level internal test behaviour before designing a clinically valid study.
Prefer parsimony
Audit the direct 34-to-6 classifier before adding hidden neurons to an already separable benchmark.
Localize errors
Move beyond aggregate accuracy and identify the pityriasis rosea–seborrheic dermatitis confusion.
Define the evidence gap
Separate internal random-split performance from calibration, external validity and clinical utility.
2. Cohort, measurements and endpoint
The analysis uses the UCI Dermatology data set (DOI 10.24432/C5FK5P), accessed 13 August 2026 and distributed under CC BY 4.0. UCI reports 366 records and 34 features. Patients were first assessed clinically and skin samples were subsequently evaluated microscopically for histopathological features.
The processed CSV used here contains 11 clinical variables, 22 histopathological variables, age and the six-class endpoint. Except for binary family history and linear age, UCI encodes feature degree from 0 (absent) to 3 (largest amount). These are ordinal assessments, not continuous physical measurements.
| Source class | Machine label | Records | Share |
|---|---|---|---|
| Chronic dermatitis | cronic_dermatitis | 52 | 14.2% |
| Lichen planus | lichen_planus | 72 | 19.7% |
| Pityriasis rubra pilaris | pitiriasis_rubra_pilaris | 20 | 5.5% |
| Pityriasis rosea | pityriasis_rosea | 49 | 13.4% |
| Psoriasis | psoriasis | 112 | 30.6% |
| Seborrheic dermatitis | seboreic_dermatitis | 61 | 16.7% |
The machine-readable labels preserve spelling in the supplied project for exact reproducibility; standard clinical spelling is used in the narrative.

| Subset | Rows | Chronic | Lichen | PRP | Pityriasis rosea | Psoriasis | Seborrheic |
|---|---|---|---|---|---|---|---|
| Training | 220 | 31 | 39 | 11 | 33 | 69 | 37 |
| Selection | 73 | 7 | 19 | 6 | 6 | 27 | 8 |
| Testing | 73 | 14 | 14 | 3 | 10 | 16 | 16 |
The stored project uses a random 60/20/20 row split. The audit found no duplicate complete rows or duplicate input vectors. Eight records have missing age; the project applies mean replacement.

3. Model
All 34 inputs are scaled and connected directly to six softmax outputs. There is no hidden layer, so the model is a multinomial logistic classifier with 210 trainable parameters: 204 input weights and six biases.
This compact structure is appropriate to audit first: the data set is small, the inputs are expert-scored and the classes are already strongly separable. Additional hidden neurons would increase flexibility and selection burden without evidence that the base model needs it.

4. Training strategy
The stored project minimizes multiclass cross-entropy with L2 regularization weight 0.01 using the quasi-Newton optimizer. Training ends on minimum loss decrease after 35 stored epochs (0–34).
Training cross-entropy falls from 2.0043 to 0.0544. Selection cross-entropy falls from 0.2776 to its minimum of 0.0743 at epoch 7 and ends at 0.0784. The modest post-minimum increase is an overfitting signal; the exported final state is not documented as a restored best-selection checkpoint.

5. Model selection and baseline
No input selection, neuron selection or architecture search was performed. The project contains default configuration panels for growing inputs and growing neurons, but no corresponding selection task or result. The direct 34-to-6 classifier is both the initial and final model.
| Testing reference | Accuracy | Macro sensitivity | Model complexity |
|---|---|---|---|
| Always predict psoriasis | 21.9% (16/73) | 16.7% | No fitted parameters |
| Fixed direct softmax model | 95.9% (70/73) | 95.0% | 210 parameters |
The base model substantially exceeds the transparent majority-class baseline. Because the result is already strong on this internal split, the next useful experiment is not a wider network: it is a leakage-free, repeated or nested evaluation followed by independent external validation.
6. Clinical validation
The exact Python export was recomputed against the 73 records marked as testing in the project after applying the project-wide age mean. This reproduces the supplied confusion table: 70 correct and three incorrect classifications.
Calibration, multiclass ROC AUC and PR AUC were not evaluated. Class counts range from only three pityriasis rubra pilaris records to 16 psoriasis and 16 seborrheic dermatitis records, so apparently perfect class sensitivities remain uncertain.
Decision rule and confusion matrix
No binary or class-specific threshold was selected. The reported operating point assigns the label with the largest of the six softmax scores (argmax); no threshold was optimized on the test subset.
| Actual / predicted | Chronic | Lichen | PRP | Pityriasis rosea | Psoriasis | Seborrheic | Total |
|---|---|---|---|---|---|---|---|
| Chronic dermatitis | 14 | 0 | 0 | 0 | 0 | 0 | 14 |
| Lichen planus | 0 | 14 | 0 | 0 | 0 | 0 | 14 |
| Pityriasis rubra pilaris | 0 | 0 | 3 | 0 | 0 | 0 | 3 |
| Pityriasis rosea | 0 | 0 | 0 | 7 | 0 | 3 | 10 |
| Psoriasis | 0 | 0 | 0 | 0 | 16 | 0 | 16 |
| Seborrheic dermatitis | 0 | 0 | 0 | 0 | 0 | 16 | 16 |
| Total | 14 | 14 | 3 | 7 | 16 | 19 | 73 |
| One-vs-rest label | Test prevalence | TP / FN / FP / TN | Sensitivity (95% Wilson CI) | Specificity | Observed precision / PPV |
|---|---|---|---|---|---|
| Chronic dermatitis | 14/73 (19.2%) | 14 / 0 / 0 / 59 | 100% (78.5–100) | 100% | 100% |
| Lichen planus | 14/73 (19.2%) | 14 / 0 / 0 / 59 | 100% (78.5–100) | 100% | 100% |
| Pityriasis rubra pilaris | 3/73 (4.1%) | 3 / 0 / 0 / 70 | 100% (43.9–100) | 100% | 100% |
| Pityriasis rosea | 10/73 (13.7%) | 7 / 3 / 0 / 63 | 70.0% (39.7–89.2) | 100% | 100% |
| Psoriasis | 16/73 (21.9%) | 16 / 0 / 0 / 57 | 100% (80.6–100) | 100% | 100% |
| Seborrheic dermatitis | 16/73 (21.9%) | 16 / 0 / 3 / 54 | 100% (80.6–100) | 94.7% | 84.2% |
| Overall test summary | Value |
|---|---|
| Accuracy | 95.9% (70/73) |
| Macro sensitivity / balanced multiclass recall | 95.0% |
| Macro precision | 97.4% |
| Macro F1 | 0.956 |
| Majority-class baseline accuracy | 21.9% (16/73) |
7. Workflow and reproducibility
A responsible workflow for this artifact is a controlled research reproduction:
Use batch inference only with the supplied schema. A valid record requires all 34 ordered inputs, including 22 biopsy-derived histopathological assessments. Reject unknown encodings, out-of-range ordinal values and missing fields outside the documented age-imputation reproduction path.
No interactive patient calculator is provided: manual entry of 34 expert-scored findings would invite transcription errors and imply a clinical use that this evidence does not support.
Verified reference calculation
The exact export was checked with this ordered 34-value vector:
[2, 1, 3, 3, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 2, 0, 2, 0, 2, 0, 2, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 1, 0, 36]
| Source label | Model score |
|---|---|
| Chronic dermatitis | 0.967870 |
| Lichen planus | 0.002199 |
| Pityriasis rubra pilaris | 0.001761 |
| Pityriasis rosea | 0.003867 |
| Psoriasis | 0.006463 |
| Seborrheic dermatitis | 0.017840 |
The Python package contains the exact export, ordered input schema, reference vector and expected scores. The Neural Designer package contains the project and exact processed CSV, preserving the stored split, base model and regenerated analyses.
8. Safety, generalizability and governance
- Internal evidence only: there is no independent-site, temporal or prospective validation and no workflow-impact study.
- Preprocessing leakage: the scaler uses complete-data statistics, so the random test subset is not an untouched estimate.
- Historical provenance: the records were donated in 1997; the public metadata does not establish contemporary diagnostic criteria, recruitment, geography, reader agreement or adjudication for every record.
- Small class samples: the test set contains only three pityriasis rubra pilaris and ten pityriasis rosea records. Perfect observed results do not imply negligible error rates.
- Calibration and thresholds: calibration, multiclass ROC AUC, PR AUC, decision-curve utility and clinically prespecified operating points are absent.
- Expert-dependent inputs: 22 variables require microscopy of skin samples and many fields use subjective ordinal grading. Inter-reader and inter-site reproducibility were not evaluated.
- Missingness: eight ages are missing and reproduced using a complete-data mean. Deployment missingness mechanisms may differ.
- Subgroups and spectrum: performance by age, sex, skin tone, disease stage, treatment status and coexisting conditions is not available.
- Human oversight: any future study requires specialist review by dermatology and dermatopathology professionals, locked preprocessing, audit logging, an abstention pathway and confirmatory clinical procedures.
Evidence required before any clinical research transition
- Reconstruct the reference standard, inclusion criteria, acquisition sites and reader process from primary records.
- Repeat development with stratification and training-only imputation/scaling, using nested or repeated resampling for model comparison.
- Prespecify clinically meaningful operating points and evaluate calibration, uncertainty and abstention.
- Validate externally across sites, time periods, acquisition workflows and relevant demographic and disease subgroups.
- Compare against dermatologist and dermatopathologist assessment and evaluate prospective workflow effects before considering regulated use.
References
- İlter N, Güvenir H. Dermatology. UCI Machine Learning Repository; 1998. DOI: 10.24432/C5FK5P.
- Güvenir HA, Demiröz G, İlter N. Learning differential diagnosis of erythemato-squamous diseases using voting feature intervals. Artificial Intelligence in Medicine. 1998;13(3):147–165. DOI: 10.1016/S0933-3657(98)00028-1.



