Skip to content
Learning

Primary tumor site classification

Classify recorded primary tumor sites

Use 17 recorded patient characteristics and metastatic-site indicators to classify the primary tumor site in the UCI Primary Tumor dataset. The supplied project selects an eight-neuron hidden layer and evaluates 67 internally held-out records.

339Source records
31Encoded input values
67Testing-role records
21Model outputs

1. Clinical question and intended use

This retrospective biomedical benchmark studies associations between patient descriptors, observed metastatic locations and a recorded tumor-site label. It supports reproducible research; it is not prospective metastasis prediction or a diagnostic instrument.

Recorded site labels

Model the observed primary-site annotation.

Selection chronology

Compare the initial and selected network architectures.

Rare-site review

Inspect per-site denominators and missed cases.

Computational oncologyBiostatisticsBiomedical research
Biomedical research only. Specialist review and a confirmatory clinical reference standard remain necessary.

2. Cohort, measurements and endpoint

UCI Primary Tumor contains 339 records from the Ljubljana Oncology Institute. The adapted CSV decodes source categories. There are 21 observed target categories in this copy; the upstream coding scheme lists 22 possible sites. Question-mark values are retained as explicit categories in several inputs, rather than interpreted as measured clinical findings.

Source: Primary Tumor. Dataset license: CC-BY-4.0. The downloadable ZIP includes the adapted data and attribution notices.

Dataset measureSaved value
Analysis unitpatient record
Records339
Raw variables18
Encoded model inputs31
Model outputs21
Training roles205
Validation / selection roles67
Testing roles67
Unused roles0
FieldRoleTypeCategories
classTargetCategoricalbladder, breast, cervix_uteri, colon, corpus_uteri, duodenum_and_small_intestine, esophagus, gallbladder, head_and_neck, kidney, liver, lung, ovary, pancreas, prostate, rectum, salivary_glands, stomach, testis, thyroid, vagina
ageInputCategorical30_to_59, 60_or_over, under_30
sexInputCategorical?, female, male
histologic_typeInputCategorical?, adeno, anaplastic, epidermoid
degree_of_differentiationInputCategorical?, fairly, poorly, well
boneInputBinaryno, yes
bone_marrowInputBinaryno, yes
lungInputBinaryno, yes
pleuraInputBinaryno, yes
peritoneumInputBinaryno, yes
liverInputBinaryno, yes
brainInputBinaryno, yes
skinInputCategorical?, no, yes
neckInputBinaryno, yes
supraclavicularInputBinaryno, yes
axillarInputCategorical?, no, yes
mediastinumInputBinaryno, yes
abdominalInputBinaryno, yes

Target class distribution table

MeasureSamplesShare (%)
bladder20.59
breast247.08
cervix_uteri20.59
colon144.13
corpus_uteri61.77
duodenum_and_small_intestine10.29
esophagus92.65
gallbladder164.72
head_and_neck205.90
kidney247.08
liver72.06
lung8424.78
ovary298.55
pancreas288.26
prostate102.95
rectum61.77
salivary_glands20.59
stomach3911.50
testis10.29
thyroid144.13
vagina10.29
The saved project assigns rows to training, validation and testing as shown above. This is internal record-level evaluation; it does not demonstrate separation by subject, device, site or acquisition batch.

3. Model

The model has 31 encoded inputs and 21 outputs. The initial architecture is shown below; the selected architecture follows the selection experiment. Native diagrams group categorical features by source field, so their 17 input nodes and one target node represent 31 encoded inputs and 21 outputs.

LayerInput shapeOutput shapeActivation
Scaling3131
Dense313Tanh
Dense321Sigmoid

Output values are uncalibrated model scores. Use the output encodings and decision rule documented with this project; do not assume independent sigmoid scores sum to one.

Primary tumor site classification — initial network architecture
Initial architecture from the saved report.

4. Training strategy

The saved training configuration uses QuasiNewton with CrossEntropy.

No separate completed training task is stored in this report. The neuron-selection experiment below contains the saved fitting results.

5. Model selection and baseline

The saved growing-neurons experiment selected eight hidden neurons. The initial network has three hidden neurons; the selected network is shown after the selection results, and precedes testing.

Growing neurons results

MeasureValue
Optimal neurons number8
Optimum training error1.924
Optimum selection error2.679
Epochs number10
Stopping criterionMaximum neurons reached
Elapsed time00:00:01
Growing neurons training/selection errors plot
Growing neurons training/selection errors plot. Native Neural Designer report for this project.
Selected layerInput shapeOutput shapeActivation
Scaling3131
Dense318Tanh
Dense821Sigmoid
Primary tumor site classification — selected network architecture
Network architecture. Native Neural Designer report for this project.

On this test subset a majority-class baseline would correctly classify 98.51% of the records. This baseline does not detect both classes and is not an optimized model.

6. Clinical validation

The figures and tables below refer to the current project’s saved testing analysis. The subset uses testing role 2; it contains 67 source records.

The native report provides one-versus-rest confusion counts at threshold 0.5 and a separate multiple-classification summary. Some test classes have zero support. A single clinical sensitivity, specificity, precision or ROC AUC is not asserted across the 21 sites; the per-site counts define the evaluation.

These operating-point metrics are calculated from the saved confusion counts at threshold 0.5. They apply to bladder; the remaining outputs have separate one-versus-rest tables.

Confusion table: bladder

MeasurePredicted positivePredicted negativeTotal
Actual positive0 (0.0%)1 (1.5%)1 (1.5%)
Actual negative0 (0.0%)66 (98.5%)66 (98.5%)
Total0 (0.0%)67 (100.0%)67 (100.0%)
Inspect every output confusion table

Confusion table: breast

MeasurePredicted positivePredicted negativeTotal
Actual positive3 (4.5%)1 (1.5%)4 (6.0%)
Actual negative2 (3.0%)61 (91.0%)63 (94.0%)
Total5 (7.5%)62 (92.5%)67 (100.0%)

Confusion table: cervix_uteri

MeasurePredicted positivePredicted negativeTotal
Actual positive0 (0.0%)0 (0.0%)0 (0.0%)
Actual negative0 (0.0%)67 (100.0%)67 (100.0%)
Total0 (0.0%)67 (100.0%)67 (100.0%)

Confusion table: colon

MeasurePredicted positivePredicted negativeTotal
Actual positive0 (0.0%)3 (4.5%)3 (4.5%)
Actual negative0 (0.0%)64 (95.5%)64 (95.5%)
Total0 (0.0%)67 (100.0%)67 (100.0%)

Confusion table: corpus_uteri

MeasurePredicted positivePredicted negativeTotal
Actual positive0 (0.0%)0 (0.0%)0 (0.0%)
Actual negative0 (0.0%)67 (100.0%)67 (100.0%)
Total0 (0.0%)67 (100.0%)67 (100.0%)

Confusion table: duodenum_and_small_intestine

MeasurePredicted positivePredicted negativeTotal
Actual positive0 (0.0%)0 (0.0%)0 (0.0%)
Actual negative0 (0.0%)67 (100.0%)67 (100.0%)
Total0 (0.0%)67 (100.0%)67 (100.0%)

Confusion table: esophagus

MeasurePredicted positivePredicted negativeTotal
Actual positive0 (0.0%)0 (0.0%)0 (0.0%)
Actual negative0 (0.0%)67 (100.0%)67 (100.0%)
Total0 (0.0%)67 (100.0%)67 (100.0%)

Confusion table: gallbladder

MeasurePredicted positivePredicted negativeTotal
Actual positive0 (0.0%)6 (9.0%)6 (9.0%)
Actual negative0 (0.0%)61 (91.0%)61 (91.0%)
Total0 (0.0%)67 (100.0%)67 (100.0%)

Confusion table: head_and_neck

MeasurePredicted positivePredicted negativeTotal
Actual positive3 (4.5%)0 (0.0%)3 (4.5%)
Actual negative0 (0.0%)64 (95.5%)64 (95.5%)
Total3 (4.5%)64 (95.5%)67 (100.0%)

Confusion table: kidney

MeasurePredicted positivePredicted negativeTotal
Actual positive1 (1.5%)3 (4.5%)4 (6.0%)
Actual negative0 (0.0%)63 (94.0%)63 (94.0%)
Total1 (1.5%)66 (98.5%)67 (100.0%)

Confusion table: liver

MeasurePredicted positivePredicted negativeTotal
Actual positive0 (0.0%)0 (0.0%)0 (0.0%)
Actual negative0 (0.0%)67 (100.0%)67 (100.0%)
Total0 (0.0%)67 (100.0%)67 (100.0%)

Confusion table: lung

MeasurePredicted positivePredicted negativeTotal
Actual positive13 (19.4%)9 (13.4%)22 (32.8%)
Actual negative5 (7.5%)40 (59.7%)45 (67.2%)
Total18 (26.9%)49 (73.1%)67 (100.0%)

Confusion table: ovary

MeasurePredicted positivePredicted negativeTotal
Actual positive3 (4.5%)2 (3.0%)5 (7.5%)
Actual negative0 (0.0%)62 (92.5%)62 (92.5%)
Total3 (4.5%)64 (95.5%)67 (100.0%)

Confusion table: pancreas

MeasurePredicted positivePredicted negativeTotal
Actual positive0 (0.0%)6 (9.0%)6 (9.0%)
Actual negative0 (0.0%)61 (91.0%)61 (91.0%)
Total0 (0.0%)67 (100.0%)67 (100.0%)

Confusion table: prostate

MeasurePredicted positivePredicted negativeTotal
Actual positive0 (0.0%)0 (0.0%)0 (0.0%)
Actual negative0 (0.0%)67 (100.0%)67 (100.0%)
Total0 (0.0%)67 (100.0%)67 (100.0%)

Confusion table: rectum

MeasurePredicted positivePredicted negativeTotal
Actual positive0 (0.0%)3 (4.5%)3 (4.5%)
Actual negative0 (0.0%)64 (95.5%)64 (95.5%)
Total0 (0.0%)67 (100.0%)67 (100.0%)

Confusion table: salivary_glands

MeasurePredicted positivePredicted negativeTotal
Actual positive0 (0.0%)1 (1.5%)1 (1.5%)
Actual negative0 (0.0%)66 (98.5%)66 (98.5%)
Total0 (0.0%)67 (100.0%)67 (100.0%)

Confusion table: stomach

MeasurePredicted positivePredicted negativeTotal
Actual positive0 (0.0%)7 (10.4%)7 (10.4%)
Actual negative1 (1.5%)59 (88.1%)60 (89.6%)
Total1 (1.5%)66 (98.5%)67 (100.0%)

Confusion table: testis

MeasurePredicted positivePredicted negativeTotal
Actual positive0 (0.0%)0 (0.0%)0 (0.0%)
Actual negative0 (0.0%)67 (100.0%)67 (100.0%)
Total0 (0.0%)67 (100.0%)67 (100.0%)

Confusion table: thyroid

MeasurePredicted positivePredicted negativeTotal
Actual positive0 (0.0%)2 (3.0%)2 (3.0%)
Actual negative0 (0.0%)65 (97.0%)65 (97.0%)
Total0 (0.0%)67 (100.0%)67 (100.0%)

Confusion table: vagina

MeasurePredicted positivePredicted negativeTotal
Actual positive0 (0.0%)0 (0.0%)0 (0.0%)
Actual negative0 (0.0%)67 (100.0%)67 (100.0%)
Total0 (0.0%)67 (100.0%)67 (100.0%)
Test measureValue
Testing records67
Positive cases1
Positive prevalence1.49%
Accuracy98.51%
Sensitivity / recall0.00%
Specificity100.00%
Precision / PPVUndefined (zero denominator)
F1 score0.000

Multiple classification tests

MeasurePrecisionRecallF1 score
bladder0.0000.0000.000
breast0.6000.7500.667
cervix_uteri0.0000.0000.000
colon0.0000.0000.000
corpus_uteri0.0000.0000.000
duodenum_and_small_intestine0.0000.0000.000
esophagus0.0000.0000.000
gallbladder0.6000.5000.545
head_and_neck1.0001.0001.000
kidney0.6670.5000.571
liver0.0000.0000.000
lung0.7390.7730.756
ovary0.3640.8000.500
pancreas0.3330.1670.222
prostate0.0000.0000.000
rectum0.0000.0000.000
salivary_glands0.0000.0000.000
stomach0.1670.2860.211
testis0.0000.0000.000
thyroid0.0000.0000.000
vagina0.0000.0000.000
Macro average0.2130.2270.213
Weighted average0.4910.5220.495
Biomedical research only. Specialist review and a confirmatory clinical reference standard remain necessary.

7. Workflow and reproducibility

Open the downloaded project in Neural Designer, inspect the dataset roles and preprocessing, then review the saved task report. Use the same input schema and category order when calculating outputs. The ZIP contains the exact current .nd, its source data and the applicable dataset notices.

Research workflow: eligible record or image → measurement and schema checks → model score → expert review → confirmatory clinical method where appropriate. No model output should be used as an autonomous diagnosis or treatment instruction.

8. Safety, generalizability and governance

The 67-row test subset leaves some tumor sites absent or represented by very few records. The saved network uses separate sigmoid outputs; its one-versus-rest tables must not be mistaken for a normalized probability distribution. No external validation or calibrated clinical probabilities are established.

Biomedical research only. Specialist review and a confirmatory clinical reference standard remain necessary.

References