Classifying recorded obesity levels
This research example classifies seven recorded weight-status categories from anthropometric and lifestyle variables. It evaluates multiclass errors rather than treating class codes as a continuous regression target.
1. Clinical question and intended use
The research question is whether recorded lifestyle and physical characteristics distinguish seven obesity levels. Since height and weight are inputs, this is a classification of current recorded status.
Seven categories
Inspect the complete class-specific confusion matrix.
Recorded measurements
Height and weight help define the classification context.
Research scope
Synthetic records and internal testing limit generalization claims.
2. Cohort, measurements and endpoint
The UCI dataset contains 2,111 records associated with Colombia, Peru and Mexico. The repository states that 77% were generated synthetically and 23% were collected through a web platform. Height and weight are inputs, so this is classification of recorded status rather than prediction of a future health outcome.
The downloadable project, saved report and supplied source CSV define the exact version used here. Repository: original dataset/source record.
| Subset | Records |
|---|---|
| Training | 1267 |
| Validation / selection | 422 |
| Testing | 422 |
| Unused | 0 |
| Variable | Role | Type | Encoding | Unit |
|---|---|---|---|---|
| gender | Input | Binary | Female; Male | As supplied |
| age | Input | Numeric | years | |
| height | Input | Numeric | m | |
| weight | Input | Numeric | kg | |
| family_history_with_overweight | Input | Binary | no; yes | As supplied |
| caloric_food | Input | Binary | no; yes | As supplied |
| vegetables | Input | Numeric | As supplied | |
| number_meals | Input | Numeric | As supplied | |
| food_between_meals | Input | Numeric | As supplied | |
| smoke | Input | Binary | no; yes | As supplied |
| water | Input | Numeric | As supplied | |
| calories | Input | Binary | no; yes | As supplied |
| activity | Input | Numeric | As supplied | |
| technology | Input | Numeric | As supplied | |
| alcohol | Input | Numeric | As supplied | |
| transportation | Input | Categorical | automobile; bike; motorbike; public_transportation; walking | As supplied |
| obesity_level | Target | Categorical | Normal_weight; Obese_I; Obese_II; Obese_III; Overweight_I; Overweight_II; Underweight | As supplied |
Interactive chart: obesity_level distribution pie chart. Enable JavaScript to explore it.
Interactive chart: obesity_level Pearson correlations chart. Enable JavaScript to explore it.
3. Model
The final model has 20 encoded input features and 7 outputs. The following dimensions describe the final saved network.
| Layer | Input shape | Output shape | Activation |
|---|---|---|---|
| Scaling | 20 | 20 | — |
| Dense | 20 | 3 | Tanh |
| Dense | 3 | 7 | Softmax |
Output semantics: one score per class in this order: Normal_weight, Obese_I, Obese_II, Obese_III, Overweight_I, Overweight_II, Underweight. The predicted label is the largest score. Calibration has not been evaluated, so scores are not presented as calibrated probabilities.

4. Training strategy
The saved training configuration uses CrossEntropy with QuasiNewton. Training minimizes the recorded objective; the validation subset monitors generalization during fitting. The testing subset is used for the evaluation below.
Interactive chart: Quasi-Newton method error history. Enable JavaScript to explore it.
Quasi-Newton method results
| Measure | Value |
|---|---|
| Epochs number | 185 |
| Elapsed time | 00:00:00 |
| Stopping criterion | Maximum validation error increases |
| Training error | 0.336 |
| Validation error | 0.343 |
5. Model selection and baseline
No model selection experiment is recorded in this supplied project. The displayed architecture is the trained model used for testing; earlier article claims about a different selected architecture do not apply to this version.
A transparent test-set comparator is the majority-class rule, with accuracy 17.8%. This is a baseline for interpretation, not an alternative model fitted on the test labels.
6. Clinical validation
The final classifier is evaluated on 422 testing records. The confusion counts below were reproduced from the saved model. Rows are actual classes and columns are predicted classes. The multiclass decision is argmax; no binary threshold is applied.
Test class prevalence is shown by the support counts. Accuracy is 87.4% and macro F1 is 0.856. A multiclass ROC AUC was not reported; the confusion matrix and per-class measures are the available evidence.
| Actual / predicted | Normal_weight | Obese_I | Obese_II | Obese_III | Overweight_I | Overweight_II | Underweight | Total |
|---|---|---|---|---|---|---|---|---|
| Normal_weight | 28 | 0 | 0 | 0 | 3 | 0 | 21 | 52 |
| Obese_I | 0 | 71 | 3 | 0 | 0 | 1 | 0 | 75 |
| Obese_II | 0 | 1 | 71 | 0 | 0 | 0 | 0 | 72 |
| Obese_III | 0 | 0 | 0 | 64 | 0 | 0 | 0 | 64 |
| Overweight_I | 7 | 1 | 0 | 0 | 41 | 7 | 0 | 56 |
| Overweight_II | 0 | 1 | 0 | 0 | 7 | 45 | 0 | 53 |
| Underweight | 1 | 0 | 0 | 0 | 0 | 0 | 49 | 50 |
| Class | Test cases | Sensitivity / recall | Specificity | Precision / PPV | F1 |
|---|---|---|---|---|---|
| Normal_weight | 52 | 53.8% | 97.8% | 77.8% | 0.636 |
| Obese_I | 75 | 94.7% | 99.1% | 95.9% | 0.953 |
| Obese_II | 72 | 98.6% | 99.1% | 95.9% | 0.973 |
| Obese_III | 64 | 100.0% | 100.0% | 100.0% | 1 |
| Overweight_I | 56 | 73.2% | 97.3% | 80.4% | 0.766 |
| Overweight_II | 53 | 84.9% | 97.8% | 84.9% | 0.849 |
| Underweight | 50 | 98.0% | 94.4% | 70.0% | 0.817 |
7. Workflow and reproducibility
Validated inputs → saved preprocessing → neural network → score or estimate → domain review. The ZIP contains the original project, source CSV, schema, test metrics and standalone interactive chart exports. The project hash in the schema identifies this exact version.
Explore the exported model
This research demonstration runs locally in your browser. Values outside the training range are outside the validated domain and are rejected. A valid input range does not guarantee that a combination is physically or operationally plausible.
This is not a diagnosis and must not guide medical treatment or donor eligibility.
8. Safety, generalizability and governance
Synthetic records and related anthropometric predictors limit interpretation of the internal split. Compare with a transparent BMI-based classification, separate original from synthetic subjects, and obtain external validation. Scores are not calibrated probabilities and must not guide individual diagnosis or treatment; clinician or specialist review is required for medical use.
No external validation or independent calibration study is included. Preprocessing statistics and model choices should be refitted within a prospective or grouped validation design. Correlations and directional responses describe associations, not causes. Human review is required before an operational decision.
Confidence intervals, subgroup performance, calibration curves and decision-cost validation are not established by these tasks. Predictive values apply to the observed test class distribution and may change when prevalence shifts.



