Classify recorded pure and impure milk spectra
The updated example uses spectral measurements and a binary Pure/Impure label. The saved model uses nine numeric inputs and is evaluated on 4,655 held-out rows from a 23,279-row adapted file.
1. Industrial challenge
Food-science analysts can study how recorded spectral measurements distinguish the dataset labels and inspect misclassified observations. Laboratory confirmation and validated quality procedures remain separate from the tutorial model.
Spectral inputs
Use eight diode measurements and the saved test field.
Label direction
Interpret the positive output as Pure.
Independent validation
Evaluate by specimen and acquisition batch.
2. Data set
The Mendeley version-1 dataset was collected for milk-impurity research with an eight-photodiode spectral sensor. The distributed CSV has 23,279 records; the repository description gives a rounded 20,000-sample overview. The model uses diode_1 through diode_8 plus the source field test, while integrationtime and referencecurrent are not inputs. The source description does not establish the physical meaning or availability timing of test.
Source: A Spectral Dataset of both Liquid and Solid Milk with Impurities for some Milk Impurity Detection. Dataset license: CC-BY-4.0. The downloadable ZIP includes the adapted data and attribution notices.
| Dataset measure | Saved value |
|---|---|
| Analysis unit | spectral measurement record |
| Records | 23,279 |
| Raw variables | 12 |
| Encoded model inputs | 9 |
| Model outputs | 1 |
| Training roles | 13969 |
| Validation / selection roles | 4655 |
| Testing roles | 4655 |
| Unused roles | 0 |
| Field | Role | Type | Categories |
|---|---|---|---|
| integrationtime | None | Constant | |
| referencecurrent | None | Constant | |
| diode_1 | Input | Numeric | |
| diode_2 | Input | Numeric | |
| diode_3 | Input | Numeric | |
| diode_4 | Input | Numeric | |
| diode_5 | Input | Numeric | |
| diode_6 | Input | Numeric | |
| diode_7 | Input | Numeric | |
| diode_8 | Input | Numeric | |
| test | Input | Numeric | |
| label | Target | Binary | Impure, Pure |

3. Model
The model has 9 encoded inputs and 1 outputs. No architecture-selection experiment is recorded in this project. The diagram shows the topology used by the saved model.
| Layer | Input shape | Output shape | Activation |
|---|---|---|---|
| Scaling | 9 | 9 | |
| Dense | 9 | 3 | Tanh |
| Dense | 3 | 1 | Sigmoid |
Output values are uncalibrated model scores. For binary evaluation, the saved positive class is Pure.

4. Training strategy
The saved training configuration uses QuasiNewton with WeightedSquaredError.
Quasi-Newton method results
| Measure | Value |
|---|---|
| Epochs number | 43 |
| Elapsed time | 00:00:01 |
| Stopping criterion | Minimum loss decrease |
| Training error | 0.005 |
| Validation error | 0.002 |

5. Model selection
No model selection experiment is recorded for this version. The validation subset guides fitting where a training report is present; it is distinct from the held-out test rows.
On this test subset a majority-class baseline would correctly classify 59.29% of the records. This baseline does not detect both classes and is not an optimized model.
6. Testing analysis
The figures and tables below refer to the current project’s saved testing analysis. The subset uses testing role 2; it contains 4655 source records.
ROC AUC describes ranking on this testing subset. Any optimal threshold shown in the saved ROC report was selected descriptively on that same subset; it is not an independently validated operating policy.
These operating-point metrics are calculated from the saved confusion counts at threshold 0.5. The positive-label coding is stated in the model section.
Confusion table
| Measure | Predicted positive | Predicted negative | Total |
|---|---|---|---|
| Actual positive | 1890 (40.6%) | 5 (0.1%) | 1895 (40.7%) |
| Actual negative | 1 (0.0%) | 2759 (59.3%) | 2760 (59.3%) |
| Total | 1891 (40.6%) | 2764 (59.4%) | 4655 (100.0%) |
| Test measure | Value |
|---|---|
| Testing records | 4655 |
| Positive cases | 1895 |
| Positive prevalence | 40.71% |
| Accuracy | 99.87% |
| Sensitivity / recall | 99.74% |
| Specificity | 99.96% |
| Precision / PPV | 99.95% |
| F1 score | 0.998 |
Area under curve
| Measure | Value |
|---|---|
| Area under curve | 1 |

7. Model deployment
Open the downloaded project in Neural Designer, inspect the dataset roles and preprocessing, then review the saved task report. Use the same input schema and category order when calculating outputs. The ZIP contains the exact current .nd, its source data and the applicable dataset notices.
Workflow: source measurements → schema and availability checks → model output → domain review. Keep model versions, validation evidence and incoming-data monitoring together.
8. Scope and limitations
The positive model class is Pure, not Impure. A high Pure score is not laboratory confirmation. Audit the undocumented test input, repeated spectra and specimen/batch grouping before interpreting the near-perfect row-level result as generalization. Independent samples, instruments and laboratory labels are needed.
References
- A Spectral Dataset of both Liquid and Solid Milk with Impurities for some Milk Impurity Detection. Pranali Modi; Janvi Bhanushali; Maahi Shah; Dhruvi Dhulia; Madhuri Barochiya. 10.17632/3gxjgxkg76.1
- Dataset terms: Creative Commons Attribution 4.0 International. Full attribution and transformations are included in
LICENSES/DATASET-LICENSE.txt. - Current Neural Designer project and saved task report, snapshot 6 October 2026.



