Skip to content
Blog

Milk impurity classification from spectral measurements

Classify recorded pure and impure milk spectra

The updated example uses spectral measurements and a binary Pure/Impure label. The saved model uses nine numeric inputs and is evaluated on 4,655 held-out rows from a 23,279-row adapted file.

23,279Source records
9Encoded input values
4,655Testing-role records
1Model outputs

1. Industrial challenge

Food-science analysts can study how recorded spectral measurements distinguish the dataset labels and inspect misclassified observations. Laboratory confirmation and validated quality procedures remain separate from the tutorial model.

Spectral inputs

Use eight diode measurements and the saved test field.

Label direction

Interpret the positive output as Pure.

Independent validation

Evaluate by specimen and acquisition batch.

Food scienceSpectral analysisQuality data review
Recorded spectral-label benchmark. It does not certify milk safety, absence of pathogens or regulatory conformity.

2. Data set

The Mendeley version-1 dataset was collected for milk-impurity research with an eight-photodiode spectral sensor. The distributed CSV has 23,279 records; the repository description gives a rounded 20,000-sample overview. The model uses diode_1 through diode_8 plus the source field test, while integrationtime and referencecurrent are not inputs. The source description does not establish the physical meaning or availability timing of test.

Source: A Spectral Dataset of both Liquid and Solid Milk with Impurities for some Milk Impurity Detection. Dataset license: CC-BY-4.0. The downloadable ZIP includes the adapted data and attribution notices.

Dataset measureSaved value
Analysis unitspectral measurement record
Records23,279
Raw variables12
Encoded model inputs9
Model outputs1
Training roles13969
Validation / selection roles4655
Testing roles4655
Unused roles0
FieldRoleTypeCategories
integrationtimeNoneConstant
referencecurrentNoneConstant
diode_1InputNumeric
diode_2InputNumeric
diode_3InputNumeric
diode_4InputNumeric
diode_5InputNumeric
diode_6InputNumeric
diode_7InputNumeric
diode_8InputNumeric
testInputNumeric
labelTargetBinaryImpure, Pure
Target class distribution pie chart
Target class distribution pie chart. Native Neural Designer report for this project.
The saved project assigns rows to training, validation and testing as shown above. This is internal record-level evaluation; it does not demonstrate separation by subject, device, site or acquisition batch.

3. Model

The model has 9 encoded inputs and 1 outputs. No architecture-selection experiment is recorded in this project. The diagram shows the topology used by the saved model.

LayerInput shapeOutput shapeActivation
Scaling99
Dense93Tanh
Dense31Sigmoid

Output values are uncalibrated model scores. For binary evaluation, the saved positive class is Pure.

Milk impurity classification from spectral measurements — initial network architecture
Topology of the saved current model; no architecture selection is recorded.

4. Training strategy

The saved training configuration uses QuasiNewton with WeightedSquaredError.

Quasi-Newton method results

MeasureValue
Epochs number43
Elapsed time00:00:01
Stopping criterionMinimum loss decrease
Training error0.005
Validation error0.002
Quasi-Newton method error history
Quasi-Newton method error history. Native Neural Designer report for this project.

5. Model selection

No model selection experiment is recorded for this version. The validation subset guides fitting where a training report is present; it is distinct from the held-out test rows.

On this test subset a majority-class baseline would correctly classify 59.29% of the records. This baseline does not detect both classes and is not an optimized model.

6. Testing analysis

The figures and tables below refer to the current project’s saved testing analysis. The subset uses testing role 2; it contains 4655 source records.

ROC AUC describes ranking on this testing subset. Any optimal threshold shown in the saved ROC report was selected descriptively on that same subset; it is not an independently validated operating policy.

These operating-point metrics are calculated from the saved confusion counts at threshold 0.5. The positive-label coding is stated in the model section.

Confusion table

MeasurePredicted positivePredicted negativeTotal
Actual positive1890 (40.6%)5 (0.1%)1895 (40.7%)
Actual negative1 (0.0%)2759 (59.3%)2760 (59.3%)
Total1891 (40.6%)2764 (59.4%)4655 (100.0%)
Test measureValue
Testing records4655
Positive cases1895
Positive prevalence40.71%
Accuracy99.87%
Sensitivity / recall99.74%
Specificity99.96%
Precision / PPV99.95%
F1 score0.998

Area under curve

MeasureValue
Area under curve1
ROC chart
ROC chart. Native Neural Designer report for this project.

7. Model deployment

Open the downloaded project in Neural Designer, inspect the dataset roles and preprocessing, then review the saved task report. Use the same input schema and category order when calculating outputs. The ZIP contains the exact current .nd, its source data and the applicable dataset notices.

Workflow: source measurements → schema and availability checks → model output → domain review. Keep model versions, validation evidence and incoming-data monitoring together.

8. Scope and limitations

The positive model class is Pure, not Impure. A high Pure score is not laboratory confirmation. Audit the undocumented test input, repeated spectra and specimen/batch grouping before interpreting the near-perfect row-level result as generalization. Independent samples, instruments and laboratory labels are needed.

References