Skip to content
Learning

MNIST handwritten digit classification

Classify handwritten digits with a convolutional network

A convolutional classifier recognizes ten handwritten digit classes in 10,000 grayscale images. The saved model achieves 96.10% accuracy on its 2,000 testing images, with all confusion counts reproduced by exported inference.

10,000Images
10Classes
2,000Saved testing-role samples
96.10%Verified test accuracy

1. Industrial challenge

Digit recognition is a compact introduction to visual pattern classification and document-processing systems. This example maps an already cropped digit image to one of ten classes; it does not locate digits in a page or read multi-digit fields.

Image preprocessing

Preserve 28 × 28 grayscale inputs and the saved scaling layer.

Pattern classification

Map an image to one of the ten digit classes.

Integration planning

Add crop detection and rejection rules before a document workflow.

Machine-learning practitionersData and analytics teams
One cropped, grayscale digit per image. The model output order is eight, five, four, nine, one, seven, six, three, two, zero.

2. Data set

The supplied collection contains 10,000 MNIST digit images at 28 × 28 pixels, with class counts matching the standard MNIST test collection. This project repartitions those images into training, validation and testing; it is not the standard 60,000-training / 10,000-test MNIST benchmark protocol.

Dataset source and access conditions. Source reviewed on 24 September 2026.

Target class distribution table
ImagesShare (%)
eight9749.74
five8928.92
four9829.82
nine100910.09
one113511.35
seven102810.28
six9589.58
three101010.10
two103210.32
zero9809.80

Target class distribution pie chart. Enable JavaScript to explore the chart.

Class distribution in the complete supplied dataset.
SubsetSamples
Training6000
Validation2000
Testing2000
Images dimension table
WidthHeightChannelsImages
Dimensions2828110000

The saved augmentation settings have zero rotation and translation, with no random reflections. Images are read in class-folder order and sorted filename order within each class.

The project stores a 60/20/20 sample partition (rounded for the 1,999-image collection). Results apply only to this example’s split. Group-level independence is not documented.

3. Model

The following configuration comes from the saved trained model. The same topology is used before and after training; no separate architecture-selection run is recorded.

LayerInput shapeOutput shapeConfiguration
Scaling28 × 28 × 128 × 28 × 1ImageMinMax
Convolutional28 × 28 × 128 × 28 × 8ReLU; 8 filters, 3 × 3
Pooling28 × 28 × 814 × 14 × 8MaxPooling
Flatten14 × 14 × 81568
Dense1568128ReLU
Dense12810Softmax

Sigmoid or softmax outputs are uncalibrated model scores. The predicted class is the label with the largest output score.

MNIST handwritten digit classification — saved neural network architecture
Static architecture diagram rendered by Neural Designer from the saved report. The topology is the same before and after training. Open the image to inspect the layer dimensions.

4. Training strategy

Training minimizes cross-entropy using Adam. The saved configuration uses a learning rate of 0.001, mini-batches of 32 and no regularization.

Adaptive moment estimation results
Value
Epochs number20
Elapsed time00:00:04
Stopping criterionMaximum epochs number
Training error0.022
Validation error0.107

Adaptive moment estimation error history. Enable JavaScript to explore the chart.

Native training and validation error history from the supplied report.

The summary table and epoch history are retained as separate native report outputs. Epoch-history endpoints need not equal the restored best-validation model; they are not additional test metrics.

5. Model selection

No architecture selection or feature selection experiment is recorded. The same network topology is used before and after training. The validation subset monitors training; the reported evaluation uses the saved testing assignment.

A constant classifier predicting the most frequent training label (one) achieves 10.65% accuracy on the same testing images, compared with 96.10% for this model. The constant label is chosen from the training images only.

6. Testing analysis

Testing uses the 2,000 images assigned to the testing subset in the saved project. The confusion matrix and class metrics agree with exported inference, and repeated reloads preserve the same evaluation images. Rows below are actual labels; columns are predicted labels.

Actual / predictedeightfivefournineonesevensixthreetwozero
eight194110100001
five116301002502
four101863000010
nine125194020601
one000021010110
seven011122070110
six220000179002
three210004020610
two300013031872
zero500000000196
ClassTesting imagesPrecisionRecallF1
eight1980.9280.9800.953
five1740.9590.9370.948
four1910.9640.9740.969
nine2110.9750.9190.946
one2130.9810.9860.984
seven2140.9540.9670.961
six1850.9890.9680.978
three2140.9280.9630.945
two1990.9790.9400.959
zero2010.9610.9750.968
Summary metricValue
Accuracy96.10%
Macro F10.961
Testing images2000

The model predicts a single digit from a cropped image. The largest class-level recall shortfall is for nine (0.919). These results use this project’s own partition of the 10,000-image collection, not the standard MNIST training/test protocol.

7. Model deployment

Recognizing handwritten digits

See how the model recognizes two handwritten digits. Each example pairs the original image with the predicted digit and the scores for all ten classes.

Example 1

Handwritten digit selected for Calculate outputs: one/1_149.bmp
Input image: one/1_149.bmp

Predicted class: one
Model score: 99.97%

ClassOutput score
eight0.000007
five0.000000
four0.000011
nine0.000001
one0.999687
seven0.000289
six0.000000
three0.000002
two0.000004
zero0.000000

Example 2

Handwritten digit selected for Calculate outputs: four/4_375.bmp
Input image: four/4_375.bmp

Predicted class: four
Model score: 93.20%

ClassOutput score
eight0.064999
five0.000000
four0.932005
nine0.002492
one0.000000
seven0.000375
six0.000088
three0.000000
two0.000037
zero0.000003

These are illustrative predictions, separate from the testing metrics above. Output scores have not been calibrated.

Download the original Neural Designer project and native figures, then relink the source dataset. The ZIP is a project archive; it does not contain the standalone Neural Engine runtime.

For batch inference, export a deployment package from Neural Designer. Keep the generated Python wrapper beside its engine/ and model/ folders. The supplied export uses Windows x64 and Python 3, with CPU inference by default.

python mnist.py sample.bmp

Preserve the trained vocabulary and sequence limit for text, or the saved image dimensions, channel handling and scaling for images. Record the model version and input quality before sending outputs for human review. The wrapper’s score output is not evidence of probability calibration.

8. Scope and limitations

Handwriting, framing, background polarity and image quality can shift outside this collection. Digit localization, rejection of non-digits and full-page optical character recognition require additional components. Writer-level separation is not documented.

The supplied artifacts do not establish independence by person, product, author, acquisition session or site. Check duplicates and group-related records before reporting performance on a new population. Calibration and external validation are not established; monitor drift and review errors before operational decisions.

References