Classify handwritten digits with a convolutional network
A convolutional classifier recognizes ten handwritten digit classes in 10,000 grayscale images. The saved model achieves 96.10% accuracy on its 2,000 testing images, with all confusion counts reproduced by exported inference.
1. Industrial challenge
Digit recognition is a compact introduction to visual pattern classification and document-processing systems. This example maps an already cropped digit image to one of ten classes; it does not locate digits in a page or read multi-digit fields.
Image preprocessing
Preserve 28 × 28 grayscale inputs and the saved scaling layer.
Pattern classification
Map an image to one of the ten digit classes.
Integration planning
Add crop detection and rejection rules before a document workflow.
2. Data set
The supplied collection contains 10,000 MNIST digit images at 28 × 28 pixels, with class counts matching the standard MNIST test collection. This project repartitions those images into training, validation and testing; it is not the standard 60,000-training / 10,000-test MNIST benchmark protocol.
Dataset source and access conditions. Source reviewed on 24 September 2026.
| Images | Share (%) | |
|---|---|---|
| eight | 974 | 9.74 |
| five | 892 | 8.92 |
| four | 982 | 9.82 |
| nine | 1009 | 10.09 |
| one | 1135 | 11.35 |
| seven | 1028 | 10.28 |
| six | 958 | 9.58 |
| three | 1010 | 10.10 |
| two | 1032 | 10.32 |
| zero | 980 | 9.80 |
Target class distribution pie chart. Enable JavaScript to explore the chart.
| Subset | Samples |
|---|---|
| Training | 6000 |
| Validation | 2000 |
| Testing | 2000 |
| Width | Height | Channels | Images | |
|---|---|---|---|---|
| Dimensions | 28 | 28 | 1 | 10000 |
The saved augmentation settings have zero rotation and translation, with no random reflections. Images are read in class-folder order and sorted filename order within each class.
3. Model
The following configuration comes from the saved trained model. The same topology is used before and after training; no separate architecture-selection run is recorded.
| Layer | Input shape | Output shape | Configuration |
|---|---|---|---|
| Scaling | 28 × 28 × 1 | 28 × 28 × 1 | ImageMinMax |
| Convolutional | 28 × 28 × 1 | 28 × 28 × 8 | ReLU; 8 filters, 3 × 3 |
| Pooling | 28 × 28 × 8 | 14 × 14 × 8 | MaxPooling |
| Flatten | 14 × 14 × 8 | 1568 | |
| Dense | 1568 | 128 | ReLU |
| Dense | 128 | 10 | Softmax |
Sigmoid or softmax outputs are uncalibrated model scores. The predicted class is the label with the largest output score.

4. Training strategy
Training minimizes cross-entropy using Adam. The saved configuration uses a learning rate of 0.001, mini-batches of 32 and no regularization.
| Value | |
|---|---|
| Epochs number | 20 |
| Elapsed time | 00:00:04 |
| Stopping criterion | Maximum epochs number |
| Training error | 0.022 |
| Validation error | 0.107 |
Adaptive moment estimation error history. Enable JavaScript to explore the chart.
The summary table and epoch history are retained as separate native report outputs. Epoch-history endpoints need not equal the restored best-validation model; they are not additional test metrics.
5. Model selection
No architecture selection or feature selection experiment is recorded. The same network topology is used before and after training. The validation subset monitors training; the reported evaluation uses the saved testing assignment.
A constant classifier predicting the most frequent training label (one) achieves 10.65% accuracy on the same testing images, compared with 96.10% for this model. The constant label is chosen from the training images only.
6. Testing analysis
Testing uses the 2,000 images assigned to the testing subset in the saved project. The confusion matrix and class metrics agree with exported inference, and repeated reloads preserve the same evaluation images. Rows below are actual labels; columns are predicted labels.
| Actual / predicted | eight | five | four | nine | one | seven | six | three | two | zero |
|---|---|---|---|---|---|---|---|---|---|---|
| eight | 194 | 1 | 1 | 0 | 1 | 0 | 0 | 0 | 0 | 1 |
| five | 1 | 163 | 0 | 1 | 0 | 0 | 2 | 5 | 0 | 2 |
| four | 1 | 0 | 186 | 3 | 0 | 0 | 0 | 0 | 1 | 0 |
| nine | 1 | 2 | 5 | 194 | 0 | 2 | 0 | 6 | 0 | 1 |
| one | 0 | 0 | 0 | 0 | 210 | 1 | 0 | 1 | 1 | 0 |
| seven | 0 | 1 | 1 | 1 | 2 | 207 | 0 | 1 | 1 | 0 |
| six | 2 | 2 | 0 | 0 | 0 | 0 | 179 | 0 | 0 | 2 |
| three | 2 | 1 | 0 | 0 | 0 | 4 | 0 | 206 | 1 | 0 |
| two | 3 | 0 | 0 | 0 | 1 | 3 | 0 | 3 | 187 | 2 |
| zero | 5 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 196 |
| Class | Testing images | Precision | Recall | F1 |
|---|---|---|---|---|
| eight | 198 | 0.928 | 0.980 | 0.953 |
| five | 174 | 0.959 | 0.937 | 0.948 |
| four | 191 | 0.964 | 0.974 | 0.969 |
| nine | 211 | 0.975 | 0.919 | 0.946 |
| one | 213 | 0.981 | 0.986 | 0.984 |
| seven | 214 | 0.954 | 0.967 | 0.961 |
| six | 185 | 0.989 | 0.968 | 0.978 |
| three | 214 | 0.928 | 0.963 | 0.945 |
| two | 199 | 0.979 | 0.940 | 0.959 |
| zero | 201 | 0.961 | 0.975 | 0.968 |
| Summary metric | Value |
|---|---|
| Accuracy | 96.10% |
| Macro F1 | 0.961 |
| Testing images | 2000 |
The model predicts a single digit from a cropped image. The largest class-level recall shortfall is for nine (0.919). These results use this project’s own partition of the 10,000-image collection, not the standard MNIST training/test protocol.
7. Model deployment
Recognizing handwritten digits
See how the model recognizes two handwritten digits. Each example pairs the original image with the predicted digit and the scores for all ten classes.
Example 1

one/1_149.bmpPredicted class: one
Model score: 99.97%
| Class | Output score |
|---|---|
| eight | 0.000007 |
| five | 0.000000 |
| four | 0.000011 |
| nine | 0.000001 |
| one | 0.999687 |
| seven | 0.000289 |
| six | 0.000000 |
| three | 0.000002 |
| two | 0.000004 |
| zero | 0.000000 |
Example 2

four/4_375.bmpPredicted class: four
Model score: 93.20%
| Class | Output score |
|---|---|
| eight | 0.064999 |
| five | 0.000000 |
| four | 0.932005 |
| nine | 0.002492 |
| one | 0.000000 |
| seven | 0.000375 |
| six | 0.000088 |
| three | 0.000000 |
| two | 0.000037 |
| zero | 0.000003 |
These are illustrative predictions, separate from the testing metrics above. Output scores have not been calibrated.
Download the original Neural Designer project and native figures, then relink the source dataset. The ZIP is a project archive; it does not contain the standalone Neural Engine runtime.
For batch inference, export a deployment package from Neural Designer. Keep the generated Python wrapper beside its engine/ and model/ folders. The supplied export uses Windows x64 and Python 3, with CPU inference by default.
python mnist.py sample.bmpPreserve the trained vocabulary and sequence limit for text, or the saved image dimensions, channel handling and scaling for images. Record the model version and input quality before sending outputs for human review. The wrapper’s score output is not evidence of probability calibration.
8. Scope and limitations
Handwriting, framing, background polarity and image quality can shift outside this collection. Digit localization, rejection of non-digits and full-page optical character recognition require additional components. Writer-level separation is not documented.
The supplied artifacts do not establish independence by person, product, author, acquisition session or site. Check duplicates and group-related records before reporting performance on a new population. Calibration and external validation are not established; monitor drift and review errors before operational decisions.