Skip to content
Learning

Amazon review sentiment classification

Turn product reviews into sentiment labels

Classify short English product reviews as bad or good with an attention-based text model. This example uses 1,000 labelled reviews and reproduces 67.5% accuracy on the 200 documents assigned to testing.

1,000Documents
2Classes
200Saved testing-role samples
67.50%Verified test accuracy

1. Business decision

Customer-experience teams need to organize review text before examining recurring product and service issues. A sentiment classifier can prioritize manual review and summarize feedback, provided that uncertain or ambiguous messages remain open to human interpretation.

Review triage

Group favourable and unfavourable feedback for a reviewer.

Error inspection

Inspect false positives and false negatives before defining a workflow.

Reproducible baseline

Start with the supplied small model and retain the same text preprocessing.

Machine-learning practitionersData and analytics teams
Sentence-level English sentiment classification. The model does not identify topics, explain dissatisfaction or estimate the effect of a marketing intervention.

2. Data set

The local file is the Amazon subset of the UCI Sentiment Labelled Sentences collection, with the binary labels expressed as Bad and Good. These are review sentences, not synthetic examples. The collection is balanced by design; it does not estimate the natural proportion of satisfied customers.

Dataset source and access conditions. Source reviewed on 24 September 2026.

Target class distribution table
DocumentsShare (%)
Bad50050.00
Good50050.00

Target class distribution pie chart. Enable JavaScript to explore the chart.

Class distribution in the complete supplied dataset.
SubsetSamples
Training600
Validation200
Testing200

Inputs are UTF-8 text, one document per row, followed by its label in the tab-separated source. The saved word-level tokenizer provides padding, unknown, start and end tokens. Its vocabulary and sequence limit must travel with the trained model.

Vocabulary
Value
Vocabulary size828
Words824
Reserved tokens4
Words appearing once0
Words never used0
Sequence length
Tokens
Median13
90th percentile26
99th percentile34
Longest42
Limit42
Padding (%)65.73
Unknown words
SamplesTokensUnknownUnknown (%)
Training60085336127.17
Validation20030402468.09
Testing20028202037.2

Vocabulary coverage curve. Enable JavaScript to explore the chart.

Native vocabulary coverage curve for the saved text dataset.
The project stores a 60/20/20 sample partition (rounded for the 1,999-image collection). Results apply only to this example’s split. Group-level independence is not documented.

3. Model

The following configuration comes from the saved trained model. The same topology is used before and after training; no separate architecture-selection run is recorded.

LayerInput shapeOutput shapeConfiguration
Embedding4242 × 16Vocabulary 828; positional encoding
MultiHeadAttention42 × 1642 × 162 attention heads; non-causal
Pooling3d42 × 1616AveragePooling
Dense1616ReLU
Dense161Sigmoid

Sigmoid or softmax outputs are uncalibrated model scores. For this binary model, the positive label is good. The 0.5 threshold assigns the positive class.

Amazon review sentiment classification — saved neural network architecture
Static architecture diagram rendered by Neural Designer from the saved report. The topology is the same before and after training. Open the image to inspect the layer dimensions.

4. Training strategy

Training minimizes cross-entropy using Adam. The saved configuration uses a learning rate of 0.001, mini-batches of 64 and L2 regularization with weight 0.001.

Adaptive moment estimation results
Value
Epochs number40
Elapsed time00:00:02
Stopping criterionMaximum epochs number
Training error0.409
Validation error0.462

Adaptive moment estimation error history. Enable JavaScript to explore the chart.

Native training and validation error history from the supplied report.

The summary table and epoch history are retained as separate native report outputs. Epoch-history endpoints need not equal the restored best-validation model; they are not additional test metrics.

5. Model selection

No architecture selection or feature selection experiment is recorded in this report. The model-selection settings in the project are configuration options, not evidence that a search was run. The validation subset monitors training; the testing subset is intended for subsequent evaluation.

A constant classifier choosing the most frequent training label (bad) achieves 49.00% accuracy on these same testing rows, compared with 67.50% for the supplied model. The baseline label is chosen only from the training rows; this is a simple comparator, not a separately tuned model.

6. Testing analysis

The exported model was rerun on the 200 documents whose saved sample role is testing. Predicted labels reproduce the saved confusion counts. Rows are actual classes and columns are predicted classes.

MetricValue
Testing documents200
Accuracy67.50%
Macro F10.675
Actual / predictedbadgood
bad6830
good3567
ClassSupportPrecisionRecallF1
bad980.6600.6940.677
good1020.6910.6570.673

At threshold 0.5, good is positive: TP = 67, FN = 35, FP = 30 and TN = 68. Sensitivity is 65.69%, specificity 69.39%, precision 69.07% and NPV 66.02%. Positive-label prevalence is 102/200 = 51%; predictive values refer to this benchmark composition.

Area under curve
Value
Area under curve0.785

The saved ROC AUC is 0.785. Its test-derived threshold of 0.633 is exploratory and is not used for the 0.5-threshold confusion matrix.

ROC chart. Enable JavaScript to explore the chart.

Native ROC curve from the saved report. The highlighted optimal threshold is exploratory, chosen on that evaluation subset.

7. Model deployment

Understanding customer feedback

See how the model distinguishes favourable and unfavourable customer reviews. Each example shows the original review and the score assigned to each sentiment.

Example 1

Input text

The only thing that I think could improve is the sound leaks out from the headset.

Predicted class: bad
Model score: 70.61%

ClassOutput score
bad0.706127
good0.293873

Example 2

Input text

Excellent sound, battery life and inconspicuous to boot!.

Predicted class: good
Model score: 62.78%

ClassOutput score
bad0.372200
good0.627800

These are illustrative predictions, separate from the testing metrics above. Output scores have not been calibrated.

Download the original Neural Designer project and native figures, then relink the source dataset. The ZIP is a project archive; it does not contain the standalone Neural Engine runtime.

For batch inference, export a deployment package from Neural Designer. Keep the generated Python wrapper beside its engine/ and model/ folders. The supplied export uses Windows x64 and Python 3, with CPU inference by default.

python amazon_reviews.py "A short text to classify"

Preserve the trained vocabulary and sequence limit for text, or the saved image dimensions, channel handling and scaling for images. Record the model version and input quality before sending outputs for human review. The wrapper’s score output is not evidence of probability calibration.

8. Evidence and limitations

Short reviews, unfamiliar vocabulary, negation, sarcasm and mixed sentiment can cause errors. The balanced benchmark does not represent live customer sentiment prevalence. Evaluate on a later, product-specific sample before routing reviews automatically.

The supplied artifacts do not establish independence by person, product, author, acquisition session or site. Check duplicates and group-related records before reporting performance on a new population. Calibration and external validation are not established; monitor drift and review errors before operational decisions.

Vocabulary statistics cover the saved dataset. A training-only vocabulary fit is not documented, so these results do not establish a fully isolated preprocessing evaluation. Refit preprocessing inside the training partition when comparing models on a strict benchmark.

References