Skip to content
Learning

Pulsar detection with machine learning

Identify pulsar candidates from signal summaries

Classify HTRU2 candidates from eight integrated pulse-profile and dispersion-measure statistics. The supplied network is a signal-screening benchmark evaluated on 3,579 held-out records.

17,898Source records
8Encoded input values
3,579Testing-role records
1Model outputs

1. Scientific objective

Radio-astronomy screening produces many candidates that need prioritization. The model learns the HTRU2 annotation from summary statistics, supporting candidate ranking and error inspection.

Signal screening

Rank candidate patterns for scientific review.

Rare-class errors

Inspect missed candidates and false positives.

Reproducibility

Retain the saved split and signal-feature definitions.

Radio astronomySignal analysisScientific computing
Candidate classification within the HTRU2 benchmark; a model score does not confirm an astronomical discovery.

2. Data and provenance

The UCI HTRU2 dataset contains 17,898 records with eight numeric features and one binary class. The saved positive class is 1, representing a pulsar candidate.

Source: HTRU2. Dataset license: CC-BY-4.0. The downloadable ZIP includes the adapted data and attribution notices.

Dataset measureSaved value
Analysis unitcandidate record
Records17,898
Raw variables9
Encoded model inputs8
Model outputs1
Training roles10740
Validation / selection roles3579
Testing roles3579
Unused roles0
FieldRoleTypeCategories
profile_meanInputNumeric
profile_stdevInputNumeric
profile_skewnessInputNumeric
profile_kurtosisInputNumeric
dm_meanInputNumeric
dm_stdevInputNumeric
dm_skewnessInputNumeric
dm_kurtosisInputNumeric
classTargetBinary0, 1
Target class distribution pie chart
Target class distribution pie chart. Native Neural Designer report for this project.
The saved project assigns rows to training, validation and testing as shown above. This is internal record-level evaluation; it does not demonstrate separation by subject, device, site or acquisition batch.

3. Model

The model has 8 encoded inputs and 1 outputs. No architecture-selection experiment is recorded in this project. The diagram shows the topology used by the saved model.

LayerInput shapeOutput shapeActivation
Scaling88
Dense83Tanh
Dense31Sigmoid

Output values are uncalibrated model scores. For binary evaluation, the saved positive class is 1.

Pulsar detection with machine learning — initial network architecture
Topology of the saved current model; no architecture selection is recorded.

4. Training strategy

The saved training configuration uses QuasiNewton with WeightedSquaredError.

Quasi-Newton method results

MeasureValue
Epochs number117
Elapsed time00:00:00
Stopping criterionMinimum loss decrease
Training error0.106
Validation error0.103
Quasi-Newton method error history
Quasi-Newton method error history. Native Neural Designer report for this project.

5. Model selection and baseline

No model selection experiment is recorded for this version. The validation subset guides fitting where a training report is present; it is distinct from the held-out test rows.

On this test subset a majority-class baseline would correctly classify 90.98% of the records. This baseline does not detect both classes and is not an optimized model.

6. Scientific validation

The figures and tables below refer to the current project’s saved testing analysis. The subset uses testing role 2; it contains 3579 source records.

ROC AUC describes ranking on this testing subset. Any optimal threshold shown in the saved ROC report was selected descriptively on that same subset; it is not an independently validated operating policy.

These operating-point metrics are calculated from the saved confusion counts at threshold 0.5. The positive-label coding is stated in the model section.

Confusion table

MeasurePredicted positivePredicted negativeTotal
Actual positive295 (8.2%)28 (0.8%)323 (9.0%)
Actual negative97 (2.7%)3159 (88.3%)3256 (91.0%)
Total392 (11.0%)3187 (89.0%)3579 (100.0%)
Test measureValue
Testing records3579
Positive cases323
Positive prevalence9.02%
Accuracy96.51%
Sensitivity / recall91.33%
Specificity97.02%
Precision / PPV75.26%
F1 score0.825

Area under curve

MeasureValue
Area under curve0.975
ROC chart
ROC chart. Native Neural Designer report for this project.
Candidate classification within the HTRU2 benchmark; a model score does not confirm an astronomical discovery.

7. Inference and reproducibility

Open the downloaded project in Neural Designer, inspect the dataset roles and preprocessing, then review the saved task report. Use the same input schema and category order when calculating outputs. The ZIP contains the exact current .nd, its source data and the applicable dataset notices.

Workflow: source measurements → schema and availability checks → model output → domain review. Keep model versions, validation evidence and incoming-data monitoring together.

8. Validity, uncertainty and limitations

Class prevalence affects precision. Internal row-level testing does not establish performance on another telescope, observing campaign or detection pipeline. Review candidates with independent observational evidence and check calibration before interpreting scores as probabilities.

References