Skip to content
Blog

Assess death risk after heart failure using machine learning

Studying mortality after heart failure

Can recorded clinical measurements distinguish patients who died during follow-up from those who did not? This example trains an 11-input classifier in Neural Designer and examines its discrimination, missed events and false positives on a held-out test subset.

299Source records
11Final input features
59Test observations
0.776Test ROC AUC

1. Research question and intended use

The modelling question concerns the recorded outcome in this cohort. It is a reproducible classification exercise for studying clinical data, with explicit attention to score direction and the errors at a chosen threshold.

Recorded outcome

DEATH_EVENT=1 identifies a death observed during follow-up.

Clinical measurements

The model uses 11 demographic, laboratory and clinical inputs.

Held-out evidence

Inspect the ROC curve alongside the 8 missed deaths and 7 false positives at a threshold of 0.5.

Clinical-data researchModel evaluationHealth analytics
Research demonstration: the score describes this recorded binary outcome and does not establish an individual treatment decision.

2. Cohort, measurements and endpoint

The prepared Heart Failure Clinical Records table contains 299 patient records, with 96 recorded deaths and 203 records without a death event. The endpoint is DEATH_EVENT: 1 means death observed during follow-up; 0 means no death observed during that period. The prepared CSV has 11 predictors and excludes the original follow-up-time column. Follow-up duration varies in the source cohort, so this binary endpoint is not a common fixed-horizon mortality outcome.

The downloadable project, saved report and supplied source CSV define the exact version used here. Repository: original dataset/source record.

SubsetRecords
Training181
Validation / selection59
Testing59
Unused0
VariableRoleTypeEncodingUnit
ageInputNumericyears
anaemiaInputBinary0; 10=no; 1=yes
creatinine_phosphokinaseInputNumericmcg/L (source convention)
diabetesInputBinary0; 10=no; 1=yes
ejection_fractionInputNumeric%
high_blood_pressureInputBinary0; 10=no; 1=yes
plateletsInputNumerickiloplatelets/mL (source convention)
serum_creatinineInputNumericmg/dL
serum_sodiumInputNumericmEq/L
sexInputBinary0; 10=woman; 1=man
smokingInputBinary0; 10=no; 1=yes
DEATH_EVENTTargetBinary0; 10=no recorded death; 1=recorded death

Interactive chart: DEATH_EVENT pie chart. Enable JavaScript to explore it.

DEATH_EVENT pie chart. Exported with Neural Designer from the saved task report.

Interactive chart: DEATH_EVENT Pearson correlations chart. Enable JavaScript to explore it.

DEATH_EVENT Pearson correlations chart. Exported with Neural Designer from the saved task report.
The downloadable project, saved report and supplied source CSV define the exact version used here. Repository: original dataset/source record. This is internal validation using the saved record-level split. Grouped or temporal independence has not been established.

3. Model

The final model has 11 encoded input features and 1 outputs. The following dimensions describe the final saved network.

LayerInput shapeOutput shapeActivation
Scaling1111—
Dense113Tanh
Dense31Sigmoid

Output semantics: the sigmoid score increases toward 1; 0 is the other class. Calibration has not been evaluated, so scores are not presented as calibrated probabilities.

Studying mortality after heart failure: initial Neural Designer architecture
Architecture used for this model; no architecture-selection experiment is recorded. The model has 11 inputs, a hidden layer with 3 tanh neurons and one sigmoid output.

4. Training strategy

Training uses weighted squared error, L2 regularization (weight 0.01) and the Quasi-Newton optimizer. Training minimizes the recorded objective; the validation subset monitors generalization during fitting. The testing subset is used for the evaluation below.

Interactive chart: Quasi-Newton method error history. Enable JavaScript to explore it.

Quasi-Newton method error history. Exported with Neural Designer from the saved task report.

Quasi-Newton method results

MeasureValue
Epochs number111
Elapsed time00:00:00
Stopping criterionMaximum validation error increases
Training error0.592
Validation error0.741

5. Model selection and baseline

No model selection experiment is recorded in this supplied project. The displayed architecture is the trained model used for testing; earlier article claims about a different selected architecture do not apply to this version.

A transparent test-set comparator is the majority-class rule, with accuracy 67.8%. This is a baseline for interpretation, not an alternative model fitted on the test labels.

6. Testing and interpretation

In the test subset, 19 of 59 patients have DEATH_EVENT=1 (32.2%). At the 0.5 threshold, the model identifies 11 of these events and misses 8; it also flags 7 of the 40 records without a death event. For the death-event class, sensitivity is 57.9%, specificity 82.5% and positive predictive value 61.1%.

The final classifier is evaluated on 59 testing records. The confusion counts below were reproduced from the saved model. Rows are actual classes and columns are predicted classes. The decision threshold is 0.5 on the score for 1.

Test class prevalence is shown by the support counts. Accuracy is 74.6% and macro F1 is 0.705. ROC AUC is 0.776. The native ROC optimal threshold is descriptive of this test set and is not an independently validated operating policy.

Actual / predicted01Total
033740
181119
ClassTest casesSensitivity / recallSpecificityPrecision / PPVF1
04082.5%57.9%80.5%0.815
11957.9%82.5%61.1%0.595

The native ROC task reports AUC 0.776 with 90% confidence limits 0.661–0.891. Its test-derived threshold of 0.382 gives sensitivity 73.7% and specificity 75.0% on these same cases. This exploratory cutoff was not selected on an independent validation set; the confusion matrix above retains the default threshold of 0.5.

Interactive chart: ROC chart. Enable JavaScript to explore it.

ROC chart. Exported with Neural Designer from the saved task report.
This is internal record-level evidence. Discrimination does not establish calibration, clinical utility or benefit to patients.

7. Workflow and reproducibility

Check inputs → apply the saved preprocessing and network → obtain a DEATH_EVENT score → review the research result. The ZIP contains the original project, source CSV, schema, test metrics and standalone interactive chart exports. The project hash in the schema identifies this exact version.

Explore the exported model

This research demonstration runs locally in your browser. Values outside the saved input range are outside the validated domain and are rejected. Plausible individual values do not guarantee that their combination is represented in the cohort.

The output is an uncalibrated score for DEATH_EVENT=1, not a personal mortality probability. It must not guide medical treatment.

8. Safety, generalizability and governance

The 59-record test subset is small and comes from the same historical cohort as the training records. This binary classifier does not model event time or censoring and cannot provide a validated 30-day or one-year mortality risk. Before any medical use, establish predictor availability at the assessment time, refit preprocessing within training, evaluate on an independent cohort and obtain clinician review.

No external validation or independent calibration study is included. The saved scaling statistics cover the cohort rather than a separately verified training-only preprocessing fit. Preprocessing statistics and model choices should be refitted within a prospective or grouped validation design. Correlations and directional responses describe associations, not causes. Human review is required before an operational decision.

The saved ROC task provides an AUC confidence interval, but subgroup performance, calibration curves and decision-cost validation are not established by these tasks. Predictive values apply to the observed test class distribution and may change when prevalence shifts.

Expert review and separate external validation are required before any medical use.

References