Skip to content
Learning

Hydraulic system anomaly detection

Detect internal pump-leakage patterns in hydraulic test-rig cycles

A 17-input autoencoder learns a normal baseline from no-leak hydraulic test-rig cycles and ranks held-out cycles by reconstruction error. The final saved model reaches ROC AUC 0.902. Its automatic threshold is conservative: it flags 246 of 984 labelled leakage cycles and 6 of 245 no-leak cycles, so the score is useful for prioritization but the default operating point misses most leakage examples.

2,20560-second cycles
17cycle-mean signals
984labelled leak cycles
0.902test ROC AUC

1. Industrial challenge

Internal leakage reduces a hydraulic pump’s volumetric efficiency and can change several measurements at once rather than creating one universal sensor limit. A reconstruction model can rank cycles whose combined pressure, flow, temperature, vibration and motor-power pattern differs from the no-leak baseline. The score supports condition-monitoring review; fault confirmation and maintenance decisions still require operating context and engineering diagnostics.

Focus the use case

Rank cycles for internal pump-leakage review rather than claiming a general hydraulic diagnosis.

Expose severity

Report weak and severe leakage separately so aggregate recall does not hide missed early-stage cases.

Control the operating point

Choose threshold and persistence from the acceptable inspection workload, then monitor drift and false alerts over time.

Reliability engineeringMaintenanceData science
This model detects patterns associated with the packaged internal pump-leakage label. It does not diagnose cooler, valve or accumulator condition, quantify leakage flow, estimate remaining useful life or replace hydraulic inspection.

2. Data set

The packaged hydraulic_anomaly_detection.csv contains 2,205 complete rows, 17 numeric inputs and one binary evaluation label. A direct comparison with the official UCI files confirms that every row represents one of the test rig’s 60-second load cycles and that each input is the arithmetic mean of one complete sensor channel during that cycle. The packaged rows are reordered, but all 2,205 official cycles are represented once, with no missing or exact duplicate rows.

Input familyPackaged variablesPhysical quantityOfficial sampling rate
Pressurepressure_1_bar–pressure_6_barSix pressure channels (bar)100 Hz
Electricalmotor_power_wMotor power (W)100 Hz
Flowflow_1_l_min–flow_2_l_minVolume flow (L/min)10 Hz
Temperaturetemperature_1_c–temperature_4_cFour temperatures (°C)1 Hz
Vibrationvibration_mm_sVibration (mm/s)1 Hz
Derived process signalscooling_efficiency_pct, cooling_power_kw, efficiency_factor_pctCooling efficiency, cooling power and efficiency factor1 Hz

The binary label also matches the official cycle annotation exactly: pump-leakage class 0 becomes anomaly = 0; weak leakage (class 1) and severe leakage (class 2) become anomaly = 1. Cooler, valve and accumulator conditions are present in the source experiment but are not the reference target in this derived file.

SubsetNo leakWeak leakageSevere leakageTotal
Training73200732
Validation24400244
Testing2454924921,229
Total1,2214924922,205
Training and validation contain only no-leak cycles; all weak- and severe-leak cycles are reserved for testing. This is appropriate for demonstrating novelty detection, but every row comes from the same laboratory rig and the test set is deliberately enriched to 80.1% leakage. The reported precision and accuracy therefore do not estimate field alarm rates. Averaging each 60-second signal also removes waveform, transient and spectral information that may matter in a deployed monitoring system.

3. Model

The final autoencoder has 17 → 8 → 4 → 8 → 17 units and 373 trainable parameters. Inputs are standardized with the training means and standard deviations; the three hidden stages use ReLU and the reconstruction output is linear before values are returned to their original units. The pump-leakage label is used only for evaluation and never enters the reconstruction path.

The four-unit bottleneck forces the network to represent the normal operating pattern compactly. A high error indicates poor reconstruction of the complete sensor pattern; it does not identify a physical root cause by itself.

Saved hydraulic autoencoder layer widths from input to reconstruction
Native Report neural network output for the 17 → 8 → 4 → 8 → 17 hydraulic autoencoder; select the diagram to inspect it at full resolution.

4. Training strategy

The final saved model minimizes mean absolute error with Adam (learning rate 0.001, batch size 32) and no regularization. It reaches the configured maximum of 1,000 epochs with standardized training error 0.107 and validation error 0.119 on the normal-only subsets. The close final values do not establish robustness across random initializations; they show only that this run fits the two stored normal blocks similarly.

Native Perform training result from the final saved model. Training and validation contain only no-leak cycles.

5. Model selection

No executed architecture-selection comparison is saved. The 17 → 8 → 4 → 8 → 17 network is therefore the evaluated configuration, not a demonstrated optimum. A professional study would compare bottleneck sizes and random seeds on a separate normal validation period, freeze the selected model, and only then evaluate leakage cycles.

Threshold sensitivity

The automatic rule uses mean training error plus one standard deviation. The candidate table comes from the final Neural Designer threshold task. Validation-normal and test columns are consequences of each rule; test performance must not be used to select a threshold and then reported as independent evidence.

Candidate ruleThresholdValidation normal flaggedTest cycles flaggedLeak recallPrecision
Training P951.1425.33%48.3%59.2%98.1%
Mean + 1 SD (reported)2.3893.28%20.5%25.0%97.6%
Training P993.0481.23%18.8%22.9%97.4%
Mean + 2 SD4.2120.82%18.8%22.9%97.4%
The default threshold favours a small review queue and high test precision, but it is not a useful high-sensitivity alarm in this run. A field threshold needs a normal-only calibration period, realistic fault prevalence and a persistence rule tied to the acceptable inspection workload.

6. Testing analysis

The final anomaly score is mean absolute reconstruction error in the variables’ original units. Across 732 normal training cycles its mean is approximately 0.567 and its standard deviation is approximately 1.823, giving the exported automatic mean-plus-one-standard-deviation threshold 2.38958. The model flags 252 of 1,229 test cycles.

The score still ranks leakage-labelled cycles reasonably well (ROC AUC 0.902), but the chosen operating point recovers only 25.0% of them. This distinction matters: ROC AUC describes ranking across all thresholds, while the confusion matrix describes one specific review threshold.

The score averages reconstruction errors measured in bar, watts, L/min, °C, mm/s, percent and kW, so it changes if units or feature definitions change. Use the standardized sample-reconstruction charts to inspect contributing signals and recalibrate the score for any production feature pipeline.
Test measureValueCount or context
ROC AUC0.902Score ranking across 1,229 packaged test cycles
Leak recall25.0%246 of 984 weak or severe leakage cycles
No-leak specificity97.6%239 of 245 no-leak cycles
Precision on this enriched subset97.6%246 of 252 flagged cycles; prevalence-dependent
F1 score39.8%At exported threshold 2.38958
Accuracy39.5%485 of 1,229 test cycles
Official pump-leakage stateFlaggedNot flaggedDetection rate
No leak (class 0)62392.4% flagged
Weak leakage (class 1)10538721.3% detected
Severe leakage (class 2)14135128.7% detected
Native Calculate anomaly threshold export for the final training run and its 732 normal training cycles. The dashed line is the reported mean-plus-one-standard-deviation threshold.
Native Calculate anomaly threshold export for the final training run. The test subset is deliberately leakage-enriched; 252 of 1,229 cycles exceed the threshold.
Native Calculate anomaly ROC curve result from the final model and the packaged 1,229-cycle test subset.
Native Plot sample reconstruction export for no-leak sample 1743. Its reproduced score is approximately 0.414, below the 2.38958 threshold. The reconstruction pair shares the same vertical scale.
Native Plot sample reconstruction export for weak-leak sample 892. Its reproduced score is approximately 0.969, below the 2.38958 threshold, so this example exposes a false negative at the reported operating point.
Native per-signal absolute standardized residuals for no-leak sample 1743. The residual pair uses the same vertical scale.
Native per-signal absolute standardized residuals for weak-leak sample 892. Its moderate residual pattern remains below the automatic threshold, illustrating why the model misses many weak-leak cycles.

7. Model deployment

In Neural Designer, run Perform training to reproduce the learning curve, Calculate anomaly threshold to inspect the training and testing score distributions, Calculate anomaly ROC curve to assess ranking performance and Calculate anomaly detection tests to recover the operating-point metrics. For a selected cycle, Plot sample reconstruction reports both standardized input-versus-reconstruction bars and per-variable reconstruction errors, which is the most useful diagnostic view for signals with different units.

  1. AcquirePreserve asset, cycle, timestamp, duty point and sensor-health metadata.
  2. TransformReproduce the same 60-second aggregation, variable names, units and ordering.
  3. ScoreStore the versioned reconstruction error and per-variable residuals.
  4. GateApply a threshold and persistence rule calibrated to the acceptable inspection workload.
  5. ConfirmReview trends and contributing signals with hydraulic diagnostics before maintenance action.

Export anomaly detector is appropriate only after validating that complete chain on separate equipment or test campaigns.

8. Scope and limitations

This is an internal benchmark on one laboratory test rig. The split does not isolate a different rig, asset, campaign or duty profile, and the test subset contains far more leakage cycles than a typical monitoring stream. Cycle means discard within-cycle dynamics. The binary target also combines weak and severe leakage; at the reported threshold the detector recovers only 21.3% of weak-leak and 28.7% of severe-leak cycles.

The reconstruction score is unit-sensitive and is not a leakage magnitude, remaining-useful-life estimate or universal alarm limit. Before field use, validate independent assets and operating regimes, retain the original severity labels, calibrate threshold and persistence on realistic normal operation, quantify false alerts per asset-hour, evaluate detection lead time, test sensor failures and drift, and define an engineering confirmation procedure. The detector must not command an automatic shutdown or replace hydraulic inspection.

References

Project evidence: the current saved hydraulic.nd project, hydraulic_anomaly_detection.csv, native Neural Designer tasks executed on that saved model and a row-level comparison against the official sensor matrices and profile.txt. Primary source: Helwig, Pignanelli and Schütze, Condition Monitoring of Hydraulic Systems, UCI Machine Learning Repository, DOI 10.24432/C5CW21. See Neural Designer model types.