Estimating vehicle CO₂ emissions from rated specifications
Vehicle specifications and fuel-consumption ratings provide a structured basis for an emissions estimator. This example models CO₂ emissions and examines how predictions vary within the supplied vehicle dataset.
1. Industrial challenge
Vehicle-rating analysis can use engine characteristics and fuel-consumption ratings to estimate a corresponding CO₂ rating. The close relationship between consumption and emissions shapes how these results should be interpreted.
Vehicle specifications
Combine numerical attributes with make, model and vehicle class.
Rated emissions
Estimate the target in g/km on the prepared rating dataset.
Catalogue analysis
Use the model within its recorded specification range.
2. Data set
The prepared dataset contains 7,385 vehicle-rating records. Inputs include make, model, vehicle class, engine characteristics and fuel consumption. Categorical expansion produces 2,127 input features. Fuel consumption and CO₂ ratings are closely linked; the result is a rating-data estimator rather than an independent measurement of road emissions.
The downloadable project, saved report and supplied source CSV define the exact version used here. Repository: original dataset/source record.
| Subset | Records |
|---|---|
| Training | 4431 |
| Validation / selection | 1477 |
| Testing | 1477 |
| Unused | 0 |
| Variable | Role | Type | Encoding | Unit |
|---|---|---|---|---|
| brand | Input | Categorical | 42 categories; see schema | As supplied |
| model | Input | Categorical | 2053 categories; see schema | As supplied |
| vehicle_class | Input | Categorical | 16 categories; see schema | As supplied |
| engine_size | Input | Numeric | L | |
| cylinders | Input | Numeric | As supplied | |
| transmission | Input | Categorical | A; AM; AS; AV; M | As supplied |
| fuel_type | Input | Categorical | D; E; N; X; Z | As supplied |
| fuel_consumption_city | Input | Numeric | L/100 km | |
| fuel_consumption_hwy | Input | Numeric | As supplied | |
| fuel_consumption_comb(l/100km) | Input | Numeric | As supplied | |
| fuel_consumption_comb(mpg) | Input | Numeric | As supplied | |
| co2_emissions | Target | Numeric | g CO₂/km |
Interactive chart: co2_emissions distribution. Enable JavaScript to explore it.
Interactive chart: co2_emissions Pearson correlations chart. Enable JavaScript to explore it.
Interactive chart: co2_emissions vs. fuel_consumption_comb(mpg) scatter chart. Enable JavaScript to explore it.
3. Model
The final model has 2127 encoded input features and 1 outputs. The following dimensions describe the final saved network.
| Layer | Input shape | Output shape | Activation |
|---|---|---|---|
| Scaling | 2127 | 2127 | — |
| Dense | 2127 | 3 | Tanh |
| Dense | 3 | 1 | Identity |
| Unscaling | 1 | 1 | — |
| Clamping | 1 | 1 | — |
The final output is a continuous estimate. Scaling, unscaling and any bounds follow the saved implementation; the calculator does not silently impose physical constraints.

4. Training strategy
The saved training configuration uses NormalizedSquaredError with QuasiNewton. Training minimizes the recorded objective; the validation subset monitors generalization during fitting. The testing subset is used for the evaluation below.
Interactive chart: Quasi-Newton method error history. Enable JavaScript to explore it.
Quasi-Newton method results
| Measure | Value |
|---|---|
| Epochs number | 199 |
| Elapsed time | 00:00:32 |
| Stopping criterion | Maximum validation error increases |
| Training error | 0.006 |
| Validation error | 0.004 |
5. Model selection
No model selection experiment is recorded in this supplied project. The displayed architecture is the trained model used for testing; earlier article claims about a different selected architecture do not apply to this version.
6. Testing analysis
The final model is evaluated on 1,477 testing observations. MAE, RMSE and prediction R² below are computed from the exact testing pairs stored by Neural Designer. Errors retain each target’s original units.
| Target | Unit | MAE | RMSE | Prediction R² | Pearson r² |
|---|---|---|---|---|---|
| co2_emissions | g CO₂/km | 2.689 | 5.157 | 0.992 | 0.993 |
The native goodness-of-fit task reports squared Pearson correlation (r²). Prediction R² here is 1 − squared prediction error / squared deviation from the test mean. They measure different properties: a strongly correlated prediction can still have a bias or an incorrect scale.
Interactive chart: co2_emissions goodness-of-fit chart. Enable JavaScript to explore it.
7. Model deployment
Validated inputs → saved preprocessing → neural network → score or estimate → domain review. The ZIP contains the original project, source CSV, schema, test metrics and standalone interactive chart exports. The project hash in the schema identifies this exact version.
Directional response at a fixed reference point
Inspect the saved reference point
| Measure | Variable | Value |
|---|---|---|
| 1 | brand | JAGUAR |
| 2 | model | F-TYPE R AWD Convertible |
| 3 | vehicle_class | TWO-SEATER |
| 4 | engine_size | 5 |
| 5 | cylinders | 8 |
| 6 | transmission | AS |
| 7 | fuel_type | Z |
| 8 | fuel_consumption_city | 15.2 |
| 9 | fuel_consumption_hwy | 9.8 |
| 10 | fuel_consumption_comb(l/100km) | 12.7 |
| 11 | fuel_consumption_comb(mpg) | 22 |
Interactive chart: co2_emissions – fuel_consumption_city directional output. Enable JavaScript to explore it.
Explore the exported model
This research demonstration runs locally in your browser. Values outside the training range are outside the validated domain and are rejected. A valid input range does not guarantee that a combination is physically or operationally plausible.
This is not a certified control or protection system.
8. Scope and limitations
Related vehicle variants may cross the row-based split. Hold out model families and model years to test transfer. Do not interpret directional curves as causal effects, or vary city consumption while assuming every correlated specification can remain physically unchanged. Road conditions and driving behaviour are outside this dataset.
No external validation or independent calibration study is included. Preprocessing statistics and model choices should be refitted within a prospective or grouped validation design. Correlations and directional responses describe associations, not causes. Human review is required before an operational decision.



