Chapter 1 DL Classification of Weather Patterns over Europe

Author: Ziyu Mu, Laura Schlueter

Supervisor: Henri Funk

Degree: Master

1.1 Abstract

Daily weather in the mid-latitudes is dominated by large-scale atmospheric circulation types (CTs), which influence regional climate variability and extremes. Accurate classification of these CTs is crucial for diagnosing long-term climate dynamics and analysing future changes. This study implements and extends a novel deep learning-based classification method for European CTs, building on the Hess & Brezowsky (HB) framework.

Using a convolutional neural network (CNN) architecture (inspired by applications in climate science and optimization techniques from), we trained the model on high-resolution ERA5 reanalysis data (1950–1980), with sea-level pressure (SLP) and geopotential height at 500 hPa (z500) fields as key predictors. To ensure robustness, a nested cross-validation approach was employed, achieving an overall accuracy of 53.17% and a macro F1 score of 55.14%, outperforming traditional classification methods.

Applied to the CanESM2 ensemble projections, which simulate data with assumptions that climate conditions were consistent with history, our results show significant future frequency shifts in CTs, notably an increase in Icelandic High, Cyclonic (HNZ) occurrences during both summer and winter half-years. This methodological advancement not only enhances classification accuracy but also offers robust ensemble analyses critical for climate-informed decision-making.

1.2 Introduction

In the mid latitudes, daily weather is dominated by large-scale atmospheric circulation patterns, defined by the position of high and low pressure centres (Häckel 2021; Mittermeier et al. 2022). Due to steep temperature and pressure gradients along the frontal zone between polar and tropical air masses (around 45° latitude), the jet stream is formed in about 10 km altitude. It is characterized by mean wind speeds of 300 km/h and the long-term average wind direction is westerly (Häckel 2021). Here, dynamic high and low pressure systems develop (Häckel 2021; Mittermeier et al. 2022). However, the westerly jet stream is not stationary, but deflected with varying amplitude creating the so-called Rossby waves, which influence the position of high and low pressure systems (Häckel 2021). While an infinite number of atmospheric conditions would be conceivable, recurrent weather patterns with similar meteorological features are observed in practice. Therefore, it is possible to classify weather patterns according to synoptic properties (Bissolli 2001).

The first concept for a classification of different weather patterns was introduced by Baur, Hess, and Nagel (1944). They used surface pressure charts from 1881 to 1939 and assigned the respective weather situation to a specific class for each day based on the spatial distribution of atmospheric pressure and the position of the frontal zones. This lead to a classification into 21 Grosswetterlagen for Europe and East Atlantic (Baur, Hess, and Nagel 1944). In 1963 Baur defined Grosswetterlage as the mean spatial air pressure distribution of a large area, at least the size of Europe, over a period of several successive days. This classification has been revised and improved several times. In 1977, Hess and Brezowsky published their third edition of a revised catalogue of Grosswetterlagen. Here, the authors determined that a Grosswetterlage should only be classified if it can be recognised on at least three successive days. Furthermore, the number was extended to 29 different Grosswetterlagen. In addition to the surface pressure chart, the classification was also based on the geopotential height in 500hPa (Bissolli and Dittmann 2001; Bissolli 2001; Werner and Gerstengarbe 2010). In the following, Werner and Gerstengarbe published seven editions of a catalogue of the Grosswetterlagen in Europe with updated data. The last edition (2010) covers the years 1881 to 2009 and provides daily information on the Grosswetterlage over Europe. Today, the German Weather Service (DWD) still constantly updates the catalogue of European Grosswetterlagen and publishes the results monthly. Since 1944, the classification of Grosswetterlagen is conducted by experts, hence it is subjectively biased (Hess 2005). In the following, the classification according to Hess & Brezwosky will be referred to as HB CT.

HB CTs can be attributed to certain weather situation at certain locations in Europe (Werner and Gerstengarbe 2010). Hence in the past, the classification of HB CTs has been used to identify regularities in frequency and duration of occurrences of certain weather conditions. Bissolli and Dittmann (2001) describes a relationship between frequency and duration of the HB CTs and mean annual air temperature and precipitation. Moreover, extreme weather events—such as heavy rainfall, floods, and heat waves—can often be linked to specific HB CTs (Mittermeier et al. 2022). Such kinds of application could not only be useful to analyse past climate. Future climate projections could also be analysed using the HB CTs, allowing conclusions to be drawn about future developments. However, analyzing future CTs requires many model ensembles to account for internal variability (Wyser et al. 2021). Manual classification of HB CTs for many ensemble members is no longer possible. It is therefore necessary to consider the possibility of automatization (Mittermeier et al. 2022). Mittermeier et al. (2022) have found a way to classify HB CTs automatically using a deep learning classifier.

1.3 Data

1.3.1 Data Sources

This section describes in detail the datasets utilized for model training, validation, and testing, providing comprehensive explanations of their characteristics, purposes, and roles within the study. Due to data availability constraints, we replaced the original study’s ERA-20C (Poli et al. 2016) and SMHI-LENS (Wyser et al. 2021) data with ERA5 (European Centre for Medium-Range Weather Forecasts 2017) (higher resolution, 1950–1980) for training and CanESM2 (Hua et al. 2015; Government of Canada 2025) (use only 12 of 50 ensemble members) for projections.

1.3.1.1 Historical Data - Training and Validation

We use reanalysis dataset, which is the historical atmospheric data reconstructed by combining observational data and model simulations, for training and validation. Data is provided by European Centre for Medium-Range Weather Forecasts (ECMWF).

We only use two variables mentioned above:

  1. SLP: Atmospheric pressure measured at Earth’s surface, vital for identifying weather systems such as cyclones and anticyclones.
  2. z500: Height of the 500 hPa pressure level in the atmosphere, indicative of atmospheric wave patterns and mid-tropospheric dynamics.

Due to reanalysis data’s high-quality, accurate, and detailed historical atmospheric representation, it is enabled to get a robust and precise training of the deep learning model.

The difference of ERA5 (Figure 1.1) and ERA-20C (Appendix Figure 1.7) lies on:

Difference between two reanalysis dataset
Reanalysis Dataset Temporal Coverage Spatial Resolution
ERA-20C (paper) 1900–1980 Approximately 125 km
ERA5 (implementation) 1950–1980 High-resolution, about 31 km
Synoptic patterns of the 29 HB CTs created using ERA5 data averaged over the period 1950 - 1980, showing SLP on the left and z500 on the right side.

FIGURE 1.1: Synoptic patterns of the 29 HB CTs created using ERA5 data averaged over the period 1950 - 1980, showing SLP on the left and z500 on the right side.

1.3.1.2 Future Climate Projections - Change Analysis

We use ensemble datasets which projected future climate conditions under specific scenarios, to evaluate future changes in atmospheric circulation, considering internal variability, essential for accurate and reliable climate change projections. The data should be preprocessed with the same variables, scale, domain, resolution as training data.

  1. CanESM2 (Government of Canada 2025)
    • Name: CanESM2 (Canadian Earth System Model version 2)
    • Provider: Canadian Centre for Climate Modelling and Analysis (CCCma)
    • Model Used: Climate Model Intercomparison Project phase 5 (CMIP5)
    • Ensemble Members: 50 (due to downloading limitation, we only used 12 ensemble)
    • Scenario: historical, simulate the condition as observed past climate from 1850 to at least 2005 (Friedlingstein et al. 2008).
    • Temporal Coverage:
      • Historical Reference: 1971–2000
      • Future Projections: 2031–2060
    • Spatial Resolution: Around 2.8°
  2. SMHI-LENS (Wyser et al. 2021)
    • Name: SMHI-LENS (Swedish Meteorological and Hydrological Institute Large Ensemble)
    • Provider: Swedish Meteorological and Hydrological Institute
    • Model Used: EC-Earth3 (Climate Model Intercomparison Project phase 6, CMIP6)
    • Ensemble Members: 50 (to effectively capture internal variability)
    • Scenario: SSP3-RCP7.0, representing a high GHG emission pathway
    • Temporal Coverage:
      • Historical Reference: 1991–2020
      • Future Projections: 2071–2100
    • Spatial Resolution: Around 0.7°

As a result with higher greenhouse gas emission under SSP3-RCP7.0 scenario, and with a more distant future in the paper:

  1. Future warming and circulation shifts might be more pronounced than our implementation.
  2. Frequency changes in HB CTs may appear larger in the paper, simply because the climate change signal is stronger.

1.3.2 Data Preprocessing Steps

  1. Reset spatial domain: only use data in Europe and the North Atlantic region (30°N–75°N latitude, 65°W–45°E longitude)
  2. Spatial Regridding: All datasets were interpolated to a common 5° grid resolution, ensuring consistency and computational efficiency
  3. Seasonal Centering: Applied to remove seasonal biases, helping the model better recognize atmospheric patterns independently of seasonal influences.

1.4 Models and Training

1.4.1 CNN Architecture

The classification model used in this study is based on a CNN, specifically designed to handle the image-like structure of atmospheric data (SLP and z500). The CNN architecture (Figure 1.2) comprises:

  1. Input Layer: Two separate input channels corresponding to SLP and z500.
  2. Convolutional Layers: Two convolutional layers applied to extract spatial patterns and features from input atmospheric fields. These layers use convolutional kernels to identify local and spatially coherent patterns.
  3. Dropout Layer: Included for regularization to mitigate overfitting by randomly dropping out units during training, thus improving the model’s generalization.
  4. Fully Connected Layers: Two fully connected layers following the convolutional and dropout layers, used for integrating learned spatial features into final predictions of HB CTs.
  5. Output Layer: Outputs softmax probabilities for each of the 29 HB CTs, which are then converted into classifications.
Schematic diagram of the CNN network

FIGURE 1.2: Schematic diagram of the CNN network

1.4.2 Training Protocol

1.4.2.1 Adam Optimizer

Adam (Adaptive Moment Estimation) (Mehta, Paunwala, and Vaidya 2019) is an optimization algorithm widely used for training neural networks. It combines the benefits of two other methods—AdaGrad (Ward, Wu, and Bottou 2020), which adapts the learning rate for each parameter, and RMSProp (F. Zou et al. 2019), which considers the exponentially decaying average of squared gradients—to adjust the learning rates adaptively during training.

Adam optimizer efficiently handles noisy and sparse gradients, common in deep learning tasks, facilitating faster convergence and better handling of large parameter spaces typical of CNNs. This makes Adam particularly suitable for our CNN-based classification model, ensuring effective and efficient training.

1.4.2.2 Bayesian Optimization for Hyperparameter Tuning

To achieve an optimal configuration of hyperparameters for our CNN, we employed Bayesian optimization (Garnett 2023), specifically using the Tree-structured Parzen Estimator (TPE) method (Ozaki et al. 2022) implemented in the Optuna library. Bayesian optimization is particularly suited for deep learning hyperparameter tuning due to its efficiency and ability to intelligently balance exploration and exploitation of the hyperparameter space, thus significantly reducing computational cost compared to grid or random search.

Specifically, the following hyperparameters were optimized:

  1. Learning Rate: The learning rate controls the magnitude of updates to the network weights during training. We allowed this hyperparameter to vary logarithmically between \(10^{-4}\) and \(10^{-2}\), enabling the optimizer to effectively explore both subtle and more substantial updates in network parameters.
  2. Weight Decay: Serving as a regularization term to prevent overfitting, the weight decay was also sampled on a logarithmic scale ranging from \(10^{-5}\) to \(10^{-3}\). This enabled fine control over model complexity and encouraged better generalization.
  3. Dropout Rate: Dropout prevents overfitting by randomly dropping neurons during training, forcing the model to learn robust representations. We tuned dropout rate uniformly within a range from 0.2 to 0.6, allowing optimization of the model’s regularization intensity.
  4. Convolutional Layer Parameters:
    • Number of Output Channels in First Convolutional Layer (out_channels1): Chosen categorically from [4, 8, 16], optimizing the depth and feature extraction capability of initial convolutional operations.
    • Number of Output Channels in Second Convolutional Layer (out_channels2): Selected from [8, 16, 32], further refining the network’s capability to learn hierarchical spatial patterns.
    • Kernel Size: Kernel size impacts the spatial context the model considers when learning features. We considered kernel sizes of [3, 5, 7], balancing between capturing detailed local patterns and broader spatial structures.
    • Fully Connected Layer Size: The number of neurons in the fully connected layer was chosen from [32, 64, 128], directly influencing the model’s ability to integrate learned spatial features into higher-level abstractions suitable for classification.

The Bayesian optimization was performed over 100 trials, systematically exploring these hyperparameters. Each trial evaluated a distinct combination of parameters based on validation accuracy. The final selection of hyperparameters was determined by the combination that produced the highest validation accuracy, ensuring the CNN model’s optimal performance on unseen data while minimizing the computational burden associated with hyperparameter tuning.

1.4.2.3 Nested Cross-Validation

Nested cross-validation (Zhong, Chalise, and He 2023) involves two cross-validation loops:

  1. Outer Loop: Evaluates the generalization performance of the model on independent test sets.
  2. Inner Loop: Used for hyperparameter tuning and model selection within each training fold.

We use 5 folds of outer loop and 5 folds of inner loop.

1.4.2.4 Epochs and Early Stopping

Training was set to a maximum of 35 epochs, with early stopping criteria implemented, patience of 6 epochs without improvement, to prevent overfitting and reduce training time.

1.4.3 Predicting Technique

1.4.3.1 Ensemble Learning

Ensemble learning (Dietterich et al. 2002) combines predictions from multiple models to enhance predictive performance. In this case, multiple CNN models trained with different initial conditions. A deep ensemble of 30 independently trained CNN models, each initialized with different random weights, is utilized.

Predictions from individual CNN models are aggregated using a weighted averaging method, where each model’s contribution is proportional to its validation performance. This enhances robustness, reduces prediction uncertainty, and provides stable and reliable classification results by capturing a wider range of atmospheric variability.

1.4.3.2 Transitional Smoothing

According to the definition, a HB CT must last at least three days (Hess and Brezowsky 1952) to represent a stable atmospheric pattern. Without enforcing this constraint, the raw neural network predictions may contain short, noisy transitions that are meteorologically implausible. Thus, transition smoothing is necessary to:

  1. Remove unrealistic, very short-lived classifications.
  2. Produce physically consistent and stable time series of HB CTs.
  3. Ensure that predictions align with domain-specific expectations.

According to the original paper:

  1. Step 1: Identify all transitions where the predicted HB CT lasts fewer than three days.
  2. Step 2: Check neighborhood consistency:
    • If the HB CT before and after the short transition is the same, assign this type to the transition days.
  3. Step 3: If different types occur before and after:
    • Compare the predicted probabilities (confidence scores) for each neighboring type.
    • Assign the transition days to the neighbor with the higher predicted probability (i.e., stronger membership).

This method ensures that short transitions are corrected in a physically meaningful way, based on both label consistency and model confidence. Since the original code is not available, our transition_smoothing function follows the paper’s logic but introduces minor adaptations for greater robustness and flexibility:

Transitional smoothing implementation
Aspect Paper Description Our Code Implementation Reason for Change
Minimum duration 3 days fixed min_duration parameter (default=3) Makes the method flexible for other settings.
Neighbor block length trust Not specified clearly min_neighbor_run_length parameter (default=2) Ensures smoothing only when neighboring segments are stable enough.
Handling first and last segments Implicitly ignored Explicitly handled Avoids out-of-bound errors at sequence edges.
Probability averaging Single-point comparison Average probability across transition segment More robust by considering entire transition region, not a single timestep.
Fallback behavior when neighbors are not strong Not detailed Fallback to the only trusted neighbor if one side is strong enough Avoids wrong smoothing when one neighbor is unreliable.

1.5 Results

1.5.1 Classification Performance

1.5.1.1 F1-score

The classification performance of our model was evaluated using a confusion matrix, F1-scores per class, macro F1-score (Opitz and Burst 2019), and overall accuracy (Alberg et al. 2004), following the approach used in the reference study (Mittermeier et al. 2022). The overall accuracy achieved by our model on the cross-validation test set is 53.17%, and the macro F1-score is 55.14%. Among all HB CTs (Appendix Figure 1.9), the best performing classes in terms of F1-score are WZ (70.21), HM (63.7), and HB (62.78). These types exhibit distinct spatial patterns, aiding CNN identification. The lowest F1-scores are observed for classes like NA (33.61) and SEZ (44.37). These are likely due to class imbalance caused by rare occurrence. Generally, accuracy for each type is quite high, which may imply overfitting.

The confusion matrix (Appendix Figure 1.17) shows that most misclassifications occur among HB CTs with similar dynamical structures, consistent with expectations from atmospheric science. Diagonal dominance is visible but moderate, reflecting the challenging nature of this classification task.

Compared to the paper results, which reported an overall accuracy of 41.1% and a macro F1-score of 38.3% (Mittermeier et al. 2022), our model achieves slightly higher overall accuracy and macro F1-score.

Metric comparison
Metric Paper (%) Our Model (%)
Macro F1-score 38.3 55.1
Overall accuracy 41.1 53.2

The boxplots of F1-score under each HB CT (Figure 1.6; Appendix Figure 1.13) confirm our conclusion:

  1. Original study has considerably lower overall F1-scores (around 0.3–0.5), indicating challenges in reliable classification.
  2. Our model demonstrates much higher stability with F1-scores generally between 0.80–0.95, suggesting improved reliability and robustness in identifying circulation patterns.

1.5.1.2 RMSE

To further evaluate the quality of the predictions, the Root-Mean-Square Error (RMSE) between the predicted HB CT composites and the true label composites was calculated for variable SLP.

Equation for the calculation of the RMSE with I being the predicted image (in our case: signature plot of the deep learning classifier) and K being the reference image (in our case: signature plot of the labels). M are number of rows and N the number of columns of the pictures to compare. The RMSE thus compares the pixel-wise values of two images. A value of zero indicates a perfect match (Mittermeier et al. 2022; Mueller et al. 2020). \[\text{RMSE} = \sqrt{\frac{1}{M\times N}\sum_{i=0,j=0}^{M-1,N-1}[I(i,j)-K(i,j)]^2}\]

Additionally, RMSEs for false positives and false negatives were computed separately.

The average RMSE for our implementation’s predictions across all HB CTs is 0.88, while the false positives and false negatives yielded RMSEs of 1.08 and 1.23, respectively.

RMSE comparison
RMSE Category Paper Implementation
Prediction RMSE 0.89 0.74
False positives RMSE 1.09 1.45
False negatives RMSE 1.28 0.69

Our implementation achieves a lower prediction RMSE (0.74) compared to the original paper (0.89) (Mittermeier et al. 2022). This indicates that, on average, the spatial patterns of the correctly predicted HB CTs are more similar to the true labels in our model than in the reference study.

However, the false positives RMSE is much higher (1.45) than in the paper (1.09) (Mittermeier et al. 2022). This suggests that when our implementation wrongly predicts a HB CT, the resulting spatial pattern is less similar to the true pattern than in the paper.

Our implementation shows a lower false negatives RMSE (0.69) compared to the paper (1.28) (Mittermeier et al. 2022). This means that when the model misses a true HB CT, the spatial signature remains relatively closer to the correct pattern than in the paper.

The overall pattern suggests that:

  1. Our implementation is better at maintaining correct structures when making predictions or missing true types.
  2. However, when making incorrect predictions (false positives), our implementation model tends to produce worse distortions than the model used in the paper.

1.5.2 Signature Plot Analysis

Figure 1.3 shows the signature plots for four HB CTs selected based on our RMSE results:

  1. Two HB CTs with lowest RMSE: WZ and TRW (best performance)
  2. Two HB CTs with highest RMSE: NA and HNFA (worst performance)

For each HB CT, four different composite plots are displayed:

  1. Labels: Composite of true labels.
  2. Predictions: Composite of model predictions.
  3. False Positives: Days predicted as the HB CT but labeled differently.
  4. False Negatives: Days labeled as the HB CT but predicted differently.

RMSE values are shown below each panel, comparing each composite to the true label composite.

Signature plots of four selected circulation patterns at slp. Column 1: labels showing the indicated HB CT, Column 2: deep ensemble predictions showing the indicated HB CT, Column 3: signature pattern, when the deep ensemble predicts the indicated HB CT while labels state differently, Column 4: labels stating the indicated HB CT while deep ensemble predicts differently. The RMSE values are calculated by comparing the respective signature plot to the signature plot of the labels (column 1).

FIGURE 1.3: Signature plots of four selected circulation patterns at slp. Column 1: labels showing the indicated HB CT, Column 2: deep ensemble predictions showing the indicated HB CT, Column 3: signature pattern, when the deep ensemble predicts the indicated HB CT while labels state differently, Column 4: labels stating the indicated HB CT while deep ensemble predicts differently. The RMSE values are calculated by comparing the respective signature plot to the signature plot of the labels (column 1).

Based on our results, we can observe that:

  1. WZ:
    • The predicted composite closely matches the label pattern (RMSE = 0.25).
    • False positives and false negatives show moderate RMSE (0.81 and 0.36), indicating stable identification of WZ.
  2. TRW:
    • Predictions for TRW are also accurate (RMSE = 0.32).
    • False positives (0.92) and false negatives (0.46) suggest TRW is generally recognized reliably.
  3. NA:
    • Predictions for NA have high RMSE (2.89), indicating difficulty in capturing the correct structure.
    • False positives (5.15) show very large errors, and false negatives (0.78) suggest missing NA patterns is common.
  4. HNFA:
    • Predictions for HNFA also show high RMSE (1.20).
    • False positives (2.80) and false negatives (1.09) are large, implying HNFA is a difficult type for the model to classify.

Differences from the original paper (Appendix Figure 1.10) arise primarily from data variations, model differences, and random initialization.

Although the specific CTs selected differ, the general trend remains consistent with the findings of the paper:

  1. Some HB CTs are easier to predict (low RMSE), others are harder (high RMSE).
  2. False positives tend to have higher RMSE than correct predictions.
  3. A larger discrepancy in the color distribution indicates a higher complexity of HB CT, and false negatives vary depending on HB CT complexity.

Thus, while specific HB CTs differ, our results fundamentally agree with the conclusions of the paper that certain patterns are systematically easier or harder to classify, and signature plots are effective to visualize these tendencies.

1.5.3 Frequency Distribution of HB CTs

Figure 1.4 shows the frequency distribution of the 29 HB CTs, expressed as the average number of days per year for the period 1950–1980. For each HB CT, the frequency derived from the true labels (blue bars) is compared to the frequency derived from the network predictions (yellow bars).

Frequency distribution of the 29 HB CTs in number of days per year for the training period 1950–1980 (blue) and for predictions of the CNN (yellow).

FIGURE 1.4: Frequency distribution of the 29 HB CTs in number of days per year for the training period 1950–1980 (blue) and for predictions of the CNN (yellow).

Overall, the CNN reliably reproduces observed HB CT frequencies:

  1. Dominant HB CTs (e.g., WZ, HM) are captured accurately, reflecting robust model learning.
  2. Rarer HB CTs (e.g., NA, SZ) exhibit minor deviations, suggesting challenges in accurately learning patterns due to fewer examples, which may indicate slight overfitting.

In comparison with the original study (Appendix Figure 1.11), despite differences in dataset periods (our study: 1950–1980; original: 1900–1980), frequency patterns remain consistent. Slight discrepancies in rarer HB CTs are likely due to dataset differences, emphasizing the importance of extended or more balanced training datasets for better generalization.

The agreement between the label frequencies and the network predictions, as well as the similarity to the paper’s findings, indicates that our model effectively learns the relative occurrence rates of the different HB CTs. Overall, the frequency distribution results support the conclusions of the original study, confirming that the network can reliably reproduce the occurrence rates of various HB CTs.

1.5.4 Future Change

Figure 1.5 and Figure 1.6 show the relative changes in the frequency of occurrence of the 29 HB CTs between the future period (2031–2060) and the historical reference period (1971–2000), for the entire year, winter half-year, and summer half-year, based on 12 climate model realizations. Figure 1.5 is the conclusion directly from predicted results, while Figure 1.6 is obtained after transition smoothing. Positive values indicate an increase in occurrence; negative values indicate a decrease. The lower panels show the F1-scores of the classification model ensemble. Higher F1-scores indicate better classification stability and reliability. The grouping by wind direction allows visual comparison of performance across different HB CT categories.

Boxplots of frequency change and f1 score without transitional smoothing. Upper plots show the change in the relative frequency of occurrence (%) of the HB CTs between the future 2031–2060 and the reference period 1971–2000 for the entire year, the winter half-year (ONDJFM) and the summer half-year (AMJJAS).  Lower plots illustrate the spread of F1-scores.

FIGURE 1.5: Boxplots of frequency change and f1 score without transitional smoothing. Upper plots show the change in the relative frequency of occurrence (%) of the HB CTs between the future 2031–2060 and the reference period 1971–2000 for the entire year, the winter half-year (ONDJFM) and the summer half-year (AMJJAS). Lower plots illustrate the spread of F1-scores.

Boxplots of frequency change and f1 score after transitional smoothing. Upper plots show the change in the relative frequency of occurrence (%) of the HB CTs between the future 2031–2060 and the reference period 1971–2000 for the entire year, the winter half-year (ONDJFM) and the summer half-year (AMJJAS).  Lower plots illustrate the spread of F1-scores.

FIGURE 1.6: Boxplots of frequency change and f1 score after transitional smoothing. Upper plots show the change in the relative frequency of occurrence (%) of the HB CTs between the future 2031–2060 and the reference period 1971–2000 for the entire year, the winter half-year (ONDJFM) and the summer half-year (AMJJAS). Lower plots illustrate the spread of F1-scores.

Without transition smoothing (Figure 1.5):

  1. The boxplots show relatively smaller spreads for most HB CTs.
  2. Most relative changes are within ±50%.
  3. Fewer extreme outliers are observed, and the whiskers are relatively short.
  4. The overall signal is dominated by noise, making it hard to detect consistent climate change trends.

With transition smoothing (Figure 1.6):

  1. The boxplots show larger spreads, particularly for rare HB CTs.
  2. Relative changes frequently extend beyond ±50%, and even approach ±100% in some cases.
  3. More extreme outliers appear.
  4. F1-score distributions are lower and broader, indicating less stable classification performance.

Overall, transition smoothing, while theoretically beneficial, introduces higher internal variability in our implementation, complicating clear signal identification. Variability differences from the original study stem from methodological factors (ensemble size, smoothing implementation) and underlying climate model characteristics.

Our findings reaffirm the value of a multi-ensemble approach and careful consideration of methodological choices when projecting future circulation trends, underscoring the need for robust ensemble analysis techniques to better discern climate-driven circulation changes from internal variability.

The original study (Appendix Figure 1.13) shows relatively controlled variability, mostly within ±50% (Mittermeier et al. 2022), and fewer extreme outliers, indicating relatively stable and moderate shifts in HB CTs. Our results (Figure 1.5) exhibit larger variability, particularly notable in specific HB CTs such as NA, SEA, HNFZ, and SZ, with frequency changes frequently exceeding ±50%, even reaching ±100%.

  1. Entire Year:
    • Original study shows moderate shifts centered around zero, with clear signals for types like WA and SEA (Mittermeier et al. 2022).
    • Our implementation shows notably higher variability with clear signals (e.g., strong increase in HNZ, HM and significant decreases in HB and NEZ).
  2. Winter Half-Year:
    • Original study depicts clear moderate increases for WW and HFA with relatively smaller variability (Mittermeier et al. 2022).
    • Our data demonstrate significant changes with pronounced variability and stronger extremes, especially in HNZ and rare CTs like NA and SZ.
  3. Summer Half-Year:
    • Original data indicate mild variability and clear, moderate increases and decreases in specific types (e.g., decreases in HNFZ, increases in WA) (Mittermeier et al. 2022).
    • Our results again illustrate increased variability with substantial increases (e.g., HNZ, HM) and decreases (e.g., HB, NEZ).

In both cases, the summer half-year has a greater influence on the entire year, despite the winter half-year showing higher variability. Rare CTs such as NA and SZ are observed with significant changes. It would be interesting to analyze how rare conditions evolve, however, due to the unbalanced data, the results are not very reliable.

Key CT Differences:

  1. WA (Westerly Anticyclonic):
    • Our results show only slight increases, with less pronounced signals for WA, especially compared to significant changes in other types.
    • Original study consistently shows WA increasing clearly across all seasons (Mittermeier et al. 2022).
  2. HNZ (Icelandic High Cyclonic):
    • In our results, HNZ consistently shows marked increases, particularly pronounced in the summer and winter half-years.
    • Original study presents less dramatic decreasing changes for HNZ, reflecting a critical methodological or scenario-driven difference (Mittermeier et al. 2022).
  3. NA and SZ (Northerly Anticyclonic and Southerly Cyclonic):
    • Our results exhibit significantly larger variability and extreme reductions in NA, while SZ shows extreme negative changes particularly in summer.
    • Original study maintains relatively modest decreasing changes for these types (Mittermeier et al. 2022).

The differences in projected frequency changes between our results and the original study primarily stem from:

  1. Time period difference (2031-2060 vs. 2071-2100).
  2. Scenario differences: CanESM2 model (historical) in our study vs. original SMHI-LENS (SSP3-RCP7.0).
  3. Model and methodological distinctions: Implementation differences (e.g., transitional smoothing, dataset variations, and resolution).
  4. Ensemble size effects: Our smaller ensemble size (12 vs. original 50) could amplify internal variability signals.

These differences underline the sensitivity of CT analysis to methodological choices, highlighting the importance of transparent documentation of scenario assumptions and model details to ensure the accurate interpretation and comparability of future climate studies.

1.6 Conclusion and Outlook

In this study, we successfully implemented a CNN-based deep learning approach to classify large-scale CTs over Europe, building upon the established HB framework. Using high-resolution ERA5 data for model training, our method achieved an overall accuracy of 53.17% and a macro F1-score of 55.14%, clearly outperforming traditional classification approaches.

The CNN model effectively learned dominant circulation patterns, yet faced challenges classifying rare or inherently complex CTs. These limitations indicate potential issues of class imbalance and inherent labeling uncertainties associated with manual classifications. One potential improvement to the statistical analysis would be reducing the number of subdivisions of CTs. Subdividing CT according to the wind direction of the jet stream, for example, could increase the robustness of the results. Additionally, the original study found that the most frequent misclassifications occurred between pairs of anticyclonic and cyclonic circulation types (Mittermeier et al. 2022), which could be addressed through refined classification methods. Furthermore, future research should be conducted to quantify human-level errors in labelling to better understand the impacts of manual classification biases.

Besides the primary training session and results shown in this report, we conducted another model training experiment using a narrower range of hyperparameters. The results, which are illustrated in the presentation, demonstrated noticeably different classification performance compared to the setup we introduced here. In particular, I suspect that the size of the fully connected linear layers is the mean reason for this difference. Since in earlier experiments with narrower range of hyperparameters, we set the fully connected layer size as fixed 50 dimensions, however, in the latest optimization round, where more extensive hyperparameter tuning were applied, the model consistently favored a dimension size of 128, which leads to high accuracy for each HB CTs. Therefore, the model shows good performance on the current dataset, it is important to note that this does not necessarily guarantee the absence of overfitting.

Due to limited computational resources, we were unable to perform a more thorough investigation to validate these observations or further optimize the model architecture. Future work should focus on systematically assessing the impact of fully connected layer sizes and further regularization strategies to ensure model robustness and generalization.

The analysis of frequency changes in both our implementation and original paper reveals that for most HB CTs, the changes fall within a range of ±5 days (relatively ±50%) (Appendix Figure 1.12; Figure 1.6; Appendix Figure 1.14; Appendix Figure 1.13). In the original study, the most significant changes in frequency were observed in the WA circulation type (Appendix Figure 1.13) (Mittermeier et al. 2022), while in the own-built model, the HM and HNZ circulation types experienced the largest shifts in frequency. Notably, our results differ from previous studies, primarily due to distinct methodological choices such as dataset selection, scenario assumptions, and smoothing techniques. These variations underline the sensitivity of HB CTs analysis and stress the importance of transparent methodological reporting.

When focusing only on the frequency of HB CTs, without considering their duration, it is important to highlight that the duration of a HB CT significantly influences the amplitude of weather extremes. This underscores the need to integrate both frequency and duration for a more comprehensive understanding of weather patterns. Duration of a HB CT impacts the amplitude of a weather extreme.

To further evaluate the uncertainties in frequency changes, it would be beneficial to combine multi-model and single-model ensembles under different forcing scenarios. This approach could help assess the uncertainties arising from various climate models and the assumptions in different forcing scenarios.

For future studies, we recommend quantifying human-level errors in labeling to understand better the impacts of manual classification biases. Additionally, models that explicitly capture temporal continuity, such as Deep Hidden Markov Models (D. Yu et al. 2015) or temporal-aware CNN architectures (e.g., ConvLSTM (Moishin et al. 2021)), could enhance prediction accuracy by directly modeling the persistence characteristic inherent in circulation patterns (Mittermeier et al. 2022).

1.7 Appendix

1.7.1 Statement

The implementation was carried out using Python version 3.9.18.

1.7.2 Figures

Synoptic patterns of the 29 HB CTs using ERA20C (Figure 1.7) (Mittermeier et al. 2022)

Synoptic patterns of the 29 HB CTs created using ERA-20C data averaged over the period 1900 -1980, showing SLP on the left and z500 on the right side.

FIGURE 1.7: Synoptic patterns of the 29 HB CTs created using ERA-20C data averaged over the period 1900 -1980, showing SLP on the left and z500 on the right side.

HB CTs (Figure 1.8)

HB CTs

FIGURE 1.8: HB CTs

F1 score comparison table (Figure 1.9)

Macro F1

FIGURE 1.9: Macro F1

Signature plots of paper - SLP (Figure 1.10) (Mittermeier et al. 2022)

Signature plots of four selected circulation patterns at slp. Column 1: labels showing the indicated HB CT, Column 2: deep ensemble predictions showing the indicated HB CT, Column 3: signature pattern, when the deep ensemble predicts the indicated HB CT while labels state differently, Column 4: labels stating the indicated HB CT while deep ensemble predicts differently. The RMSE values are calculated by comparing the respective signature plot to the signature plot of the labels (column 1). The four HB CTs are chosen as positive (green) and negative (red) examples.

FIGURE 1.10: Signature plots of four selected circulation patterns at slp. Column 1: labels showing the indicated HB CT, Column 2: deep ensemble predictions showing the indicated HB CT, Column 3: signature pattern, when the deep ensemble predicts the indicated HB CT while labels state differently, Column 4: labels stating the indicated HB CT while deep ensemble predicts differently. The RMSE values are calculated by comparing the respective signature plot to the signature plot of the labels (column 1). The four HB CTs are chosen as positive (green) and negative (red) examples.

Frequency distribution of paper (Figure 1.11) (Mittermeier et al. 2022)

Frequency distribution of the 29 HB CTs in number of days per year for the training period 1900–1980 (blue) and for predictions of the CNN (yellow).

FIGURE 1.11: Frequency distribution of the 29 HB CTs in number of days per year for the training period 1900–1980 (blue) and for predictions of the CNN (yellow).

Boxplot of absolute future change and F1 score of our implementation (Figure 1.12)

Boxplots of frequency change and f1 score. Upper plots show the change in the absolute frequency of occurrence (days) of the HB CTs between the future 2031–2060 and the reference period 1971–2000 for the entire year, the winter half-year (ONDJFM) and the summer half-year (AMJJAS).  Lower plots illustrate the spread of F1-scores.

FIGURE 1.12: Boxplots of frequency change and f1 score. Upper plots show the change in the absolute frequency of occurrence (days) of the HB CTs between the future 2031–2060 and the reference period 1971–2000 for the entire year, the winter half-year (ONDJFM) and the summer half-year (AMJJAS). Lower plots illustrate the spread of F1-scores.

Boxplot of future change and F1 score of paper (Figure 1.13) (Mittermeier et al. 2022)

Boxplots of frequency change and f1 score. Upper plots show the change in the relative frequency of occurrence (%) of the HB CTs between the future 2031–2060 and the reference period 1971–2000 for the entire year, the winter half-year (ONDJFM) and the summer half-year (AMJJAS).  Lower plots illustrate the spread of F1-scores.

FIGURE 1.13: Boxplots of frequency change and f1 score. Upper plots show the change in the relative frequency of occurrence (%) of the HB CTs between the future 2031–2060 and the reference period 1971–2000 for the entire year, the winter half-year (ONDJFM) and the summer half-year (AMJJAS). Lower plots illustrate the spread of F1-scores.

Boxplot of absolute future change and F1 score of paper (Figure 1.14) (Mittermeier et al. 2022)

Boxplots of frequency change and f1 score. Upper plots show the change in the absolute relative frequency of occurrence (days) of the HB CTs between the future 2031–2060 and the reference period 1971–2000 for the entire year, the winter half-year (ONDJFM) and the summer half-year (AMJJAS).  Lower plots illustrate the spread of F1-scores.

FIGURE 1.14: Boxplots of frequency change and f1 score. Upper plots show the change in the absolute relative frequency of occurrence (days) of the HB CTs between the future 2031–2060 and the reference period 1971–2000 for the entire year, the winter half-year (ONDJFM) and the summer half-year (AMJJAS). Lower plots illustrate the spread of F1-scores.

RMSE table of our implementation - SLP (Figure 1.15)

RMSE - slp

FIGURE 1.15: RMSE - slp

RMSE table of paper - SLP (Figure 1.16)

RMSE - slp

FIGURE 1.16: RMSE - slp

Confusion matrix of our implementation (Figure 1.17)

Confusion matrix

FIGURE 1.17: Confusion matrix

Confusion matrix of paper (Figure 1.18) (Mittermeier et al. 2022)

Confusion matrix

FIGURE 1.18: Confusion matrix

Full signature plot of our implementation - SLP (Figure 1.19)

Signature plot

FIGURE 1.19: Signature plot

Full signature plot of paper - SLP (Figure 1.20) (Mittermeier et al. 2022)

Signature plot

FIGURE 1.20: Signature plot

References

Alberg, Anthony J, Ji Wan Park, Brant W Hager, Malcolm V Brock, and Marie Diener-West. 2004. “The Use of ‘Overall Accuracy’ to Evaluate the Validity of Screening or Diagnostic Tests.” Journal of General Internal Medicine 19 (5p1): 460–65.
Baur, F., P. Hess, and H. Nagel. 1944. Kalender Der Großwetterlagen Europas 1881–1939. Bad Homburg: Forschungsinstitut für langfristige Witterungsvorhersage. https://www.mdpi.com/2073-4441/12/6/1562.
Bissolli, Peter. 2001. “Wetterlagen Und Großwetterlagen Im 20. Jahrhundert.” DWD Klimastatusbericht. https://www.dwd.de/DE/leistungen/klimastatusbericht/publikationen/ksb2001_pdf/04_2001.pdf?__blob=publicationFile&v=1.
Bissolli, Peter, and Elke Dittmann. 2001. “The Objective Weather Type Classification of the German Weather Service and Its Possibilities of Application to Environmental and Meteorological Investigations.” Meteorologische Zeitschrift 10 (4): 253–60. https://doi.org/10.1127/0941-2948/2001/0010-0253.
Dietterich, Thomas G et al. 2002. “Ensemble Learning.” The Handbook of Brain Theory and Neural Networks 2 (1): 110–25.
European Centre for Medium-Range Weather Forecasts. 2017. ERA5: Data Documentation.” https://confluence.ecmwf.int/display/CKB/ERA5%3A+data+documentation.
Friedlingstein, Olivier Boucher, Mark Webb, Jonathan Gregory, et al. 2008. “A Summary of the CMIP5 Experiment Design.” https://pcmdi.llnl.gov/mips/cmip5/docs/Taylor_CMIP5_design.pdf?id=96.
Garnett, Roman. 2023. Bayesian Optimization. Cambridge University Press.
Government of Canada. 2025. “Canadian Climate Data and Scenarios - CanESM2 Predictions.” https://climate-scenarios.canada.ca/?page=pred-canesm2.
Häckel, Hans. 2021. Meteorologie. 9., vollständig überarbeitete und erweiterte Auflage. Stuttgart: Verlag Eugen Ulmer. https://katalog.ub.uni-heidelberg.de/titel/68020108.
Hess, Paul. 2005. “Katalog Der Grosswetterlagen Europas Nach Paul Hess Und Helmuth Brezowsky:(1881-2004).”
Hess, Paul, and Helmuth Brezowsky. 1952. Katalog Der Großwetterlagen Europas. Berichte Des Deutschen Wetterdienstes in Der US-Zone 33. Deutscher Wetterdienst. https://search.worldcat.org/title/Katalog-der-Grosswetterlagen-Europas/oclc/66852539.
Hua, Wenjian, Haishan Chen, Shanlei Sun, and Liming Zhou. 2015. “Assessing Climatic Impacts of Future Land Use and Land Cover Change Projected with the CanESM2 Model.” International Journal of Climatology 35 (12).
Mehta, Smit, Chirag Paunwala, and Bhaumik Vaidya. 2019. “CNN Based Traffic Sign Classification Using Adam Optimizer.” In 2019 International Conference on Intelligent Computing and Control Systems (ICCS), 1293–98. IEEE.
Mittermeier, Magdalena, Maximilian Weigert, David Rügamer, Helmut Küchenhoff, and Ralf Ludwig. 2022. “A Deep Learning Based Classification of Atmospheric Circulation Types over Europe: Projection of Future Changes in a CMIP6 Large Ensemble.” Environmental Research Letters 17 (8): 084021. https://doi.org/10.1088/1748-9326/ac8068.
Moishin, Mohammed, Ravinesh C Deo, Ramendra Prasad, Nawin Raj, and Shahab Abdulla. 2021. “Designing Deep-Based Learning Flood Forecast Model with ConvLSTM Hybrid Algorithm.” Ieee Access 9: 50982–93.
Mueller, Markus U, Nikoo Ekhtiari, Rodrigo M Almeida, and Christoph Rieke. 2020. “Super-Resolution of Multispectral Satellite Images Using Convolutional Neural Networks.” arXiv Preprint arXiv:2002.00580.
Opitz, Juri, and Sebastian Burst. 2019. “Macro F1 and Macro F1.” arXiv Preprint arXiv:1911.03347.
Ozaki, Yoshihiko, Yuki Tanigaki, Shuhei Watanabe, Masahiro Nomura, and Masaki Onishi. 2022. “Multiobjective Tree-Structured Parzen Estimator.” Journal of Artificial Intelligence Research 73: 1209–50.
Poli, Paul, Hans Hersbach, Dick Dee, John Siminski, Patrick Laloyaux, Debora Tan, Carole Peubey, et al. 2016. “ERA-20C: An Atmospheric Reanalysis of the Twentieth Century.” Journal of Climate 29 (11): 4083–97. https://doi.org/10.1175/JCLI-D-15-0556.1.
Ward, Rachel, Xiaoxia Wu, and Leon Bottou. 2020. “Adagrad Stepsizes: Sharp Convergence over Nonconvex Landscapes.” Journal of Machine Learning Research 21 (219): 1–30.
Werner, P. C., and F. W. Gerstengarbe. 2010. Katalog Der Großwetterlagen Europas (1881–2009) Nach Paul Hess Und Helmut Brezowsky. PIK Report 119. Potsdam-Institut für Klimafolgenforschung. https://www.pik-potsdam.de/en/output/publications/pikreports/.files/pr119.pdf.
Wyser, Klaus, Torben Koenigk, Ulrich Fladrich, Roberto Fuentes-Franco, Mohammad P. Karami, and Tim Kruschke. 2021. “The SMHI Large Ensemble (SMHI-LENS) with EC-Earth3.3.1.” Geoscientific Model Development 14 (9): 4781–96. https://doi.org/10.5194/gmd-14-4781-2021.
Yu, Dong, Li Deng, Dong Yu, and Li Deng. 2015. “Deep Neural Network-Hidden Markov Model Hybrid Systems.” Automatic Speech Recognition: A Deep Learning Approach, 99–116.
Zhong, Yi, Prabhakar Chalise, and Jianghua He. 2023. “Nested Cross-Validation with Ensemble Feature Selection and Classification Model for High-Dimensional Biological Data.” Communications in Statistics-Simulation and Computation 52 (1): 110–25.
Zou, Fangyu, Li Shen, Zequn Jie, Weizhong Zhang, and Wei Liu. 2019. “A Sufficient Condition for Convergences of Adam and Rmsprop.” In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 11127–35.