NASA涡轮风扇数据集CNN-LSTM模型过拟合问题求助(附代码)
问题描述
基于NASA TurboFan FD001数据集构建的CNN-LSTM混合网络存在严重过拟合,测试集MAE始终维持在50-100区间。已尝试调整训练轮数、批次大小及模型架构,但效果不佳。
数据集维度:
- 训练集:(100, 364, 14)(样本数,循环次数,特征数)
- 测试集:(100, 364, 14)
注:故障周期较短的发动机样本已被补零至最大循环数364。
现有代码
训练数据预处理
import pandas as pd import numpy as np import matplotlib.pyplot as plt from sklearn.preprocessing import MinMaxScaler training_data = pd.read_csv('../data/train1.csv') testing_data = pd.read_csv('../data/test1.csv') rul = pd.read_csv('../data/RUL1.csv') # Drop columns with constant values training_data_nconst = pd.read_csv('~/raytheon/data/train1.csv') print(f'Previous Training Shape : {training_data_nconst.shape}') for column in training_data_nconst.columns: if training_data_nconst[column].std() < 0.01: training_data_nconst.drop(column, axis=1, inplace=True) # Drop the usual suspects training_data_nconst.drop(['Unit', 'Cycle', 'RUL'], axis=1, inplace=True) print(f'New Training Shape : {training_data_nconst.shape}') training_data_nconst.describe() # from sklearn.preprocessing import MinMaxScaler scaler = MinMaxScaler() training_data_normalized = training_data.copy() training_data_normalized = scaler.fit_transform(training_data_nconst) print(training_data_normalized) # Correcting column names based on the initial setup feature_columns = column_names = ["T24", "T30", "T50", "P30", "Nf", "Nc", "Ps30", "phi", "NRf", "NRc", "BPR", "htBleed", "W31", "W32"] training_DataFrame_normalized = pd.DataFrame(training_data_normalized, columns=feature_columns) # Add back the 'Unit #' and 'Time Cycle' columns from the original training data training_DataFrame_normalized.insert(0, 'Cycle', training_data['Cycle']) training_DataFrame_normalized.insert(0, 'Unit', training_data['Unit']) training_DataFrame_normalized.insert(16, 'RUL', training_data['RUL']) # Find the largest engine subset using max_cycle. max_cycle_length = training_DataFrame_normalized['Cycle'].max() print(f'Most cycles: {max_cycle_length}') '''Break the training data into engine subsets. Scale up other subsets. Just fill in 0's for new rows. (Thinking failure - 0 sensor output..?)''' engine_data_subsets = {} for unit in training_DataFrame_normalized['Unit'].unique(): engine_subset = training_DataFrame_normalized.loc[training_DataFrame_normalized['Unit'] == unit] # Find the maximum cycle number for the current engine max_cycle_engine = engine_subset['Cycle'].max() # Calculate padding length for this engine padding_length = max_cycle_length - max_cycle_engine # Create a padding DataFrame with zeros for sensor columns (excluding 'Unit' and 'Cycle') if padding_length > 0: padding_df = pd.DataFrame(0, index=range(padding_length), columns=engine_subset.columns[2:]) padding_df['Unit'] = unit padding_df['Cycle'] = range(max_cycle_engine + 1, max_cycle_engine + 1 + padding_length) # Concatenate the padding DataFrame to the end of the engine_subset engine_subset = pd.concat([engine_subset, padding_df], ignore_index=True) engine_data_subsets[unit] = engine_subset print(engine_data_subsets[100]) engine_1_data = engine_data_subsets[1] sensor_columns = feature_columns
测试数据预处理
# Drop columns with constant values test_data_nconst = pd.read_csv('~/raytheon/data/test1.csv') print(f'Previous Testing Shape : {test_data_nconst.shape}') for column in test_data_nconst.columns: if test_data_nconst[column].std() < 0.01: test_data_nconst.drop(column, axis=1, inplace=True) # Drop the usual suspects test_data_nconst.drop(['Unit', 'Cycle'], axis=1, inplace=True) print(f'New Training Shape : {test_data_nconst.shape}') test_data_nconst.describe() # from sklearn.preprocessing import MinMaxScaler scaler = MinMaxScaler() test_data_normalized = testing_data.copy() test_data_normalized = scaler.fit_transform(test_data_nconst) print(test_data_normalized) test_DataFrame_normalized = pd.DataFrame(test_data_normalized, columns=feature_columns) # Add back the 'Unit #' and 'Time Cycle' columns from the original training data test_DataFrame_normalized.insert(0, 'Cycle', testing_data['Cycle']) test_DataFrame_normalized.insert(0, 'Unit', testing_data['Unit']) # Find the largest engine subset using max_cycle. max_test_cycle_length = test_DataFrame_normalized.groupby('Unit')['Cycle'].max().max() print(f'Most cycles: {max_test_cycle_length}') test_engine_data_subsets = {} for unit in test_DataFrame_normalized['Unit'].unique(): test_engine_subset = test_DataFrame_normalized.loc[test_DataFrame_normalized['Unit'] == unit] # Find the maximum cycle number for the current engine max_test_cycle_engine = test_engine_subset['Cycle'].max() # Calculate padding length for this engine padding_length = max_test_cycle_length - max_test_cycle_engine # Create a padding DataFrame with zeros for sensor columns (excluding 'Unit' and 'Cycle') if padding_length > 0: padding_df = pd.DataFrame(0, index=range(padding_length), columns=test_engine_subset.columns[2:]) padding_df['Unit'] = unit padding_df['Cycle'] = range(max_test_cycle_engine + 1, max_test_cycle_engine + 1 + padding_length) # Concatenate the padding DataFrame to the end of the engine_subset test_engine_subset = pd.concat([test_engine_subset, padding_df], ignore_index=True) test_engine_data_subsets[unit] = test_engine_subset test_engine_data_subsets[100] # Get the data for engine 1 test_engine_1_data = test_engine_data_subsets[1] test_sensor_columns = feature_columns
模型输入与定义
import numpy as np # Each value in engine_data_subsets is a numpy array of shape (cycle_count, num_features) batch_num = len(engine_data_subsets) # Assuming all arrays have the same shape, get the shape of the first array example_array = next(iter(engine_data_subsets.values())) cycle_count, num_features = example_array.shape X_train = np.zeros((batch_num, cycle_count, num_features)) for i, array in enumerate(engine_data_subsets.values()): X_train[i, :, :] = array y_train = training_DataFrame_normalized.groupby('Unit')['RUL'].max() # Remove non-features from INPUT X_train = X_train[:,:,2:16] y_train = np.reshape(y_train, (100, 1)) print('X Train set shape', X_train.shape) print('y Train set shape', y_train.shape) from tensorflow.keras.models import Sequential from tensorflow.keras.layers import Conv1D, MaxPooling1D, Flatten, Dense, Dropout, LSTM from tensorflow.keras import regularizers from tensorflow.keras.optimizers import Adam from tensorflow.keras.callbacks import ReduceLROnPlateau # Define the model model = Sequential() # Convolutional layers model.add(Conv1D(filters=32, kernel_size=3, activation='relu', input_shape=(X_train.shape[1], X_train.shape[2]))) model.add(MaxPooling1D(pool_size=2)) model.add(LSTM(32, return_sequences=True)) model.add(Conv1D(filters=62, kernel_size=3, activation='relu')) model.add(MaxPooling1D(pool_size=2)) # Flattening layer model.add(Flatten()) # Fully connected layers model.add(Dense(128, activation='relu')) model.add(Dense(64, activation='relu')) # Output layer model.add(Dense(1)) # Compile the model optimizer = Adam(learning_rate=0.0001) model.compile(optimizer=optimizer, loss='mean_squared_error', metrics=['mae']) # Print model summary model.summary() # Train the model history = model.fit(X_train, y_train, epochs=40, batch_size=32, validation_split=0.2) predictions = model.predict(X_train) predictions.shape from sklearn.metrics import mean_absolute_error, mean_squared_error, r2_score # Calculate Mean Absolute Error (MAE) mae = mean_absolute_error(y_train, predictions) print("Mean Absolute Error (MAE):", mae) # Calculate Mean Squared Error (MSE) mse = mean_squared_error(y_train, predictions) print("Mean Squared Error (MSE):", mse) result_train = pd.DataFrame(y_train) result_train['Predictions'] = predictions.round(0) result_train # Each value in engine_data_subsets is a numpy array of shape (cycle_count, num_features) batch_num = len(test_engine_data_subsets) # Assuming all arrays have the same shape, get the shape of the first array example_array = next(iter(test_engine_data_subsets.values())) cycle_count, num_features = example_array.shape X_test = np.zeros((batch_num, cycle_count, num_features)) for i, array in enumerate(test_engine_data_subsets.values()): X_test[i, :, :] = array y_test = rul['RUL'] #y_train = training_DataFrame_normalized.groupby('Unit')['RUL'].max() print(f' INPUT Shape: {X_test.shape}') print(f' OUTPUT Shape: {rul.shape}') # Remove non-features from INPUT X_test = X_test[:,:,2:16] print(X_test.shape) max_length = 362 # Maximum length of sequences X_test_padded = np.zeros((X_test.shape[0], max_length, X_test.shape[2])) for i in range(X_test.shape[0]): X_test_padded[i, :X_test.shape[1], :] = X_test[i, :max_length, :] print(X_test_padded.shape) predictions_test = model.predict(X_test_padded) # Calculate Mean Absolute Error (MAE) mae = mean_absolute_error(y_test, predictions_test) print("Mean Absolute Error (MAE):", mae) # Calculate Mean Squared Error (MSE) mse = mean_squared_error(y_test, predictions_test) print("Mean Squared Error (MSE):", mse) result_test = pd.DataFrame(y_test) result_test['Predictions'] = predictions_test.round(0) result_test


改进方案
1. 修复数据预处理核心问题
- 测试集归一化错误:当前测试集用
scaler.fit_transform会导致训练/测试集分布不一致,必须用训练集拟合的scaler转换测试集:# 训练集部分 scaler = MinMaxScaler() training_data_normalized = scaler.fit_transform(training_data_nconst) # 测试集部分替换原fit_transform test_data_normalized = scaler.transform(test_data_nconst) - 补零策略优化:末尾补零会引入无效噪声,建议改用每个发动机的最后N个循环序列(比如最后100个循环)作为输入,或用该发动机传感器均值填充空缺,而非零值。
2. 调整训练数据划分逻辑
validation_split=0.2随机拆分样本会破坏发动机时序相关性,应按发动机Unit划分验证集,避免数据泄露:
# 按Unit划分训练/验证集 train_units = training_DataFrame_normalized['Unit'].unique()[:80] val_units = training_DataFrame_normalized['Unit'].unique()[80:] X_train = np.array([engine_data_subsets[unit].iloc[:,2:16].values for unit in train_units]) y_train = training_DataFrame_normalized[training_DataFrame_normalized['Unit'].isin(train_units)].groupby('Unit')['RUL'].max().values.reshape(-1,1) X_val = np.array([engine_data_subsets[unit].iloc[:,2:16].values for unit in val_units]) y_val = training_DataFrame_normalized[training_DataFrame_normalized['Unit'].isin(val_units)].groupby('Unit')['RUL'].max().values.reshape(-1,1) # 训练时改用validation_data history = model.fit(X_train, y_train, epochs=40, batch_size=32, validation_data=(X_val, y_val))
3. 加入正则化抑制过拟合
利用已导入的regularizers,给模型层添加L2正则和Dropout:
model = Sequential() # 卷积层加入L2正则 model.add(Conv1D(filters=32, kernel_size=3, activation='relu', input_shape=(X_train.shape[1], X_train.shape[2]), kernel_regularizer=regularizers.l2(0.001))) model.add(MaxPooling1D(pool_size=2)) model.add(LSTM(32, return_sequences=True, kernel_regularizer=regularizers.l2(0.001))) model.add(Dropout(0.2)) # 加入Dropout model.add(Conv1D(filters=64, kernel_size=3, activation='relu', kernel_regularizer=regularizers.l2(0.001))) model.add(MaxPooling1D(pool_size=2)) model.add(Dropout(0.2)) model.add(Flatten()) model.add(Dense(128, activation='relu', kernel_regularizer=regularizers.l2(0.001))) model.add(Dropout(0.3)) model.add(Dense(64, activation='relu', kernel_regularizer=regularizers.l2(0.001))) model.add(Dropout(0.2)) model.add(Dense(1))
4. 优化训练策略
- 早停机制:加入EarlyStopping,当验证集MAE不再下降时停止训练,避免过拟合:
from tensorflow.keras.callbacks import EarlyStopping early_stop = EarlyStopping(monitor='val_mae', patience=5, restore_best_weights=True) history = model.fit(X_train, y_train, epochs=100, batch_size=32, validation_data=(X_val, y_val), callbacks=[early_stop, ReduceLROnPlateau()]) - 自定义损失函数:RUL预测中短寿命样本误差影响更大,给小RUL样本更高权重:
def weighted_mse(y_true, y_pred): # 对RUL小于50的样本权重设为2,其余为1 weights = tf.where(y_true < 50, 2.0, 1.0) return tf.reduce_mean(weights * tf.square(y_true - y_pred)) model.compile(optimizer=optimizer, loss=weighted_mse, metrics=['mae'])
5. 特征工程增强
- 提取时序特征:对每个发动机的传感器数据计算斜率、滚动均值、滚动方差等趋势特征,补充到原始特征中,提升模型对退化趋势的捕捉能力。
- 筛选关键特征:用相关性分析或特征重要性排序,去掉与RUL无关的特征,减少模型学习负担。
6. 简化模型结构
当前CNN+LSTM组合参数过多,尝试简化:
- 减少卷积层filters数量,比如把62改为64或32。
- 去掉一层全连接层,仅保留
Dense(64),降低模型容量。
内容的提问来源于stack exchange,提问作者Luke Auslender
相关产品推荐
相关产品推荐

