You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

NASA涡轮风扇数据集CNN-LSTM模型过拟合问题求助(附代码)

问题描述

基于NASA TurboFan FD001数据集构建的CNN-LSTM混合网络存在严重过拟合,测试集MAE始终维持在50-100区间。已尝试调整训练轮数、批次大小及模型架构,但效果不佳。

数据集维度:

  • 训练集:(100, 364, 14)(样本数,循环次数,特征数)
  • 测试集:(100, 364, 14)

注:故障周期较短的发动机样本已被补零至最大循环数364。


现有代码

训练数据预处理

import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
from sklearn.preprocessing import MinMaxScaler

training_data = pd.read_csv('../data/train1.csv')
testing_data = pd.read_csv('../data/test1.csv')
rul = pd.read_csv('../data/RUL1.csv')

# Drop columns with constant values
training_data_nconst = pd.read_csv('~/raytheon/data/train1.csv')
print(f'Previous Training Shape : {training_data_nconst.shape}') 

for column in training_data_nconst.columns:
    if training_data_nconst[column].std() < 0.01:
        training_data_nconst.drop(column, axis=1, inplace=True)
    
# Drop the usual suspects
training_data_nconst.drop(['Unit', 'Cycle', 'RUL'], axis=1, inplace=True)

print(f'New Training Shape : {training_data_nconst.shape}')
training_data_nconst.describe()

# from sklearn.preprocessing import MinMaxScaler
scaler = MinMaxScaler()
training_data_normalized = training_data.copy()
training_data_normalized = scaler.fit_transform(training_data_nconst)
print(training_data_normalized)

# Correcting column names based on the initial setup
feature_columns = column_names = ["T24", "T30", "T50", "P30", "Nf", "Nc", "Ps30", 
              "phi", "NRf", "NRc", "BPR", "htBleed", "W31", "W32"]
training_DataFrame_normalized = pd.DataFrame(training_data_normalized, columns=feature_columns)

# Add back the 'Unit #' and 'Time Cycle' columns from the original training data
training_DataFrame_normalized.insert(0, 'Cycle', training_data['Cycle'])
training_DataFrame_normalized.insert(0, 'Unit', training_data['Unit'])

training_DataFrame_normalized.insert(16, 'RUL', training_data['RUL'])

# Find the largest engine subset using max_cycle.
max_cycle_length = training_DataFrame_normalized['Cycle'].max()
print(f'Most cycles: {max_cycle_length}')

'''Break the training data into engine subsets.
Scale up other subsets. Just fill in 0's for new rows. (Thinking failure - 0 sensor output..?)'''
engine_data_subsets = {}

for unit in training_DataFrame_normalized['Unit'].unique():
    engine_subset = training_DataFrame_normalized.loc[training_DataFrame_normalized['Unit'] == unit]
    
    # Find the maximum cycle number for the current engine
    max_cycle_engine = engine_subset['Cycle'].max()

    # Calculate padding length for this engine
    padding_length = max_cycle_length - max_cycle_engine

    # Create a padding DataFrame with zeros for sensor columns (excluding 'Unit' and 'Cycle')
    if padding_length > 0:
        padding_df = pd.DataFrame(0, index=range(padding_length), columns=engine_subset.columns[2:])
        padding_df['Unit'] = unit
        padding_df['Cycle'] = range(max_cycle_engine + 1, max_cycle_engine + 1 + padding_length)
        
        # Concatenate the padding DataFrame to the end of the engine_subset
        engine_subset = pd.concat([engine_subset, padding_df], ignore_index=True)
        
    engine_data_subsets[unit] = engine_subset
print(engine_data_subsets[100])


engine_1_data = engine_data_subsets[1]
sensor_columns = feature_columns

测试数据预处理

# Drop columns with constant values
test_data_nconst = pd.read_csv('~/raytheon/data/test1.csv')
print(f'Previous Testing Shape : {test_data_nconst.shape}') 

for column in test_data_nconst.columns:
    if test_data_nconst[column].std() < 0.01:
        test_data_nconst.drop(column, axis=1, inplace=True)
        
# Drop the usual suspects
test_data_nconst.drop(['Unit', 'Cycle'], axis=1, inplace=True)
        
print(f'New Training Shape : {test_data_nconst.shape}')
test_data_nconst.describe()

# from sklearn.preprocessing import MinMaxScaler
scaler = MinMaxScaler()
test_data_normalized = testing_data.copy()
test_data_normalized = scaler.fit_transform(test_data_nconst)
print(test_data_normalized)

test_DataFrame_normalized = pd.DataFrame(test_data_normalized, columns=feature_columns)

# Add back the 'Unit #' and 'Time Cycle' columns from the original training data
test_DataFrame_normalized.insert(0, 'Cycle', testing_data['Cycle'])
test_DataFrame_normalized.insert(0, 'Unit', testing_data['Unit'])

# Find the largest engine subset using max_cycle.
max_test_cycle_length = test_DataFrame_normalized.groupby('Unit')['Cycle'].max().max()
print(f'Most cycles: {max_test_cycle_length}')

test_engine_data_subsets = {}

for unit in test_DataFrame_normalized['Unit'].unique():
    test_engine_subset = test_DataFrame_normalized.loc[test_DataFrame_normalized['Unit'] == unit]
    
    # Find the maximum cycle number for the current engine
    max_test_cycle_engine = test_engine_subset['Cycle'].max()

    # Calculate padding length for this engine
    padding_length = max_test_cycle_length - max_test_cycle_engine

    # Create a padding DataFrame with zeros for sensor columns (excluding 'Unit' and 'Cycle')
    if padding_length > 0:
        padding_df = pd.DataFrame(0, index=range(padding_length), columns=test_engine_subset.columns[2:])
        padding_df['Unit'] = unit
        padding_df['Cycle'] = range(max_test_cycle_engine + 1, max_test_cycle_engine + 1 + padding_length)
        
        # Concatenate the padding DataFrame to the end of the engine_subset
        test_engine_subset = pd.concat([test_engine_subset, padding_df], ignore_index=True)
        
    test_engine_data_subsets[unit] = test_engine_subset
test_engine_data_subsets[100]

# Get the data for engine 1
test_engine_1_data = test_engine_data_subsets[1]
test_sensor_columns = feature_columns

模型输入与定义

import numpy as np

# Each value in engine_data_subsets is a numpy array of shape (cycle_count, num_features)
batch_num = len(engine_data_subsets)

# Assuming all arrays have the same shape, get the shape of the first array
example_array = next(iter(engine_data_subsets.values()))
cycle_count, num_features = example_array.shape

X_train = np.zeros((batch_num, cycle_count, num_features))

for i, array in enumerate(engine_data_subsets.values()):
    X_train[i, :, :] = array

y_train = training_DataFrame_normalized.groupby('Unit')['RUL'].max()

# Remove non-features from INPUT
X_train = X_train[:,:,2:16]
y_train = np.reshape(y_train, (100, 1))
print('X Train set shape', X_train.shape)
print('y Train set shape', y_train.shape)

from tensorflow.keras.models import Sequential
from tensorflow.keras.layers import Conv1D, MaxPooling1D, Flatten, Dense, Dropout, LSTM
from tensorflow.keras import regularizers
from tensorflow.keras.optimizers import Adam
from tensorflow.keras.callbacks import ReduceLROnPlateau

# Define the model
model = Sequential()

# Convolutional layers
model.add(Conv1D(filters=32, kernel_size=3, activation='relu', input_shape=(X_train.shape[1], X_train.shape[2])))
model.add(MaxPooling1D(pool_size=2))
model.add(LSTM(32, return_sequences=True))
model.add(Conv1D(filters=62, kernel_size=3, activation='relu'))
model.add(MaxPooling1D(pool_size=2))

# Flattening layer
model.add(Flatten())

# Fully connected layers
model.add(Dense(128, activation='relu'))
model.add(Dense(64, activation='relu'))

# Output layer
model.add(Dense(1))

# Compile the model
optimizer = Adam(learning_rate=0.0001)
model.compile(optimizer=optimizer, loss='mean_squared_error', metrics=['mae'])

# Print model summary
model.summary()

# Train the model
history = model.fit(X_train, y_train, epochs=40, batch_size=32, validation_split=0.2)
predictions = model.predict(X_train)
predictions.shape

from sklearn.metrics import mean_absolute_error, mean_squared_error, r2_score
# Calculate Mean Absolute Error (MAE)
mae = mean_absolute_error(y_train, predictions)
print("Mean Absolute Error (MAE):", mae)

# Calculate Mean Squared Error (MSE)
mse = mean_squared_error(y_train, predictions)
print("Mean Squared Error (MSE):", mse)

result_train = pd.DataFrame(y_train)
result_train['Predictions'] = predictions.round(0)
result_train

# Each value in engine_data_subsets is a numpy array of shape (cycle_count, num_features)
batch_num = len(test_engine_data_subsets)

# Assuming all arrays have the same shape, get the shape of the first array
example_array = next(iter(test_engine_data_subsets.values()))
cycle_count, num_features = example_array.shape

X_test = np.zeros((batch_num, cycle_count, num_features))

for i, array in enumerate(test_engine_data_subsets.values()):
    X_test[i, :, :] = array

y_test = rul['RUL']
#y_train = training_DataFrame_normalized.groupby('Unit')['RUL'].max()
print(f' INPUT Shape: {X_test.shape}')
print(f' OUTPUT Shape: {rul.shape}')

# Remove non-features from INPUT
X_test = X_test[:,:,2:16]
print(X_test.shape)

max_length = 362  # Maximum length of sequences
X_test_padded = np.zeros((X_test.shape[0], max_length, X_test.shape[2]))
for i in range(X_test.shape[0]):
    X_test_padded[i, :X_test.shape[1], :] = X_test[i, :max_length, :]

print(X_test_padded.shape)
predictions_test = model.predict(X_test_padded)

# Calculate Mean Absolute Error (MAE)
mae = mean_absolute_error(y_test, predictions_test)
print("Mean Absolute Error (MAE):", mae)

# Calculate Mean Squared Error (MSE)
mse = mean_squared_error(y_test, predictions_test)
print("Mean Squared Error (MSE):", mse)

result_test = pd.DataFrame(y_test)
result_test['Predictions'] = predictions_test.round(0)
result_test

预测结果
训练曲线


改进方案

1. 修复数据预处理核心问题

  • 测试集归一化错误:当前测试集用scaler.fit_transform会导致训练/测试集分布不一致,必须用训练集拟合的scaler转换测试集:
    # 训练集部分
    scaler = MinMaxScaler()
    training_data_normalized = scaler.fit_transform(training_data_nconst)
    
    # 测试集部分替换原fit_transform
    test_data_normalized = scaler.transform(test_data_nconst)
    
  • 补零策略优化:末尾补零会引入无效噪声,建议改用每个发动机的最后N个循环序列(比如最后100个循环)作为输入,或用该发动机传感器均值填充空缺,而非零值。

2. 调整训练数据划分逻辑

validation_split=0.2随机拆分样本会破坏发动机时序相关性,应按发动机Unit划分验证集,避免数据泄露:

# 按Unit划分训练/验证集
train_units = training_DataFrame_normalized['Unit'].unique()[:80]
val_units = training_DataFrame_normalized['Unit'].unique()[80:]

X_train = np.array([engine_data_subsets[unit].iloc[:,2:16].values for unit in train_units])
y_train = training_DataFrame_normalized[training_DataFrame_normalized['Unit'].isin(train_units)].groupby('Unit')['RUL'].max().values.reshape(-1,1)

X_val = np.array([engine_data_subsets[unit].iloc[:,2:16].values for unit in val_units])
y_val = training_DataFrame_normalized[training_DataFrame_normalized['Unit'].isin(val_units)].groupby('Unit')['RUL'].max().values.reshape(-1,1)

# 训练时改用validation_data
history = model.fit(X_train, y_train, epochs=40, batch_size=32, validation_data=(X_val, y_val))

3. 加入正则化抑制过拟合

利用已导入的regularizers,给模型层添加L2正则和Dropout:

model = Sequential()
# 卷积层加入L2正则
model.add(Conv1D(filters=32, kernel_size=3, activation='relu', 
                 input_shape=(X_train.shape[1], X_train.shape[2]),
                 kernel_regularizer=regularizers.l2(0.001)))
model.add(MaxPooling1D(pool_size=2))
model.add(LSTM(32, return_sequences=True, kernel_regularizer=regularizers.l2(0.001)))
model.add(Dropout(0.2))  # 加入Dropout
model.add(Conv1D(filters=64, kernel_size=3, activation='relu', kernel_regularizer=regularizers.l2(0.001)))
model.add(MaxPooling1D(pool_size=2))
model.add(Dropout(0.2))

model.add(Flatten())
model.add(Dense(128, activation='relu', kernel_regularizer=regularizers.l2(0.001)))
model.add(Dropout(0.3))
model.add(Dense(64, activation='relu', kernel_regularizer=regularizers.l2(0.001)))
model.add(Dropout(0.2))
model.add(Dense(1))

4. 优化训练策略

  • 早停机制:加入EarlyStopping,当验证集MAE不再下降时停止训练,避免过拟合:
    from tensorflow.keras.callbacks import EarlyStopping
    early_stop = EarlyStopping(monitor='val_mae', patience=5, restore_best_weights=True)
    history = model.fit(X_train, y_train, epochs=100, batch_size=32, 
                        validation_data=(X_val, y_val), callbacks=[early_stop, ReduceLROnPlateau()])
    
  • 自定义损失函数:RUL预测中短寿命样本误差影响更大,给小RUL样本更高权重:
    def weighted_mse(y_true, y_pred):
        # 对RUL小于50的样本权重设为2,其余为1
        weights = tf.where(y_true < 50, 2.0, 1.0)
        return tf.reduce_mean(weights * tf.square(y_true - y_pred))
    
    model.compile(optimizer=optimizer, loss=weighted_mse, metrics=['mae'])
    

5. 特征工程增强

  • 提取时序特征:对每个发动机的传感器数据计算斜率、滚动均值、滚动方差等趋势特征,补充到原始特征中,提升模型对退化趋势的捕捉能力。
  • 筛选关键特征:用相关性分析或特征重要性排序,去掉与RUL无关的特征,减少模型学习负担。

6. 简化模型结构

当前CNN+LSTM组合参数过多,尝试简化:

  • 减少卷积层filters数量,比如把62改为64或32。
  • 去掉一层全连接层,仅保留Dense(64),降低模型容量。

内容的提问来源于stack exchange,提问作者Luke Auslender

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.25 16:45:58