时序多分类LSTM网络精度不足,求优化方案
电力系统时序故障分类LSTM优化方案
问题概述
开发了用于电力系统时序多分类的LSTM网络,数据包含15个特征(连接母线的5条输电线路三相电流),6类标签(0为正常状态,1-5对应不同线路短路故障),标签已做独热编码。当前模型最高精度仅约75%,使用MinMaxScaler会损坏数据,StandardScaler无显著效果,需要优化。
核心优化建议
- 修复时序序列构造:当前
sequence_length=1完全浪费了LSTM的时序建模能力,需根据电力故障的时间特性设置合理序列长度(如10-50个连续时间步),让模型学习故障前后的电流变化趋势。 - 改进标准化策略:改用对异常值鲁棒的
RobustScaler,且严格遵循训练集拟合、测试集转换的原则,避免数据泄露。 - 优化模型结构:替换Recurrent Dropout为普通Dropout;引入双向LSTM捕捉双向时序特征;添加Early Stopping防止过拟合;调整神经元数量适配任务复杂度。
- 处理数据不平衡:检查样本类别分布,若存在不平衡,在损失函数中加入类别权重,或采用过采样/欠采样调整分布。
- 动态训练策略:使用学习率衰减(ReduceLROnPlateau)、增加训练轮数,监控验证集指标避免过拟合。
改进后的完整代码
import os import glob import pandas as pd import numpy as np import tensorflow as tf from sklearn.model_selection import train_test_split from sklearn.preprocessing import RobustScaler from sklearn.utils.class_weight import compute_class_weight from matplotlib import pyplot from keras.layers import Input, Dropout, Dense, LSTM, Bidirectional, BatchNormalization from keras.models import Sequential from keras.utils import to_categorical from keras.callbacks import EarlyStopping, ReduceLROnPlateau # 加载数据 def load_data(train_dir, test_dir): def load_files(file_paths): data_list = [] labels_list = [] for file_path in file_paths: df = pd.read_excel(file_path, dtype=np.float64) # 提取特征和标签(根据实际列索引调整) data = df.iloc[:, 1:-3].values labels = df.iloc[:, -2].values data_list.append(data) labels_list.append(labels) return np.concatenate(data_list, axis=0), np.concatenate(labels_list, axis=0) X_train, y_train = load_files(glob.glob(train_dir)) X_test, y_test = load_files(glob.glob(test_dir)) # 独热编码标签 y_train = to_categorical(y_train) y_test = to_categorical(y_test) return X_train, y_train, X_test, y_test # 构造时序序列 def create_sequences(data, labels, seq_length): X, y = [], [] for i in range(len(data) - seq_length + 1): X.append(data[i:i+seq_length]) # 取序列最后一个时间步的标签作为样本标签 y.append(labels[i+seq_length-1]) return np.array(X), np.array(y) # 评估模型 def evaluate_model(X_train, y_train, X_test, y_test, class_weights): verbose, epochs, batch_size = 1, 50, 128 n_timesteps, n_features, n_outputs = X_train.shape[1], X_train.shape[2], y_train.shape[1] model = Sequential() # 双向LSTM捕捉双向时序特征 model.add(Bidirectional(LSTM(128, return_sequences=True), input_shape=(n_timesteps, n_features))) model.add(BatchNormalization()) model.add(Dropout(0.3)) model.add(LSTM(64, return_sequences=False)) model.add(BatchNormalization()) model.add(Dropout(0.3)) model.add(Dense(128, activation='relu')) model.add(Dense(n_outputs, activation='softmax')) model.compile(loss='categorical_crossentropy', optimizer='adam', metrics=['accuracy']) # 回调函数:早停+学习率衰减 callbacks = [ EarlyStopping(patience=10, restore_best_weights=True), ReduceLROnPlateau(factor=0.5, patience=5, min_lr=1e-6) ] # 训练模型,加入类别权重 model.fit(X_train, y_train, epochs=epochs, batch_size=batch_size, verbose=verbose, validation_split=0.1, callbacks=callbacks, class_weight=class_weights) # 评估 _, accuracy = model.evaluate(X_test, y_test, batch_size=batch_size, verbose=1) return accuracy # 主函数 def main(): # 数据路径 train_dir = '/Users/hosseinebrahimi/Documents/ThesisResults/Normal/NEW/*.xlsx' test_dir = '/Users/hosseinebrahimi/Documents/ThesisResults/Normal/NEW/test/*.xlsx' # 加载数据 X_train, y_train, X_test, y_test = load_data(train_dir, test_dir) print("原始数据形状:", X_train.shape, y_train.shape, X_test.shape, y_test.shape) # 标准化:用RobustScaler按特征维度处理,避免异常值影响 scaler = RobustScaler() X_train_scaled = scaler.fit_transform(X_train) X_test_scaled = scaler.transform(X_test) # 构造时序序列(设置序列长度为20,可根据实际调整) seq_length = 20 X_train_seq, y_train_seq = create_sequences(X_train_scaled, y_train, seq_length) X_test_seq, y_test_seq = create_sequences(X_test_scaled, y_test, seq_length) print("序列数据形状:", X_train_seq.shape, y_train_seq.shape, X_test_seq.shape, y_test_seq.shape) # 计算类别权重,处理不平衡问题 y_train_classes = np.argmax(y_train_seq, axis=1) class_weights = compute_class_weight(class_weight='balanced', classes=np.unique(y_train_classes), y=y_train_classes) class_weights = dict(enumerate(class_weights)) # 运行实验 repeats = 3 scores = [] for r in range(repeats): score = evaluate_model(X_train_seq, y_train_seq, X_test_seq, y_test_seq, class_weights) score = score * 100.0 print(f"#{r+1}: %.3f%%" % score) scores.append(score) # 汇总结果 print("测试精度:", scores) print(f"平均精度: %.3f%% (+/-%.3f)" % (np.mean(scores), np.std(scores))) if __name__ == "__main__": main()
关键改进说明
- 时序序列构造:新增
create_sequences函数,将连续时间步组合成序列,让LSTM能学习故障的时序演化特征。 - 标准化优化:改用
RobustScaler,并严格用训练集拟合后转换测试集,避免数据泄露。 - 模型结构:加入双向LSTM,替换Recurrent Dropout为普通Dropout,增强特征捕捉能力同时避免过度抑制。
- 训练策略:添加Early Stopping防止过拟合,ReduceLROnPlateau动态调整学习率,加入类别权重处理数据不平衡。
- 代码模块化:将数据加载、序列构造等功能封装成函数,提升可读性和可维护性。
内容的提问来源于stack exchange,提问作者Hossein Ebrahimi
相关产品推荐
相关产品推荐

