You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

时序多分类LSTM网络精度不足,求优化方案

电力系统时序故障分类LSTM优化方案

问题概述

开发了用于电力系统时序多分类的LSTM网络,数据包含15个特征(连接母线的5条输电线路三相电流),6类标签(0为正常状态,1-5对应不同线路短路故障),标签已做独热编码。当前模型最高精度仅约75%,使用MinMaxScaler会损坏数据,StandardScaler无显著效果,需要优化。

核心优化建议

  • 修复时序序列构造:当前sequence_length=1完全浪费了LSTM的时序建模能力,需根据电力故障的时间特性设置合理序列长度(如10-50个连续时间步),让模型学习故障前后的电流变化趋势。
  • 改进标准化策略:改用对异常值鲁棒的RobustScaler,且严格遵循训练集拟合、测试集转换的原则,避免数据泄露。
  • 优化模型结构:替换Recurrent Dropout为普通Dropout;引入双向LSTM捕捉双向时序特征;添加Early Stopping防止过拟合;调整神经元数量适配任务复杂度。
  • 处理数据不平衡:检查样本类别分布,若存在不平衡,在损失函数中加入类别权重,或采用过采样/欠采样调整分布。
  • 动态训练策略:使用学习率衰减(ReduceLROnPlateau)、增加训练轮数,监控验证集指标避免过拟合。

改进后的完整代码

import os
import glob
import pandas as pd
import numpy as np
import tensorflow as tf
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import RobustScaler
from sklearn.utils.class_weight import compute_class_weight
from matplotlib import pyplot
from keras.layers import Input, Dropout, Dense, LSTM, Bidirectional, BatchNormalization
from keras.models import Sequential
from keras.utils import to_categorical
from keras.callbacks import EarlyStopping, ReduceLROnPlateau

# 加载数据
def load_data(train_dir, test_dir):
    def load_files(file_paths):
        data_list = []
        labels_list = []
        for file_path in file_paths:
            df = pd.read_excel(file_path, dtype=np.float64)
            # 提取特征和标签(根据实际列索引调整)
            data = df.iloc[:, 1:-3].values
            labels = df.iloc[:, -2].values
            data_list.append(data)
            labels_list.append(labels)
        return np.concatenate(data_list, axis=0), np.concatenate(labels_list, axis=0)

    X_train, y_train = load_files(glob.glob(train_dir))
    X_test, y_test = load_files(glob.glob(test_dir))
    
    # 独热编码标签
    y_train = to_categorical(y_train)
    y_test = to_categorical(y_test)
    return X_train, y_train, X_test, y_test

# 构造时序序列
def create_sequences(data, labels, seq_length):
    X, y = [], []
    for i in range(len(data) - seq_length + 1):
        X.append(data[i:i+seq_length])
        # 取序列最后一个时间步的标签作为样本标签
        y.append(labels[i+seq_length-1])
    return np.array(X), np.array(y)

# 评估模型
def evaluate_model(X_train, y_train, X_test, y_test, class_weights):
    verbose, epochs, batch_size = 1, 50, 128
    n_timesteps, n_features, n_outputs = X_train.shape[1], X_train.shape[2], y_train.shape[1]
    
    model = Sequential()
    # 双向LSTM捕捉双向时序特征
    model.add(Bidirectional(LSTM(128, return_sequences=True), input_shape=(n_timesteps, n_features)))
    model.add(BatchNormalization())
    model.add(Dropout(0.3))
    model.add(LSTM(64, return_sequences=False))
    model.add(BatchNormalization())
    model.add(Dropout(0.3))
    model.add(Dense(128, activation='relu'))
    model.add(Dense(n_outputs, activation='softmax'))
    
    model.compile(loss='categorical_crossentropy', optimizer='adam', metrics=['accuracy'])
    
    # 回调函数:早停+学习率衰减
    callbacks = [
        EarlyStopping(patience=10, restore_best_weights=True),
        ReduceLROnPlateau(factor=0.5, patience=5, min_lr=1e-6)
    ]
    
    # 训练模型,加入类别权重
    model.fit(X_train, y_train, 
              epochs=epochs, 
              batch_size=batch_size, 
              verbose=verbose,
              validation_split=0.1,
              callbacks=callbacks,
              class_weight=class_weights)
    
    # 评估
    _, accuracy = model.evaluate(X_test, y_test, batch_size=batch_size, verbose=1)
    return accuracy

# 主函数
def main():
    # 数据路径
    train_dir = '/Users/hosseinebrahimi/Documents/ThesisResults/Normal/NEW/*.xlsx'
    test_dir = '/Users/hosseinebrahimi/Documents/ThesisResults/Normal/NEW/test/*.xlsx'
    
    # 加载数据
    X_train, y_train, X_test, y_test = load_data(train_dir, test_dir)
    print("原始数据形状:", X_train.shape, y_train.shape, X_test.shape, y_test.shape)
    
    # 标准化:用RobustScaler按特征维度处理,避免异常值影响
    scaler = RobustScaler()
    X_train_scaled = scaler.fit_transform(X_train)
    X_test_scaled = scaler.transform(X_test)
    
    # 构造时序序列(设置序列长度为20,可根据实际调整)
    seq_length = 20
    X_train_seq, y_train_seq = create_sequences(X_train_scaled, y_train, seq_length)
    X_test_seq, y_test_seq = create_sequences(X_test_scaled, y_test, seq_length)
    print("序列数据形状:", X_train_seq.shape, y_train_seq.shape, X_test_seq.shape, y_test_seq.shape)
    
    # 计算类别权重,处理不平衡问题
    y_train_classes = np.argmax(y_train_seq, axis=1)
    class_weights = compute_class_weight(class_weight='balanced', classes=np.unique(y_train_classes), y=y_train_classes)
    class_weights = dict(enumerate(class_weights))
    
    # 运行实验
    repeats = 3
    scores = []
    for r in range(repeats):
        score = evaluate_model(X_train_seq, y_train_seq, X_test_seq, y_test_seq, class_weights)
        score = score * 100.0
        print(f"#{r+1}: %.3f%%" % score)
        scores.append(score)
    
    # 汇总结果
    print("测试精度:", scores)
    print(f"平均精度: %.3f%% (+/-%.3f)" % (np.mean(scores), np.std(scores)))

if __name__ == "__main__":
    main()

关键改进说明

  1. 时序序列构造:新增create_sequences函数,将连续时间步组合成序列,让LSTM能学习故障的时序演化特征。
  2. 标准化优化:改用RobustScaler,并严格用训练集拟合后转换测试集,避免数据泄露。
  3. 模型结构:加入双向LSTM,替换Recurrent Dropout为普通Dropout,增强特征捕捉能力同时避免过度抑制。
  4. 训练策略:添加Early Stopping防止过拟合,ReduceLROnPlateau动态调整学习率,加入类别权重处理数据不平衡。
  5. 代码模块化:将数据加载、序列构造等功能封装成函数,提升可读性和可维护性。

内容的提问来源于stack exchange,提问作者Hossein Ebrahimi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 15:39:52