You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

LSTM训练外汇时间序列重采样后Loss为NaN、精度骤降问题排查

问题描述

我用外汇时间序列数据训练LSTM模型,使用原始1分钟数据时模型精度正常、Loss表现稳定;但将数据重采样为1小时或24小时周期后,模型直接失效:Loss变为NaN,精度暴跌,训练日志如下:

Epoch 5/5
110/110 [==============================] - 4s 32ms/step - loss: nan - val_loss: nan - accuracy: 0.4437 - val_accuracy: 0.3453

我的数据加载函数如下:

def readCSV(path, candleTime):
    # Load data
    #print("Loading: " + path)
    data = pd.read_csv(path, names=['Date_Time', 'open', 'high', 'low', 'close', 'volume'], sep=";", index_col=0)

    # Convert the index to datetime
    data.index = pd.to_datetime(data.index, format='%Y%m%d %H%M%S%f')

    if(candleTime == CandleTime.hour):
        data = data.resample('1H').agg({'open': 'first', 
                                 'high': 'max', 
                                 'low': 'min', 
                                 'close': 'last'})
    if(candleTime == CandleTime.day):
        data = data.resample('24H').agg({'open': 'first', 
                            'high': 'max', 
                            'low': 'min', 
                            'close': 'last'})
            
    return data

补充说明

完整训练函数如下(当前仅执行一次迭代):

def trainModel(trainCandles, prediction_minutes = 60, model_name = 'lstm_1m_10_model'):
tf.keras.backend.clear_session()
#Prepare Data
print("Preparing data..")
x_train = []
y_train = []
normalizedCandles = trainCandles[['open', 'high', 'low', 'close']].to_numpy(copy=True)
for x in range(prediction_minutes, len(normalizedCandles)):
    xdata = normalizedCandles[x-prediction_minutes:x]
    predictionData = []
    for candleX in xdata:
        predictionData.append([candleX[0], candleX[1], candleX[2], candleX[3]])
    candleY = normalizedCandles[x]
    x_train.append(predictionData)
    y_train.append([candleY[0], candleY[1], candleY[2], candleY[3]])

print("Spliting..")
# split train and test
x_toSplit, y_toSplit = x_train, y_train
sizeOf70percentage = int(len(x_toSplit)/100*70)
x_test = np.array(x_toSplit[sizeOf70percentage:len(x_toSplit)])
y_test = np.array(y_toSplit[sizeOf70percentage:len(x_toSplit)])
x_train = np.array(x_toSplit[0: sizeOf70percentage])
y_train = np.array(y_toSplit[0: sizeOf70percentage])


print("Total size of samples: " + str(len(x_train)))
model=None

if (os.path.isdir(model_name)): # you won't have a model for first iteration
    print("Loading model..")
    model = load_model(model_name)
else:
    print("Creatng model..")
    model = Sequential()
    model.add(LSTM(units=50, return_sequences=True, input_shape=(x_train.shape[1], x_train.shape[2])))
    model.add(Dropout(0.2))
    model.add(LSTM(units=50, return_sequences=True))
    model.add(Dropout(0.2))
    model.add(LSTM(units=50))
    model.add(Dropout(0.2))
    model.add(Dense(units=4))
    model.compile(optimizer='Adam', loss='mean_squared_error', metrics=["accuracy"])


history = model.fit(
    x_train, 
    y_train, 
    validation_data=(x_test, y_test), 
    epochs=5, 
    batch_size=32)

model.save(model_name)
问题排查与解决方案

1. 重采样后的数据缺失值问题

重采样1H/24H时,若原始1分钟数据存在时间段缺失(比如周末休市),resample会生成空行,导致数据中出现NaN,后续训练时模型计算Loss会直接变为NaN。

修复方法:
重采样后添加缺失值处理逻辑,比如删除空行或填充:

# 重采样后添加该行代码
data = data.dropna()
# 或用前后值填充缺失值
# data = data.fillna(method='ffill').fillna(method='bfill')

2. 未做数据归一化(核心问题)

代码中normalizedCandles仅将数据转为numpy数组,完全未做归一化处理!1分钟数据价格波动范围小,模型尚可勉强训练;但1H/24H数据价格跨度大,数值范围差异会引发梯度爆炸,直接导致Loss变为NaN。

修复方法:
使用MinMaxScaler对OHLC数据做归一化,注意保留缩放器用于后续预测:

from sklearn.preprocessing import MinMaxScaler

def trainModel(trainCandles, seq_length = 24, model_name = 'lstm_1h_model'):
    tf.keras.backend.clear_session()
    print("Preparing data..")
    
    # 初始化归一化器,对OHLC四列做缩放
    scaler = MinMaxScaler(feature_range=(0,1))
    normalizedCandles = scaler.fit_transform(trainCandles[['open', 'high', 'low', 'close']])
    
    x_train = []
    y_train = []
    for x in range(seq_length, len(normalizedCandles)):
        x_train.append(normalizedCandles[x-seq_length:x])
        y_train.append(normalizedCandles[x])
    
    # 后续拆分、模型构建代码不变...

3. 输入序列长度与数据量不匹配

你用prediction_minutes=60作为输入序列长度,在1分钟数据里对应60根K线(1小时);但在1H数据里,60根K线对应60小时,重采样后的1H数据总条数可能远小于这个值,导致训练样本极少,模型无法收敛甚至出现计算异常。

修复方法:
根据不同K线周期调整输入序列长度,比如1H数据用24根(1天),24H数据用7根(1周):

# 可在函数参数中添加序列长度参数,或根据周期自动调整
def trainModel(trainCandles, seq_length = 24, model_name = 'lstm_1h_model'):
    # 替换原prediction_minutes为seq_length
    for x in range(seq_length, len(normalizedCandles)):
        x_train.append(normalizedCandles[x-seq_length:x])
        y_train.append(normalizedCandles[x])

4. 回归任务误用accuracy指标

你用MSE作为Loss(回归任务),但同时用accuracy作为评估指标,这完全不匹配——accuracy是分类任务的指标,对回归问题毫无意义,所以看到的0.44精度是无效值。

修复方法:
将metrics替换为回归任务适用的指标,比如mae(平均绝对误差):

model.compile(optimizer='Adam', loss='mean_squared_error', metrics=["mae"])

内容的提问来源于stack exchange,提问作者Luboš Hájek

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 13:15:33