使用LSTM预测苹果收盘价时触发pandas InvalidIndexError报错
问题说明
使用过去50天的历史股价数据预测苹果公司(Apple Inc.)股票收盘价,运行代码时抛出如下报错:
pandas.errors.InvalidIndexError: (slice(115770, None, None), slice(None, None, None))
涉及的核心实现代码如下:
#Import the librarires import math import pandas_datareader as web import numpy as np import pandas as pd from sklearn.preprocessing import MinMaxScaler from keras.models import Sequential from keras.layers import Dense, LSTM import matplotlib.pyplot as plt plt.style.use('fivethirtyeight') #Get the stock quote df =web.DataReader ('AAPL', data_source='yahoo', start='2017-01-01', end='2022-05-31') ..... #Create the data sets x_test and y_test x_test = [] y_test = dataset[training_data_len:, :] for i in range(50, len(test_data)): x_test.append(test_data[i-50:i, 0]) #Convert Data to a numpy array x_test = np.array(x_test) #Reshape the data x_test = np.reshape(x_test, (x_test.shape[0], x_test.shape[1], 1 )) #Get the models predicte price values predictions = model.predict(x_test) #inverse transform the prediction # basically unscaling the values and have the prediction contain the same values as the y_testdataset, given the x_dataset predictions = scaler.inverse_transform(predictions)
报错原因
这个错误是直接对pandas的DataFrame/Series对象使用了numpy风格的多维位置切片语法导致的:
- pandas对象原生不支持
[行切片, 列位置]这种numpy数组的切片写法,直接使用时pandas会把传入的整数切片当成索引标签去匹配,而不是按位置取数,当传入的位置值不在索引标签里时就会抛出索引错误。 - 代码中
y_test = dataset[training_data_len:, :]、循环内的test_data[i-50:i, 0]都属于这类问题,如果dataset、test_data是pandas结构而非numpy数组,就会触发报错。
修复方案
- 所有做数值切片、送入缩放器/模型之前,先把pandas的表格结构转为numpy数组,不要直接对DataFrame/Series使用numpy式多维切片。
- 构造测试集时注意要从训练集结束位置往前预留50天的窗口,保证第一个测试样本能取到完整的50天历史输入。
修正后的对应代码片段如下:
# 读取行情数据后,单独提取收盘价列并转为numpy数组 df = web.DataReader('AAPL', data_source='yahoo', start='2017-01-01', end='2022-05-31') # 提取收盘价,转numpy数组 close_data = df.filter(['Close']).values # 归一化处理,此时scaled_data输出为numpy数组 scaler = MinMaxScaler(feature_range=(0,1)) scaled_data = scaler.fit_transform(close_data) # --- 中间划分训练集的逻辑保持不变,计算得到training_data_len --- # 构造测试集:从训练集结束位置往前推50天,保证窗口完整 test_data = scaled_data[training_data_len - 50: , :] x_test = [] # y_test直接取对应位置的真实收盘价,numpy数组切片不会触发pandas索引错误 y_test = close_data[training_data_len:, :] for i in range(50, len(test_data)): x_test.append(test_data[i-50:i, 0]) # 后续转数组、reshape、预测、逆缩放逻辑无需修改 x_test = np.array(x_test) x_test = np.reshape(x_test, (x_test.shape[0], x_test.shape[1], 1)) predictions = model.predict(x_test) predictions = scaler.inverse_transform(predictions)
内容的提问来源于stack exchange,提问作者DeepestLore
相关产品推荐
相关产品推荐

