Transformer时间序列预测中解决IndexError: tuple index out of range问题
解决Transformer时间序列预测中的IndexError: tuple index out of range问题
你的输入维度是(None, 30)(2D张量),但Keras的MultiHeadAttention层要求输入必须是3D张量,形状为(batch_size, sequence_length, feature_dim)——它需要明确区分序列长度和每个时间步的特征维度,否则在内部处理时会因无法找到特征维度的索引(张量只有2个维度,索引只能到1,但代码需要访问索引2)而触发IndexError。
解决方案
1. 将输入张量从2D转为3D
你需要给每个时间步添加一个特征维度,把(None, 30)变成(None, 30, 1)(适用于单变量时间序列场景),有两种实现方式:
- 预处理数据时调整形状:
# 假设X_train原形状是(samples, 30) X_train = X_train.reshape(-1, 30, 1) X_test = X_test.reshape(-1, 30, 1) input_shape = X_train.shape[1:] # 现在变为(30, 1) - 在模型内部通过Reshape层调整:
在build_model函数的输入层后添加维度转换:def build_model(...): inputs = keras.Input(shape=input_shape) # 把2D输入转为3D x = layers.Reshape((input_shape[0], 1))(inputs) for _ in range(num_transformer_blocks): x = transformer_encoder(x, head_size, num_heads, ff_dim, dropout) # ... 其余代码
2. 修正GlobalAveragePooling1D的参数
你使用了data_format="channels_first",但调整后的3D输入是channels_last格式((batch, seq_len, features)),这会导致池化层处理错误。需要改为channels_last(或直接省略,默认值即为该格式):
x = layers.GlobalAveragePooling1D()(x)
修改后的完整代码示例
from tensorflow import keras from tensorflow.keras import layers def transformer_encoder(inputs, head_size, num_heads, ff_dim, dropout=0): # Attention and Normalization print(inputs.shape) x = layers.MultiHeadAttention( key_dim=head_size, num_heads=num_heads, dropout=dropout )(inputs, inputs) x = layers.Dropout(dropout)(x) x = layers.LayerNormalization(epsilon=1e-6)(x) res = x + inputs # Feed Forward Part x = layers.Conv1D(filters=ff_dim, kernel_size=1, activation="relu")(res) x = layers.Dropout(dropout)(x) x = layers.Conv1D(filters=inputs.shape[-1], kernel_size=1)(x) x = layers.LayerNormalization(epsilon=1e-6)(x) return x + res def build_model( input_shape, head_size, num_heads, ff_dim, num_transformer_blocks, mlp_units, dropout=0, mlp_dropout=0, n_classes=1 # 单变量预测场景下设为1 ): inputs = keras.Input(shape=input_shape) # 自动处理2D转3D if len(input_shape) == 1: x = layers.Reshape((input_shape[0], 1))(inputs) else: x = inputs for _ in range(num_transformer_blocks): x = transformer_encoder(x, head_size, num_heads, ff_dim, dropout) x = layers.GlobalAveragePooling1D()(x) for dim in mlp_units: x = layers.Dense(dim, activation="relu")(x) x = layers.Dropout(mlp_dropout)(x) outputs = layers.Dense(n_classes)(x) return keras.Model(inputs, outputs) # 预处理数据(假设X_train原形状是(samples,30)) X_train = X_train.reshape(-1, 30, 1) input_shape = X_train.shape[1:] # 定义优化器和损失函数 adam = keras.optimizers.Adam() root_mean_squared_error = keras.losses.RootMeanSquaredError() model_mlp = build_model( input_shape, head_size=256, num_heads=1, ff_dim=1, num_transformer_blocks=4, mlp_units=[128], mlp_dropout=0.4, dropout=0.25, ) model_mlp.compile(optimizer=adam, loss=root_mean_squared_error) model_mlp.summary()
内容的提问来源于stack exchange,提问作者assa
相关产品推荐
相关产品推荐

