Transformer时间序列分类模型输入形状不兼容问题求助
时间序列Transformer分类模型输入形状不匹配问题解决
问题根源分析
报错ValueError: Input 0 of layer "model" is incompatible with the layer: expected shape=(None, 41671, 43), found shape=(None, 43)的核心原因是:
- 你将整个训练集的形状
X_train.shape(即(41671, 43))传给了build_model的input_shape参数,但Keras的Input(shape=...)需要的是单个样本的形状,而非数据集的整体形状。 - 模型错误地将数据集的样本数量(41671)当成了时间序列的步长,导致期望输入维度为
(批量大小, 41671, 43),但实际输入的每个样本仅为(43,),维度不匹配。
此外代码还存在其他几个关键问题:
GlobalAveragePooling1D的data_format="channels_first"设置错误,与输入特征维度位置不匹配。- 输出层激活函数与损失函数不匹配:二分类任务用
softmax+sparse_categorical_crossentropy是错误组合。 - 评估时变量名错误:使用了未定义的
model而非训练好的model5。
针对性解决方案
1. 修正输入形状与数据维度
- 将
build_model的input_shape参数改为单个样本的序列形状:如果每个样本是单时间步+43个特征,则设置为input_shape=(1, X_train.shape[1])。 - 同时将训练/验证数据增加一个时间步维度,从
(样本数, 43)转换为(样本数, 1, 43),可通过np.expand_dims实现。
2. 调整池化层参数
移除GlobalAveragePooling1D的data_format="channels_first"参数,使用默认的channels_last(特征维度在最后),确保池化后输出形状与MLP层匹配。
3. 匹配输出层与损失函数
如果是二分类任务:
- 输出层改为
Dense(1, activation="sigmoid") - 损失函数用
binary_crossentropy,评估指标用binary_accuracy
如果是多分类任务(标签为整数索引):
- 输出层神经元数设为类别总数,激活用
softmax - 损失函数保持
sparse_categorical_crossentropy,指标用sparse_categorical_accuracy
4. 修复变量名错误
将model.evaluate(X_valid, y_valid, verbose=1)改为model5.evaluate(X_valid, y_valid, verbose=1)
修正后的完整代码
# Model 5: Time-Series Transformer for Classification import tensorflow as tf from tensorflow import keras from keras import layers import numpy as np import pandas as pd tickers = ['AAPL', 'GOOG', 'MSFT', 'INTC', 'AMZN'] def transformer_encoder(inputs, head_size, num_heads, ff_dim, dropout=0): x = layers.LayerNormalization(epsilon=1e-6)(inputs) x = layers.MultiHeadAttention( key_dim=head_size, num_heads=num_heads, dropout=dropout )(x, x) x = layers.Dropout(dropout)(x) res = x + inputs # Feed Forward Part x = layers.LayerNormalization(epsilon=1e-6)(res) x = layers.Conv1D(filters=ff_dim, kernel_size=1, activation="relu")(x) x = layers.Dropout(dropout)(x) x = layers.Conv1D(filters=inputs.shape[-1], kernel_size=1)(x) return x + res def build_model( input_shape, head_size, num_heads, ff_dim, num_transformer_blocks, mlp_units, dropout=0, mlp_dropout=0, ): inputs = keras.Input(shape=input_shape) x = inputs for _ in range(num_transformer_blocks): x = transformer_encoder(x, head_size, num_heads, ff_dim, dropout) # 使用默认channels_last格式 x = layers.GlobalAveragePooling1D()(x) for dim in mlp_units: x = layers.Dense(dim, activation="relu")(x) x = layers.Dropout(mlp_dropout)(x) # 二分类任务用sigmoid激活 outputs = layers.Dense(1, activation="sigmoid")(x) return keras.Model(inputs, outputs) # Model Training for i in range(len(tickers)): train_tickers = tickers.copy() train_tickers.pop(i) print(train_tickers) df_train = pd.DataFrame() for train_ticker in train_tickers: df = pd.read_csv(f"Spoofing-Injected DataFrames/{train_ticker}_segmentsummary_spoofed_bidside_{FACTOR}_{INTERVAL_START}_{INTERVAL_END}.csv", index_col=0) df_train = pd.concat([df_train, df], axis=0) X_train = df_train.drop("Classification", axis=1) y_train = df_train["Classification"] # 增加时间步维度,转换为(样本数, 1, 43) X_train = np.expand_dims(X_train.values, axis=1) df_valid = pd.read_csv(f"Spoofing-Injected DataFrames/{tickers[i]}_segmentsummary_spoofed_bidside_{FACTOR}_{INTERVAL_START}_{INTERVAL_END}.csv", index_col=0) X_valid = df_valid.drop("Classification", axis=1) y_valid = df_valid["Classification"] # 验证集同步增加维度 X_valid = np.expand_dims(X_valid.values, axis=1) model5 = build_model( # 传入单个样本的形状:(1, 43) input_shape=(1, X_train.shape[2]), head_size=256, num_heads=4, ff_dim=4, num_transformer_blocks=4, mlp_units=[128], mlp_dropout=0.4, dropout=0.25 ) # 二分类对应损失函数与指标 model5.compile( loss='binary_crossentropy', optimizer=keras.optimizers.Adam(learning_rate=1e-4), metrics=['binary_accuracy'] ) model5.summary() callbacks = [keras.callbacks.EarlyStopping(patience=10, restore_best_weights=True)] history = model5.fit( X_train, y_train, validation_split=0.2, epochs=200, batch_size=64, callbacks=callbacks, verbose=0 ) # 修正评估变量名 model5.evaluate(X_valid, y_valid, verbose=1) history_df = pd.DataFrame(history.history) # 从第5轮开始绘制曲线 history_df.loc[5:, ['loss', 'val_loss']].plot() history_df.loc[5:, ['binary_accuracy', 'val_binary_accuracy']].plot() print(("Best Validation Loss: {:0.4f}" +\ "\nBest Validation Accuracy: {:0.4f}")\ .format(history_df['val_loss'].min(), history_df['val_binary_accuracy'].max()))
额外注意事项
- 如果数据集是将连续多行作为一个时间序列样本(比如每个样本包含N个时间步,每个时间步43个特征),则需要先对原始数据进行滑窗分割,生成形状为
(样本数, N, 43)的输入,此时input_shape设置为(N, 43)即可。 - 交叉验证时,不同折的样本数不同,但只要每个样本的形状统一(比如都是
(1,43)或(N,43)),模型就能正常接受输入,无需额外修改。
内容的提问来源于stack exchange,提问作者Marwan Ismail
相关产品推荐
相关产品推荐

