You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Transformer时间序列分类模型输入形状不兼容问题求助

时间序列Transformer分类模型输入形状不匹配问题解决

问题根源分析

报错ValueError: Input 0 of layer "model" is incompatible with the layer: expected shape=(None, 41671, 43), found shape=(None, 43)的核心原因是:

  • 你将整个训练集的形状X_train.shape(即(41671, 43))传给了build_model的input_shape参数,但Keras的Input(shape=...)需要的是单个样本的形状,而非数据集的整体形状。
  • 模型错误地将数据集的样本数量(41671)当成了时间序列的步长,导致期望输入维度为(批量大小, 41671, 43),但实际输入的每个样本仅为(43,),维度不匹配。

此外代码还存在其他几个关键问题:

  1. GlobalAveragePooling1D的data_format="channels_first"设置错误,与输入特征维度位置不匹配。
  2. 输出层激活函数与损失函数不匹配:二分类任务用softmax+sparse_categorical_crossentropy是错误组合。
  3. 评估时变量名错误:使用了未定义的model而非训练好的model5。

针对性解决方案

1. 修正输入形状与数据维度

  • 将build_model的input_shape参数改为单个样本的序列形状:如果每个样本是单时间步+43个特征,则设置为input_shape=(1, X_train.shape[1])。
  • 同时将训练/验证数据增加一个时间步维度,从(样本数, 43)转换为(样本数, 1, 43),可通过np.expand_dims实现。

2. 调整池化层参数

移除GlobalAveragePooling1D的data_format="channels_first"参数,使用默认的channels_last(特征维度在最后),确保池化后输出形状与MLP层匹配。

3. 匹配输出层与损失函数

如果是二分类任务:

  • 输出层改为Dense(1, activation="sigmoid")
  • 损失函数用binary_crossentropy,评估指标用binary_accuracy

如果是多分类任务(标签为整数索引):

  • 输出层神经元数设为类别总数,激活用softmax
  • 损失函数保持sparse_categorical_crossentropy,指标用sparse_categorical_accuracy

4. 修复变量名错误

将model.evaluate(X_valid, y_valid, verbose=1)改为model5.evaluate(X_valid, y_valid, verbose=1)

修正后的完整代码

# Model 5: Time-Series Transformer for Classification
import tensorflow as tf
from tensorflow import keras
from keras import layers
import numpy as np
import pandas as pd

tickers = ['AAPL', 'GOOG', 'MSFT', 'INTC', 'AMZN']

def transformer_encoder(inputs, head_size, num_heads, ff_dim, dropout=0):
    x = layers.LayerNormalization(epsilon=1e-6)(inputs)
    x = layers.MultiHeadAttention(
        key_dim=head_size, num_heads=num_heads, dropout=dropout
    )(x, x)
    x = layers.Dropout(dropout)(x)
    res = x + inputs

    # Feed Forward Part
    x = layers.LayerNormalization(epsilon=1e-6)(res)
    x = layers.Conv1D(filters=ff_dim, kernel_size=1, activation="relu")(x)
    x = layers.Dropout(dropout)(x)
    x = layers.Conv1D(filters=inputs.shape[-1], kernel_size=1)(x)
    return x + res

def build_model(
    input_shape,
    head_size,
    num_heads,
    ff_dim,
    num_transformer_blocks,
    mlp_units,
    dropout=0,
    mlp_dropout=0,
):
    inputs = keras.Input(shape=input_shape)
    x = inputs
    for _ in range(num_transformer_blocks):
        x = transformer_encoder(x, head_size, num_heads, ff_dim, dropout)

    # 使用默认channels_last格式
    x = layers.GlobalAveragePooling1D()(x)
    for dim in mlp_units:
        x = layers.Dense(dim, activation="relu")(x)
        x = layers.Dropout(mlp_dropout)(x)
    # 二分类任务用sigmoid激活
    outputs = layers.Dense(1, activation="sigmoid")(x)
    return keras.Model(inputs, outputs)

# Model Training
for i in range(len(tickers)):
    train_tickers = tickers.copy()
    train_tickers.pop(i)
    print(train_tickers)
    df_train = pd.DataFrame()
    for train_ticker in train_tickers:
        df = pd.read_csv(f"Spoofing-Injected DataFrames/{train_ticker}_segmentsummary_spoofed_bidside_{FACTOR}_{INTERVAL_START}_{INTERVAL_END}.csv", index_col=0)
        df_train = pd.concat([df_train, df], axis=0)
    X_train = df_train.drop("Classification", axis=1)
    y_train = df_train["Classification"]
    
    # 增加时间步维度,转换为(样本数, 1, 43)
    X_train = np.expand_dims(X_train.values, axis=1)
    
    df_valid = pd.read_csv(f"Spoofing-Injected DataFrames/{tickers[i]}_segmentsummary_spoofed_bidside_{FACTOR}_{INTERVAL_START}_{INTERVAL_END}.csv", index_col=0)
    X_valid = df_valid.drop("Classification", axis=1)
    y_valid = df_valid["Classification"]
    
    # 验证集同步增加维度
    X_valid = np.expand_dims(X_valid.values, axis=1)

    model5 = build_model(
        # 传入单个样本的形状:(1, 43)
        input_shape=(1, X_train.shape[2]),
        head_size=256,
        num_heads=4,
        ff_dim=4,
        num_transformer_blocks=4,
        mlp_units=[128],
        mlp_dropout=0.4,
        dropout=0.25
    )

    # 二分类对应损失函数与指标
    model5.compile(
        loss='binary_crossentropy',
        optimizer=keras.optimizers.Adam(learning_rate=1e-4),
        metrics=['binary_accuracy']
    )

    model5.summary()

    callbacks = [keras.callbacks.EarlyStopping(patience=10, restore_best_weights=True)]
    history = model5.fit(
        X_train,
        y_train,
        validation_split=0.2,
        epochs=200,
        batch_size=64,
        callbacks=callbacks,
        verbose=0
    )

    # 修正评估变量名
    model5.evaluate(X_valid, y_valid, verbose=1)
    history_df = pd.DataFrame(history.history)
    # 从第5轮开始绘制曲线
    history_df.loc[5:, ['loss', 'val_loss']].plot()
    history_df.loc[5:, ['binary_accuracy', 'val_binary_accuracy']].plot()

    print(("Best Validation Loss: {:0.4f}" +\
        "\nBest Validation Accuracy: {:0.4f}")\
        .format(history_df['val_loss'].min(), 
                history_df['val_binary_accuracy'].max()))

额外注意事项

  • 如果数据集是将连续多行作为一个时间序列样本(比如每个样本包含N个时间步,每个时间步43个特征),则需要先对原始数据进行滑窗分割,生成形状为(样本数, N, 43)的输入,此时input_shape设置为(N, 43)即可。
  • 交叉验证时,不同折的样本数不同,但只要每个样本的形状统一(比如都是(1,43)或(N,43)),模型就能正常接受输入,无需额外修改。

内容的提问来源于stack exchange,提问作者Marwan Ismail

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.16 09:15:06