You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PyTorch实现带位置编码的Transformer回归模型遇张量尺寸不匹配错误

问题分析与修复方案

错误根源

你遇到的RuntimeError是维度不匹配导致的,核心问题集中在以下几点:

1. 模型初始化参数传错

Transformer类初始化需要max_seq_len(序列最大长度),但构建模型时误传了num_heads=8:

model = Transformer(input_dim, hidden_dim, num_layers, num_heads)  # 错误:把注意力头数当成了序列长度

这直接导致位置编码的pe张量长度固定为8,和实际输入的序列维度冲突,相加时触发维度不匹配错误。

2. 输入形状不符合Transformer预期

Transformer处理的是3维序列数据:[batch_size, seq_len, input_dim],但你的输入是2维表格数据[样本数, input_dim],经过嵌入层后变成[batch_size, hidden_dim],而位置编码是3维张量[1, max_seq_len, hidden_dim],两者维度无法对齐相加。

3. 编码器层缺失核心组件

你定义的EncoderLayer只有前馈网络,完全没有Transformer的核心——多头自注意力机制,这根本不是标准的Transformer编码器结构。

4. 损失计算时形状不匹配

y_pred的形状是[batch_size, 1],但你把y_batch强制转成[1, batch_size],两者形状不一致,会导致损失计算错误。


逐步修复代码

第一步:修正Transformer模型结构

添加多头注意力模块,重构编码器层,确保输入输出维度对齐:

import math
import torch
import torch.nn as nn

class Transformer(nn.Module):
    def __init__(self, input_dim, hidden_dim, num_layers, num_heads, max_seq_len):
        super(Transformer, self).__init__()
        # 输入嵌入:将每个token的input_dim映射到hidden_dim
        self.embedding = nn.Linear(input_dim, hidden_dim)
        # 位置编码
        self.pos_encoding = PositionalEncoding(hidden_dim, max_seq_len)
        # 标准Transformer编码器层
        encoder_layer = nn.TransformerEncoderLayer(
            d_model=hidden_dim,
            nhead=num_heads,
            dim_feedforward=hidden_dim*4,  # 前馈网络维度通常设为hidden_dim的4倍
            dropout=0.1,
            batch_first=True  # 指定输入输出形状为[batch, seq_len, hidden_dim]
        )
        self.encoder = nn.TransformerEncoder(encoder_layer, num_layers=num_layers)
        # 回归输出层
        self.output_layer = nn.Linear(hidden_dim, 1)

    def forward(self, x):
        # x形状:[batch_size, seq_len, input_dim]
        x = self.embedding(x)  # 转换为[batch_size, seq_len, hidden_dim]
        x = self.pos_encoding(x)  # 加入位置编码
        x = self.encoder(x)  # 编码器处理序列
        x = x.mean(dim=1)  # 对序列维度取平均,得到全局表征
        x = self.output_layer(x)  # 输出回归结果
        return x

class PositionalEncoding(nn.Module):
    def __init__(self, hidden_dim, max_seq_len):
        super(PositionalEncoding, self).__init__()
        self.dropout = nn.Dropout(p=0.1)
        pe = torch.zeros(max_seq_len, hidden_dim)
        position = torch.arange(0, max_seq_len, dtype=torch.float).unsqueeze(1)
        div_term = torch.exp(torch.arange(0, hidden_dim, 2).float() * (-math.log(10000.0) / hidden_dim))
        pe[:, 0::2] = torch.sin(position * div_term)
        pe[:, 1::2] = torch.cos(position * div_term)
        self.register_buffer('pe', pe.unsqueeze(0))  # 形状:[1, max_seq_len, hidden_dim]

    def forward(self, x):
        # x形状:[batch_size, seq_len, hidden_dim]
        x = x + self.pe[:, :x.size(1)]  # 取对应长度的位置编码
        x = self.dropout(x)
        return x

第二步:修正数据预处理与模型初始化

将2维表格数据转换为3维序列格式,正确传递模型参数:

from sklearn.model_selection import train_test_split

# 假设X1是表格数据,形状为[样本数, 特征数],我们将每个特征作为序列的一个token
input_dim = 1  # 每个token的维度
seq_len = X1.shape[1]  # 序列长度等于特征数
hidden_dim = 16
num_layers = 2
num_heads = 8
# 注意:hidden_dim必须能被num_heads整除,否则多头注意力会报错
assert hidden_dim % num_heads == 0, "hidden_dim必须能被num_heads整除"
max_seq_len = seq_len
lr = 1e-3
batch_size = 2
epochs = 10  # 增加epochs保证训练有效

X_train, X_val, y_train, y_val = train_test_split(X1, y1, test_size=0.2, random_state=42)

# 调整输入为3维序列:[样本数, seq_len, input_dim]
X_train = torch.tensor(X_train.values, dtype=torch.float).unsqueeze(-1)
X_val = torch.tensor(X_val.values, dtype=torch.float).unsqueeze(-1)
y_train = torch.tensor(y_train.values, dtype=torch.float).unsqueeze(1)
y_val = torch.tensor(y_val.values, dtype=torch.float).unsqueeze(1)

# 创建模型:传递正确的参数
model = Transformer(input_dim, hidden_dim, num_layers, num_heads, max_seq_len)
optimizer = torch.optim.Adam(model.parameters(), lr=lr)
criterion = nn.MSELoss()

# 数据加载器
train_dataset = torch.utils.data.TensorDataset(X_train, y_train)
train_loader = torch.utils.data.DataLoader(train_dataset, batch_size=batch_size, shuffle=True)
val_dataset = torch.utils.data.TensorDataset(X_val, y_val)
val_loader = torch.utils.data.DataLoader(val_dataset, batch_size=batch_size, shuffle=False)

# 训练循环
for epoch in range(epochs):
    model.train()
    train_loss = 0
    for X_batch, y_batch in train_loader:
        optimizer.zero_grad()
        y_pred = model(X_batch)
        # 确保y_pred和y_batch形状一致:[batch_size, 1]
        loss = criterion(y_pred, y_batch)
        loss.backward()
        optimizer.step()
        train_loss += loss.item() * X_batch.shape[0]
    train_loss /= len(train_dataset)

    model.eval()
    val_loss = 0
    with torch.no_grad():
        for X_batch, y_batch in val_loader:
            y_pred = model(X_batch)
            loss = criterion(y_pred, y_batch)
            val_loss += loss.item() * X_batch.shape[0]
        val_loss /= len(val_dataset)

    print(f"Epoch {epoch+1}/{epochs}, Train Loss: {train_loss:.4f}, Val Loss: {val_loss:.4f}")

关于改用TensorFlow的问题

完全可以改用TensorFlow/Keras实现带位置编码的Transformer回归模型,Keras内置了MultiHeadAttention和TransformerEncoder层,实现更简洁:

示例代码片段:

import tensorflow as tf
from tensorflow.keras import layers
import math

def positional_encoding(max_seq_len, hidden_dim):
    position = tf.range(max_seq_len, dtype=tf.float32)[:, tf.newaxis]
    div_term = tf.exp(tf.range(0, hidden_dim, 2, dtype=tf.float32) * (-math.log(10000.0) / hidden_dim))
    pe = tf.zeros((max_seq_len, hidden_dim))
    pe[:, 0::2] = tf.sin(position * div_term)
    pe[:, 1::2] = tf.cos(position * div_term)
    pe = pe[tf.newaxis, :, :]
    return tf.cast(pe, dtype=tf.float32)

def build_transformer_regressor(input_dim, hidden_dim, num_layers, num_heads, max_seq_len):
    inputs = layers.Input(shape=(max_seq_len, input_dim))
    # 输入嵌入
    x = layers.Dense(hidden_dim)(inputs)
    # 位置编码
    pe = positional_encoding(max_seq_len, hidden_dim)
    x = x + pe[:, :tf.shape(x)[1]]
    x = layers.Dropout(0.1)(x)
    # 堆叠编码器层
    for _ in range(num_layers):
        # 多头自注意力
        attn_output = layers.MultiHeadAttention(num_heads=num_heads, key_dim=hidden_dim)(x, x)
        x = layers.LayerNormalization(epsilon=1e-6)(x + attn_output)
        # 前馈网络
        ff_output = layers.Dense(hidden_dim*4, activation='relu')(x)
        ff_output = layers.Dense(hidden_dim)(ff_output)
        x = layers.LayerNormalization(epsilon=1e-6)(x + ff_output)
        x = layers.Dropout(0.1)(x)
    # 回归输出
    x = layers.GlobalAveragePooling1D()(x)
    outputs = layers.Dense(1)(x)
    model = tf.keras.Model(inputs, outputs)
    return model

# 初始化模型并训练
model = build_transformer_regressor(input_dim=1, hidden_dim=16, num_layers=2, num_heads=8, max_seq_len=seq_len)
model.compile(optimizer='adam', loss='mse')
model.fit(X_train, y_train, epochs=10, batch_size=2, validation_data=(X_val, y_val))

不管用PyTorch还是TensorFlow,核心逻辑一致:确保输入为序列格式、位置编码维度匹配、编码器包含多头注意力+残差连接。

内容的提问来源于stack exchange,提问作者Waleed Kh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 01:22:06