You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

GPT-2训练时nn.CrossEntropyLoss报RuntimeError:目标尺寸不匹配

GPT-2训练时CrossEntropyLoss形状不匹配错误

问题描述

训练GPT-2模型,输入为分词/填充后的张量,batch size设为32,最大序列长度343,模型自带维度768,训练循环持续抛出错误:

RuntimeError: Expected target size [32, 768], got [32, 343]

代码片段

# Create a TensorDataset from input_ids and output_ids
dataset = TensorDataset(input_tensors, output_tensors)

#Constants
batch_size = 32
num_epochs = 20
# Create a DataLoader from the dataset
dataloader = DataLoader(dataset, batch_size=batch_size, shuffle=True)

# Set the device to run on
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")

# Define the model architecture
model = transformers.GPT2Model.from_pretrained('gpt2').to(device)

# Define the loss function
loss_function = nn.CrossEntropyLoss(ignore_index=0, reduction='mean')

# Define the optimizer
optimizer = torch.optim.Adam(model.parameters(), lr=0.001)

# Set the model to training mode
model.train()
print(f"input_tensors.shape before the loop: {input_tensors.shape}")
print(f"output_tensors.shape before the loop: {output_tensors.shape}")

# Loop over the number of epochs
for epoch in range(num_epochs):
    # Initialize the epoch loss
    epoch_loss = 0
    
    # Loop over the data in the dataloader
    for input_tensors, output_tensors in dataloader:
        # Send the input and target tensors to the device
        input_tensors = input_tensors.to(device)
        output_tensors = output_tensors.type(torch.LongTensor)
        output_tensors = output_tensors.to(device)
        # Zero gradients
        optimizer.zero_grad()
        
        # Begin Forward pass
        logits = model(input_tensors)[0]
        
        print(f"logits.shape: {logits.shape}")
        print(f"input_tensors.shape: {input_tensors.shape}")
        print(f"output_tensors.shape: {output_tensors.shape}")
        
        # Compute the loss
        loss = loss_function(logits, output_tensors)

        # Backward pass
        loss.backward()

        # Update the model parameters
        optimizer.step()

        # Add the loss to the epoch loss
        epoch_loss += loss.item()
        # Print the epoch loss
    print(f'Epoch {epoch+1}: Loss = {epoch_loss}')

各张量尺寸

  • 循环前input_tensors.shape == torch.Size([2625, 343])
  • 循环前output_tensors.shape == torch.Size([2625, 343])
  • logits.shape == torch.Size([32, 343, 768])
  • 循环内input_tensors.shape == torch.Size([32, 343])
  • 循环内output_tensors.shape == torch.Size([32, 343])

解决方案

核心问题

你使用了GPT2Model而非GPT2LMHeadModel:

  • GPT2Model输出的是模型的隐藏层特征,维度为[batch_size, seq_len, hidden_size](即你的[32,343,768]),这并不是语言模型任务需要的词汇表概率分布。
  • CrossEntropyLoss默认会将输入的最后一维视为类别数,因此它期望目标张量的形状是[32,768],但你的目标是[32,343](每个位置的token ID),导致形状不匹配。

修改步骤

  1. 替换模型为GPT2LMHeadModel
    该模型自带语言模型头,会将隐藏层特征映射到词汇表大小的logits(GPT-2词汇表大小为50257),输出维度变为[32,343,50257],与目标的序列长度维度匹配。

  2. 调整张量形状以适配CrossEntropyLoss
    CrossEntropyLoss要求输入形状为[N, C](N为样本数,C为类别数),目标形状为[N],因此需要将序列维度展开。

修改后的关键代码

# 替换模型定义
model = transformers.GPT2LMHeadModel.from_pretrained('gpt2').to(device)

# 在Forward pass后调整形状并计算损失
logits = model(input_tensors)[0]
# 将logits调整为 [batch*seq_len, vocab_size]
logits = logits.view(-1, logits.size(-1))
# 将目标调整为 [batch*seq_len]
output_tensors = output_tensors.view(-1)
loss = loss_function(logits, output_tensors)

其他注意事项

  • 无需再使用squeeze/unsqueeze调整维度,上述形状调整已经完全适配损失函数要求。
  • 保持ignore_index=0可以忽略填充token(假设你用0作为填充值)的损失计算,符合语言模型训练的常规做法。

内容的提问来源于stack exchange,提问作者C_Dog

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 08:35:24