GPT-2训练时nn.CrossEntropyLoss报RuntimeError:目标尺寸不匹配
GPT-2训练时CrossEntropyLoss形状不匹配错误
问题描述
训练GPT-2模型,输入为分词/填充后的张量,batch size设为32,最大序列长度343,模型自带维度768,训练循环持续抛出错误:
RuntimeError: Expected target size [32, 768], got [32, 343]
代码片段
# Create a TensorDataset from input_ids and output_ids dataset = TensorDataset(input_tensors, output_tensors) #Constants batch_size = 32 num_epochs = 20 # Create a DataLoader from the dataset dataloader = DataLoader(dataset, batch_size=batch_size, shuffle=True) # Set the device to run on device = torch.device("cuda" if torch.cuda.is_available() else "cpu") # Define the model architecture model = transformers.GPT2Model.from_pretrained('gpt2').to(device) # Define the loss function loss_function = nn.CrossEntropyLoss(ignore_index=0, reduction='mean') # Define the optimizer optimizer = torch.optim.Adam(model.parameters(), lr=0.001) # Set the model to training mode model.train() print(f"input_tensors.shape before the loop: {input_tensors.shape}") print(f"output_tensors.shape before the loop: {output_tensors.shape}") # Loop over the number of epochs for epoch in range(num_epochs): # Initialize the epoch loss epoch_loss = 0 # Loop over the data in the dataloader for input_tensors, output_tensors in dataloader: # Send the input and target tensors to the device input_tensors = input_tensors.to(device) output_tensors = output_tensors.type(torch.LongTensor) output_tensors = output_tensors.to(device) # Zero gradients optimizer.zero_grad() # Begin Forward pass logits = model(input_tensors)[0] print(f"logits.shape: {logits.shape}") print(f"input_tensors.shape: {input_tensors.shape}") print(f"output_tensors.shape: {output_tensors.shape}") # Compute the loss loss = loss_function(logits, output_tensors) # Backward pass loss.backward() # Update the model parameters optimizer.step() # Add the loss to the epoch loss epoch_loss += loss.item() # Print the epoch loss print(f'Epoch {epoch+1}: Loss = {epoch_loss}')
各张量尺寸
- 循环前
input_tensors.shape == torch.Size([2625, 343]) - 循环前
output_tensors.shape == torch.Size([2625, 343]) logits.shape == torch.Size([32, 343, 768])- 循环内
input_tensors.shape == torch.Size([32, 343]) - 循环内
output_tensors.shape == torch.Size([32, 343])
解决方案
核心问题
你使用了GPT2Model而非GPT2LMHeadModel:
GPT2Model输出的是模型的隐藏层特征,维度为[batch_size, seq_len, hidden_size](即你的[32,343,768]),这并不是语言模型任务需要的词汇表概率分布。CrossEntropyLoss默认会将输入的最后一维视为类别数,因此它期望目标张量的形状是[32,768],但你的目标是[32,343](每个位置的token ID),导致形状不匹配。
修改步骤
替换模型为GPT2LMHeadModel
该模型自带语言模型头,会将隐藏层特征映射到词汇表大小的logits(GPT-2词汇表大小为50257),输出维度变为[32,343,50257],与目标的序列长度维度匹配。调整张量形状以适配CrossEntropyLoss
CrossEntropyLoss要求输入形状为[N, C](N为样本数,C为类别数),目标形状为[N],因此需要将序列维度展开。
修改后的关键代码
# 替换模型定义 model = transformers.GPT2LMHeadModel.from_pretrained('gpt2').to(device) # 在Forward pass后调整形状并计算损失 logits = model(input_tensors)[0] # 将logits调整为 [batch*seq_len, vocab_size] logits = logits.view(-1, logits.size(-1)) # 将目标调整为 [batch*seq_len] output_tensors = output_tensors.view(-1) loss = loss_function(logits, output_tensors)
其他注意事项
- 无需再使用
squeeze/unsqueeze调整维度,上述形状调整已经完全适配损失函数要求。 - 保持
ignore_index=0可以忽略填充token(假设你用0作为填充值)的损失计算,符合语言模型训练的常规做法。
内容的提问来源于stack exchange,提问作者C_Dog
相关产品推荐
相关产品推荐

