PyTorch训练报错:shape '[-1,14]'与输入尺寸491520不匹配
Hey there, let's break down this issue and get it sorted out!
The Root Cause
Your error stems from a simple math mismatch in tensor reshaping. The view() operation in PyTorch requires the total number of elements in the tensor to stay identical before and after reshaping. Let's crunch the numbers:
- Your
logitstensor has a shape of(32, 20, 768), so total elements are32 * 20 * 768 = 491520 - You’re trying to reshape it to
(-1, 14)— but491520 ÷ 14 ≈ 35108.57, which isn’t an integer. PyTorch can’t split a tensor into a non-integer number of chunks, hence the RuntimeError.
What You’re Missing
It looks like you’re working on a sequence-based task (like token-level classification) with 14 output labels. The problem is your model’s output is a 768-dimensional hidden state per token, not the 14-dimensional logits you need for classification. You need an extra linear layer to map that 768-dimensional space down to your 14 target labels.
Solution 1: Token-Level Classification (predict labels for every token)
If you want to assign a label to each of the 20 tokens in your sequence:
- Add a linear layer to convert 768-dim hidden states to 14-dim logits:
import torch.nn as nn # Define the classifier layer (ideally add this to your model's __init__ method) classifier = nn.Linear(768, num_labels) # num_labels=14 # Process the model outputs to get proper logits logits = classifier(outputs[0]) # Now logits shape is (32, 20, 14)
- Now reshape and calculate loss — total elements are
32*20*14=8960, which divides evenly by 14:
loss_func = nn.BCEWithLogitsLoss() loss = loss_func( logits.view(-1, num_labels), b_labels.type_as(logits).view(-1, num_labels) )
Solution 2: Sequence-Level Classification (single label per entire sequence)
If you only need one label per full sequence (not per token), first aggregate the sequence's hidden states (e.g., take the<[BOS_never_used_51bce0c785ca2f68081bfa7d91973934]> token output, or average all tokens):
# Option 1: Use the<[BOS_never_used_51bce0c785ca2f68081bfa7d91973934]> token's output (standard in transformer models) cls_token_output = outputs[0][:, 0, :] # Shape becomes (32, 768) # Option 2: Average all token outputs across the sequence avg_token_output = outputs[0].mean(dim=1) # Shape also (32, 768) # Map to 14 labels classifier = nn.Linear(768, num_labels) logits = classifier(cls_token_output) # Shape is (32, 14) # Calculate loss without reshaping (shapes match directly) loss_func = nn.BCEWithLogitsLoss() loss = loss_func(logits, b_labels.type_as(logits))
内容的提问来源于stack exchange,提问作者samer hassan

