You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PyTorch训练报错:shape '[-1,14]'与输入尺寸491520不匹配

Fixing RuntimeError: shape '[-1, 14]' is invalid for input of size 491520

Hey there, let's break down this issue and get it sorted out!

The Root Cause

Your error stems from a simple math mismatch in tensor reshaping. The view() operation in PyTorch requires the total number of elements in the tensor to stay identical before and after reshaping. Let's crunch the numbers:

  • Your logits tensor has a shape of (32, 20, 768), so total elements are 32 * 20 * 768 = 491520
  • You’re trying to reshape it to (-1, 14) — but 491520 ÷ 14 ≈ 35108.57, which isn’t an integer. PyTorch can’t split a tensor into a non-integer number of chunks, hence the RuntimeError.

What You’re Missing

It looks like you’re working on a sequence-based task (like token-level classification) with 14 output labels. The problem is your model’s output is a 768-dimensional hidden state per token, not the 14-dimensional logits you need for classification. You need an extra linear layer to map that 768-dimensional space down to your 14 target labels.

Solution 1: Token-Level Classification (predict labels for every token)

If you want to assign a label to each of the 20 tokens in your sequence:

  1. Add a linear layer to convert 768-dim hidden states to 14-dim logits:
import torch.nn as nn

# Define the classifier layer (ideally add this to your model's __init__ method)
classifier = nn.Linear(768, num_labels)  # num_labels=14

# Process the model outputs to get proper logits
logits = classifier(outputs[0])  # Now logits shape is (32, 20, 14)
  1. Now reshape and calculate loss — total elements are 32*20*14=8960, which divides evenly by 14:
loss_func = nn.BCEWithLogitsLoss()
loss = loss_func(
    logits.view(-1, num_labels),
    b_labels.type_as(logits).view(-1, num_labels)
)

Solution 2: Sequence-Level Classification (single label per entire sequence)

If you only need one label per full sequence (not per token), first aggregate the sequence's hidden states (e.g., take the<[BOS_never_used_51bce0c785ca2f68081bfa7d91973934]> token output, or average all tokens):

# Option 1: Use the<[BOS_never_used_51bce0c785ca2f68081bfa7d91973934]> token's output (standard in transformer models)
cls_token_output = outputs[0][:, 0, :]  # Shape becomes (32, 768)

# Option 2: Average all token outputs across the sequence
avg_token_output = outputs[0].mean(dim=1)  # Shape also (32, 768)

# Map to 14 labels
classifier = nn.Linear(768, num_labels)
logits = classifier(cls_token_output)  # Shape is (32, 14)

# Calculate loss without reshaping (shapes match directly)
loss_func = nn.BCEWithLogitsLoss()
loss = loss_func(logits, b_labels.type_as(logits))

内容的提问来源于stack exchange,提问作者samer hassan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 17:42:35