使用PyTorch TensorDataset类时遭遇‘1D target tensor expected’错误的原因排查、独热编码原理及解决方法咨询
Hey there! Let's break down what's going on here and get that error fixed step by step.
Why the Error Happens
First, let's look at the shape of your target tensor. When you read labels from your CSV using targets_file.iloc[1:1001].values, you end up with a 2D numpy array (shape (1000, 1)—each label is in its own row, wrapped in an extra dimension). Converting this directly to a PyTorch tensor keeps that 2D structure, so your targets tensor ends up as torch.Size([1000, 1]).
The problem lies with PyTorch's CrossEntropyLoss (which I assume you're using later in your pipeline): it expects 1D target tensors (shape (1000,)) where each element is a single class index. It interprets 2D targets as multi-label classification (even if each row only has one value), which isn't what you need here.
TensorDataset isn't the issue—it's just passing along the tensors you provide. The root cause is that extra dimension in your target tensor.
A Quick Breakdown of One-Hot Encoding
You asked about one-hot encoding, so let's clarify this key point:
- One-hot encoding converts a class index (like
2for a 3-class problem) into a binary vector (e.g.,[0, 0, 1]). - You don't need one-hot encoded targets for CrossEntropyLoss in PyTorch. The loss function internally converts 1D class indices to one-hot vectors under the hood to compute the loss. If you accidentally pass a one-hot encoded 2D tensor, you'll get the exact same "multi-target not supported" error—so stick with 1D class indices!
How to Fix the Issue
You just need to remove that extra dimension from your target tensor. Here are a few simple ways to do it:
Option 1: Use squeeze() when creating the tensor
Modify your target tensor line to:
targets = torch.tensor(targets_file.iloc[1:1001].values, dtype=torch.long).squeeze()
squeeze()removes any dimensions of size 1, turning(1000,1)into(1000,).- We set
dtype=torch.longbecause CrossEntropyLoss expects integer-type targets (class indices should be integers, not floats).
Option 2: Flatten the numpy array first
You can flatten the values before converting to a tensor:
targets = torch.tensor(targets_file.iloc[1:1001].values.flatten(), dtype=torch.long)
Option 3: Extract the column directly as a 1D array
If your labels are in a single named column (e.g., label), pull it as a 1D series:
targets = torch.tensor(targets_file.iloc[1:1001]['label'].values, dtype=torch.long)
Verify the Fix
After making the change, check your target tensor's shape to confirm it's 1D:
print(targets.shape) # Should output torch.Size([1000])
Once your targets are in the correct shape, CrossEntropyLoss will work seamlessly with your TensorDataset and DataLoaders.
内容的提问来源于stack exchange,提问作者ClarenceHD

