PyTorch声学模型训练报错:期望Long类型却得到Int
Hey there, let's get that acoustic model training properly. The error you're seeing happens because PyTorch's CrossEntropyLoss requires target labels to be of type torch.long (64-bit integers), but your current label data is stored as torch.int—a type mismatch that breaks the loss calculation. Below are the step-by-step fixes, plus some other important tweaks to your code to avoid further issues:
1. Fix the Label Type Mismatch
The quick fix is to convert your batch labels to long type when calculating the loss. Update your training loop's loss line to:
loss = criterion(outputs, y_train.long())
For a more permanent solution, you can modify your Dataloader to return labels as long type directly (e.g., adding labels = labels.long() when loading your data) so you don't have to convert it every batch.
2. Fix the Model Definition's Missing Parameter
Your DNN class references hidden1_size in its layers, but you didn't include it as a parameter in the __init__ method—this will cause an error when instantiating the model. Plus, the final layer's ReLU activation is a mistake for classification with CrossEntropyLoss (the loss function expects raw logits, not activated outputs). Here's the corrected model:
class DNN(nn.Module): def __init__(self, input_size, hidden1_size, hidden2_size, hidden3_size, output_size): super(DNN, self).__init__() self.fc1 = nn.Linear(input_size, hidden1_size) self.relu1 = nn.ReLU() self.fc2 = nn.Linear(hidden1_size, hidden2_size) self.relu2 = nn.ReLU() self.fc3 = nn.Linear(hidden2_size, hidden3_size) self.relu3 = nn.ReLU() self.fc4 = nn.Linear(hidden3_size, output_size) # Remove final ReLU: CrossEntropyLoss expects raw logits def forward(self, x): out = self.fc1(x) out = self.relu1(out) out = self.fc2(out) out = self.relu2(out) out = self.fc3(out) out = self.relu3(out) out = self.fc4(out) # No final ReLU here return out
Then update the model instantiation to include the missing hidden1_size parameter:
model = DNN(input_size, hidden1_size, hidden2_size, hidden3_size, output_size)
3. Fix Training/Testing Loop Logic Errors
Your current code uses the global labels variable instead of the batch-specific y_train/y_test when calculating accuracy, which will give completely wrong results. Also, you should switch the model to evaluation mode during testing to disable training-specific behaviors like dropout. Here's the corrected loop:
# Train the network iter = 0 for epoch in range(no_epochs): for i, (X_train, y_train) in enumerate(train_loader): optimizer.zero_grad() outputs = model(X_train) # Convert label to long type loss = criterion(outputs, y_train.long()) loss.backward() optimizer.step() iter += 1 if iter % 500 == 0: correct = 0 total = 0 # Switch to evaluation mode model.eval() # Disable gradient calculation to save memory with torch.no_grad(): for X_test, y_test in test_loader: outputs = model(X_test) _, predicted = torch.max(outputs.data, 1) total += y_test.size(0) # Match label type for comparison correct += (predicted == y_test.long()).sum().item() # Switch back to training mode model.train() accuracy = 100 * correct / total # Use modern PyTorch syntax for loss value print(f"Iteration: {iter}, Loss: {loss.item()}, Accuracy: {accuracy:.2f}%")
Key Takeaways
- The original error is solved by converting labels to
torch.longtype - Removing the final ReLU ensures your model outputs valid logits for
CrossEntropyLoss - Fixing the loop logic ensures you're calculating accuracy correctly on your test batches
内容的提问来源于stack exchange,提问作者Mohamed Nabih Ali

