CNN模型训练迭代报错:张量维度不匹配问题求助
解决CNN训练中测试准确率计算的RuntimeError问题
我在训练CNN模型时,用nn.CrossEntropyLoss()计算损失,optim.SGD作为优化器。但在训练迭代中计算测试准确率时遇到RuntimeError,错误信息如下:
RuntimeError: The size of tensor a (128) must match the size of tensor b (16) at non-singleton dimension 0
相关训练代码
损失函数与优化器定义
criterion = nn.CrossEntropyLoss() optimizer = optim.SGD(net.parameters(), lr=0.001, momentum=0.9)
训练主循环
epochs = 10 epoch_log = [] loss_log = [] accuracy_log = [] for epoch in range(epochs): print(f'starting epoch : {epoch+1}...') running_loss = 0.0 for i, data in enumerate(trainloader, 0): inputs, labels = data # move our data to GPU inputs = inputs.to(device) labels = labels.to(device) #set the gradients to zero optimizer.zero_grad() # Forward -> backprop + optimize outputs = net(inputs) loss = criterion(outputs, labels) loss.backward() optimizer.step() running_loss += loss.item() if i % 50 == 49: correct = 0 total = 0 with torch.no_grad(): for data in testloader: images, labels = data images = images.to(device) labels = labels.to(device) outputs = net(inputs) _, predicted = torch.max(outputs.data, dim = 1) total += labels.size(0) correct += (predicted == labels).sum().item() accuracy = 100 * correct / total epoch_num = epoch + 1 actual_loss = running_loss / 50 print(f"Epoch : {epoch_num}, mini-batches completed : {(i+1)}, Loss : {actual_loss:.3f}, Test Accuracy : {accuracy:.3f}%") running_loss = 0.0 # store training stats after each epoch epoch_log.append(epoch_nmum) loss_log.append(actual_loss) accuracy_log.append(accuracy) print("Training Completed")
完整错误栈
starting epoch : 1... --------------------------------------------------------------------------- RuntimeError Traceback (most recent call last) <ipython-input-21-b3f8854281bf> in <module> 44 45 total += labels.size(0) ---> 46 correct += (predicted == labels).sum().item() 47 48 accuracy = 100 * correct / total RuntimeError: The size of tensor a (128) must match the size of tensor b (16) at non-singleton dimension 0
问题原因与解决方法
核心问题
测试循环里犯了低级错误:用训练批次的inputs喂模型,而非当前测试批次的images。
训练循环中inputs是当前训练批次的张量(尺寸128),测试循环里每个批次的labels尺寸为16,两者维度不匹配,导致predicted == labels比较时触发维度不匹配错误。
修复代码
把测试循环内的outputs = net(inputs)改为outputs = net(images):
with torch.no_grad(): for data in testloader: images, labels = data images = images.to(device) labels = labels.to(device) # 替换成测试批次数据推理 outputs = net(images) _, predicted = torch.max(outputs.data, dim = 1) total += labels.size(0) correct += (predicted == labels).sum().item()
额外注意事项
- 训练循环末尾的
epoch_nmum是笔误,应改为epoch_num,否则会报未定义变量错误 - 统计日志的代码缩进错误,需放在外层epoch循环内,否则只会记录最后一次的统计值
内容的提问来源于stack exchange,提问作者Dhruv Jani
相关产品推荐
相关产品推荐

