CNN训练循环无法终止问题求助
问题:CNN训练循环无法按指定迭代次数停止
我尝试构建一个CNN并使用训练模块对其进行训练,希望指定训练的迭代次数,但发现训练循环持续运行无法停止。请问有人能帮我解决这个问题吗?
def train(model, epochs=10): optimiser = torch.optim.SGD(model.parameters(), lr=0.001) writer = SummaryWriter() batch_idx = 0 loss_total = 0 epoch = 0 for epoch in range(epochs): print('range:', range(epochs)) for batch in train_loader: features, labels = batch prediction = model(features) # cf = confusion_matrix(labels, prediction) loss = F.cross_entropy(prediction, labels) # Loss model changes label size loss_total += loss.item() loss.backward() print('loss:', loss.item()) optimiser.step() optimiser.zero_grad() writer.add_scalar('Loss', loss.item(), batch_idx) batch_idx += 1 print('epoch', epoch) epoch += 1 # why does this not stop??? print('Total loss:', loss_total/batch_idx)
完整代码可参考作者GitHub仓库中的CNN.py文件。
问题原因与解决方法
问题出在内层batch循环中手动执行了epoch += 1:
- 外层
for epoch in range(epochs)循环的逻辑是,每次迭代从range(epochs)序列中依次取值赋值给epoch,遍历完所有值后循环自动停止。 - 但你在处理每个batch时手动递增
epoch,直接打乱了外层循环的计数逻辑,导致epoch的值永远无法达到epochs设定的次数,循环也就停不下来。
解决步骤:
- 删除内层batch循环里的
epoch += 1语句 - 开头的
epoch = 0初始化可以去掉,因为外层for循环会自动为epoch赋值
修改后的核心代码片段:
def train(model, epochs=10): optimiser = torch.optim.SGD(model.parameters(), lr=0.001) writer = SummaryWriter() batch_idx = 0 loss_total = 0 for epoch in range(epochs): print('range:', range(epochs)) for batch in train_loader: features, labels = batch prediction = model(features) # cf = confusion_matrix(labels, prediction) loss = F.cross_entropy(prediction, labels) # Loss model changes label size loss_total += loss.item() loss.backward() print('loss:', loss.item()) optimiser.step() optimiser.zero_grad() writer.add_scalar('Loss', loss.item(), batch_idx) batch_idx += 1 print('epoch', epoch) print('Total loss:', loss_total/batch_idx)
内容的提问来源于stack exchange,提问作者Rory Merz
相关产品推荐
相关产品推荐

