You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于PyTorch的DNA序列CNN部署遇CUDA设备不匹配报错求助

问题分析与解决方案

报错信息

RuntimeError: Exception occured in Recorder when calling event after_batch:
iter.device(arg).is_cuda() INTERNAL ASSERT FAILED at "../aten/src/ATen/native/cuda/Loops.cuh":89, please report a bug to PyTorch. argument 1: expected a CUDA device but found cpu

问题代码

# import the necessary packages
from torch.nn import Module,Dropout,Conv1d,Linear,MaxPool1d,ReLU,Conv2d,MaxPool2d
import pytorch_lightning as pl
from torch import flatten
from typing import List
class model_test(nn.Module): # deepcre model  
def __init__(self,
             seq_len: int =1000,
             kernel_size: int = 8,
             p = 0.25): # drop out value 
    super().__init__()
    self.seq_len = seq_len
    # adjusting window size corresponding to sequence length 
    window_size = int(seq_len*(8/3000)) # 1000* = 2 
    # CNN module
    self.conv11 = Conv1d(4,64,kernel_size=(kernel_size), padding='same' )
    self.relu11 = ReLU()
    self.conv12 = Conv1d(64,64,kernel_size=kernel_size, padding='same')
    self.relu12 = ReLU()
    self.maxpool1 = MaxPool1d(kernel_size=window_size)
    self.Dropout1 = Dropout(p)
    #self.flatten = flatten()
    self.fc1 =Linear( 64*64*500 , 1)

def forward(self, x):
   # x = xb.permute(0,2,1).unsqueeze(1)
    """Forward pass."""
    print(x.size()) # 512 batch size x 1000 sequence length x 4 channel 
    print(x.dim())
    x = x.permute(0,2,1)
    print( "input : ",x.shape) # 512 batch size x 4 channel x 1000 sequence length 
    x = self.conv11(x)
    print(x.size()) 
    x = self.relu11(x)
    print(x.size()) 
    x = self.conv12(x)
    print("conv1 2nd : ",x.size())
    x = self.relu12(x)
    x = self.maxpool1(x)
    print("maxpool : ",x.size())
    x = self.Dropout1(x)
    x = x.view(batch_size_init, 64*500)
    x= x.view(-1)
    x = self.fc1(x)
    return x

def training_step(self, batch, batch_idx):
    x, y = batch
    y_hat = self.model(x)
    loss = F.cross_entropy(y_hat, y)
    self.log("train_loss", loss, on_step=True, on_epoch=True, prog_bar=True, logger=True)
    return loss

def configure_optimizers(self):
    return torch.optim.Adam(self.parameters(), lr=0.02)

核心问题与修复步骤

1. 模型继承错误

model_test继承了nn.Module,但同时实现了PyTorch Lightning的training_step等方法,导致Lightning无法正确管理设备。需改为继承pl.LightningModule:

class model_test(pl.LightningModule): # 替换原nn.Module

2. 补充缺失的依赖导入

代码中使用了nn、F、torch但未导入,补充以下代码到开头:

import torch
import torch.nn as nn
import torch.nn.functional as F

3. 修正Forward方法中的维度与变量错误

  • 未定义的batch_size_init会引发运行错误,改用动态获取batch size
  • 全连接层fc1的输入维度计算错误,根据maxpool后的维度调整
  • 错误的x.view(-1)会破坏batch维度,需保留batch维度后展平

修改后的Forward方法:

def forward(self, x):
    """Forward pass."""
    x = x.permute(0,2,1)  # 调整维度为[batch, 4, seq_len]
    x = self.conv11(x)
    x = self.relu11(x)
    x = self.conv12(x)
    x = self.relu12(x)
    x = self.maxpool1(x)  # 输出维度为[batch, 64, 500]
    x = self.Dropout1(x)
    batch_size = x.size(0)
    x = x.view(batch_size, -1)  # 动态展平为[batch, 64*500]
    x = self.fc1(x)
    return x

同时修正__init__中的全连接层:

self.fc1 = Linear(64 * 500, 1)  # 修正输入维度

4. 修正Training Step中的模型调用

self.model(x)是错误调用,当前类本身就是模型,需改为直接调用self(x):

def training_step(self, batch, batch_idx):
    x, y = batch
    y_hat = self(x)  # 替换self.model(x)
    loss = F.cross_entropy(y_hat, y)
    self.log("train_loss", loss, on_step=True, on_epoch=True, prog_bar=True, logger=True)
    return loss

5. 设备管理优化

使用PyTorch Lightning时无需手动将数据移到GPU,只需在Trainer初始化时指定GPU:

trainer = pl.Trainer(accelerator="gpu", devices=1, max_epochs=10)
model = model_test()
trainer.fit(model, train_dataloader=train_loader)

总结

以上修复解决了模型继承错误、依赖缺失、维度不匹配、设备管理异常等核心问题,可解决报错并正常运行模型。

内容的提问来源于stack exchange,提问作者Jin_soo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 13:51:18