GPU训练速度慢于CPU的问题求助(含PyTorch与CUDA配置信息)
问题:RTX4070Ti Super训练PyTorch模型速度慢于CPU的排查求助
在20k×20的数据集上训练4层PyTorch模型,CPU耗时约120分钟,GPU反而耗时130分钟。训练时通过任务管理器监测到GPU利用率可达94%,确认GPU在运行,但速度不如CPU,求排查方向。
环境与操作信息
1. PyTorch安装方式
使用以下conda命令安装:
conda install pytorch torchvision torchaudio pytorch-cuda=12.1 -c pytorch -c nvidia
2. CUDA版本异常情况
最初安装了CUDA 12.3,通过系统添加/删除程序和conda uninstall cuda卸载后重装了12.1版本,但打印版本信息时nvcc仍显示12.3。不过torch.cuda.is_available()返回true,具体版本输出如下:
print('__Python VERSION:', sys.version) print('__pyTorch VERSION:', torch.version) print('__CUDA VERSION', ) __Python VERSION: 3.11.8 | packaged by Anaconda, Inc. | (main, Feb 26 2024, 21:34:05) [MSC v.1916 64 bit (AMD64)] __pyTorch VERSION: 2.2.1 __CUDA VERSION nvcc: NVIDIA (R) Cuda compiler driver Copyright (c) 2005-2023 NVIDIA Corporation Built on Wed_Nov_22_10:30:42_Pacific_Standard_Time_2023 Cuda compilation tools, release 12.3, V12.3.107 Build cuda_12.3.r12.3/compiler.33567101_0
3. 设备迁移代码
已将数据和模型移至CUDA:
X_train_tensor_gpu = X_train_tensor.to('cuda') model = Net() model.to('cuda')
4. nvidia-smi输出
Wed Mar 6 14:59:19 2024 ±----------------------------------------------------------------------------------------+ | NVIDIA-SMI 551.61 Driver Version: 551.61 CUDA Version: 12.4 | |-----------------------------------------±-----------------------±---------------------+ | GPU Name TCC/WDDM | Bus-Id Disp.A | Volatile Uncorr. ECC | | Fan Temp Perf Pwr:Usage/Cap | Memory-Usage | GPU-Util Compute M. | | | | MIG M. | |=========================================+========================+======================| | 0 NVIDIA GeForce RTX 4070 … WDDM | 00000000:01:00.0 On | N/A | | 0% 40C P8 20W / 285W | 1746MiB / 16376MiB | 6% Default | | | | N/A | ±----------------------------------------±-----------------------±---------------------+
内容的提问来源于stack exchange,提问作者Jayden Choi
相关产品推荐
相关产品推荐

