GTX1660Ti环境下Jupyter Notebook中TensorFlow2.10调用GPU问题[已解决]
问题:使用ktrain训练图像模型时误判TensorFlow未调用GPU
环境配置
- Python v3.10.6
- TensorFlow v2.10.1
- CUDA v12.0
- CuDNN v8.1
- GPU:NVIDIA GTX 1660Ti(支持CUDA)
问题现象
在Jupyter Notebook中使用ktrain训练图像识别模型时,观察到CPU占用率飙升,误以为TensorFlow未调用GPU,耗时近2天排查。
已执行的排查与尝试
1. GPU可用性验证
运行以下代码检测GPU状态:
import tensorflow as tf print(tf.config.list_physical_devices('GPU')) print("Num GPUs Available: ", len(tf.config.list_physical_devices('GPU'))) sess = tf.compat.v1.Session(config=tf.compat.v1.ConfigProto(log_device_placement=True)) from tensorflow.python.client import device_lib print(device_lib.list_local_devices()) tf.debugging.set_log_device_placement(True) # Create some tensors a = tf.constant([[1.0, 2.0, 3.0], [4.0, 5.0, 6.0]]) b = tf.constant([[1.0, 2.0], [3.0, 4.0], [5.0, 6.0]]) c = tf.matmul(a, b) print(c)
输出明确显示GPU设备被正常识别且可用:
[PhysicalDevice(name='/physical_device:GPU:0', device_type='GPU')] Num GPUs Available: 1 Device mapping: /job:localhost/replica:0/task:0/device:GPU:0 -> device: 0, name: NVIDIA GeForce GTX 1660 Ti with Max-Q Design, pci bus id: 0000:01:00.0, compute capability: 7.5 [name: "/device:CPU:0" device_type: "CPU" memory_limit: 268435456 locality { } incarnation: 9175052246053814955 xla_global_id: -1 , name: "/device:GPU:0" device_type: "GPU" memory_limit: 4163895296 locality { bus_id: 1 links { } } incarnation: 3853720177835357027 physical_device_desc: "device: 0, name: NVIDIA GeForce GTX 1660 Ti with Max-Q Design, pci bus id: 0000:01:00.0, compute capability: 7.5" xla_global_id: 416903419 ] tf.Tensor( [[22. 28.] [49. 64.]], shape=(2, 2), dtype=float32)
2. 无效的解决尝试
- 设置CUDA环境变量:
import os os.environ["CUDA_DEVICE_ORDER"]="PCI_BUS_ID" os.environ["CUDA_VISIBLE_DEVICES"]="0" import pandas as pd import ktrain as kt from ktrain import vision as vis
- 显式指定GPU设备执行训练:
with tf.device('/device:GPU:0'): learner.autofit(1e-4)
- 更换TensorFlow和Python旧版本
- 未尝试在Jupyter Notebook外部运行代码
3. 模型配置代码
model = vis.image_classifier('pretrained_resnet50', train_data, val_data) learner = kt.get_learner(model=model, train_data=train_data, val_data=val_data, workers=4, use_multiprocessing=False, batch_size=16)
根因分析
GPU实际一直在正常参与训练,误判源于CPU占用率较高,以及对GPU负载的预期偏差。
解决方案
训练过程中,在命令行执行nvidia-smi命令,若输出显示GPU内存几乎被占满,即可确认GPU正在正常工作。
内容的提问来源于stack exchange,提问作者raitrow
相关产品推荐
相关产品推荐

