使用GPU训练Keras模型时Python内核崩溃问题求助
Docker环境下TensorFlow GPU训练触发Python内核崩溃问题
操作背景
按TensorFlow官方文档指引完成基础安装后,自定义Dockerfile构建包含Jupyter Lab、TensorFlow及个人常用工具包的镜像,操作步骤如下:
- 从NVIDIA官网下载驱动安装包,直接安装NVIDIA驱动
- 禁用nouveau旧驱动以适配老GPU
- 安装nvidia-container-toolkit
- 编写自定义Dockerfile和docker-compose配置文件
已通过命令sudo docker run --rm --runtime=nvidia --gpus all ubuntu nvidia-smi验证安装有效性,该命令可正常返回容器内NVIDIA驱动及CUDA版本信息。
环境配置详情
Dockerfile内容
FROM tensorflow/tensorflow:latest-gpu WORKDIR /tf # 安装Jupyter导出文件所需依赖 RUN apt-get update && apt-get upgrade -y && \ apt-get install texlive \ texlive-latex-extra \ texlive-xetex \ texlive-fonts-recommended \ texlive-plain-generic \ pandoc -y # 更新pip至最新版本 RUN pip install --upgrade pip # 安装数据科学相关Python包 RUN pip install ... ENTRYPOINT ["/usr/local/bin/jupyter", "lab", "--ip=0.0.0.0", "--no-browser", "--allow-root", "--notebook-dir=/tf/"]
docker-compose配置
version: "3" services: jupyter: build: . image: jupyter_tf_gpu_custom:latest container_name: jupyter-tf-gpu restart: unless-stopped labels: . . . traefik相关配置 volumes: - jupyter_data:/tf/ networks: - traefik-network deploy: resources: reservations: devices: - driver: nvidia count: all capabilities: [gpu] networks: traefik-network: external: true volumes: jupyter_data:
核心软件版本
- TensorFlow:2.13.1
- NVIDIA驱动:535.129.03(对应CUDA版本12.2)
服务器硬件与系统
- 操作系统:Debian (bullseye)
- GPU:Quadro RTX 4000
- 内存:32GB RAM
- CPU:6核/12线程Xeon
问题描述
Docker内Jupyter Lab中执行nvidia-smi可正常显示GPU、驱动及CUDA版本,但调用fit()训练任意规模Sequential模型时,Python内核直接崩溃。相同代码在CPU环境下可正常运行。
TensorFlow运行时警告
successful NUMA node read from SysFS had negative value (-1), but there must be at least one NUMA node, so returning NUMA node zero
已尝试的解决方法
- 配置GPU内存动态增长:
import tensorflow as tf gpus = tf.config.experimental.list_physical_devices('GPU') for gpu in gpus: tf.config.experimental.set_memory_growth(gpu, True)
- 安装
tensorflow[and-cuda],此前曾出现警告:Attempting to register factory for plugin cuDNN when one has already been registered - 更换过多个版本的NVIDIA驱动:530、470、370
内容的提问来源于stack exchange,提问作者Noe Guedet
相关产品推荐
相关产品推荐

