You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用GPU训练Keras模型时Python内核崩溃问题求助

Docker环境下TensorFlow GPU训练触发Python内核崩溃问题

操作背景

按TensorFlow官方文档指引完成基础安装后,自定义Dockerfile构建包含Jupyter Lab、TensorFlow及个人常用工具包的镜像,操作步骤如下:

  • 从NVIDIA官网下载驱动安装包,直接安装NVIDIA驱动
  • 禁用nouveau旧驱动以适配老GPU
  • 安装nvidia-container-toolkit
  • 编写自定义Dockerfile和docker-compose配置文件

已通过命令sudo docker run --rm --runtime=nvidia --gpus all ubuntu nvidia-smi验证安装有效性,该命令可正常返回容器内NVIDIA驱动及CUDA版本信息。

环境配置详情

Dockerfile内容

FROM tensorflow/tensorflow:latest-gpu

WORKDIR /tf

# 安装Jupyter导出文件所需依赖
RUN apt-get update && apt-get upgrade -y && \
    apt-get install texlive \
    texlive-latex-extra \
    texlive-xetex \
    texlive-fonts-recommended \
    texlive-plain-generic \
    pandoc -y

# 更新pip至最新版本
RUN pip install --upgrade pip

# 安装数据科学相关Python包
RUN pip install ... 

ENTRYPOINT ["/usr/local/bin/jupyter", "lab", "--ip=0.0.0.0", "--no-browser", "--allow-root", "--notebook-dir=/tf/"]

docker-compose配置

version: "3"
services:
  jupyter:
    build: .
    image: jupyter_tf_gpu_custom:latest
    container_name: jupyter-tf-gpu
    restart: unless-stopped
    labels:
      . . . traefik相关配置
    volumes:
      - jupyter_data:/tf/
    networks:
      - traefik-network
    deploy:
      resources:
        reservations:
          devices:
          - driver: nvidia
            count: all
            capabilities: [gpu]

networks:
  traefik-network:
    external: true

volumes:
  jupyter_data:

核心软件版本

  • TensorFlow:2.13.1
  • NVIDIA驱动:535.129.03(对应CUDA版本12.2)

服务器硬件与系统

  • 操作系统:Debian (bullseye)
  • GPU:Quadro RTX 4000
  • 内存:32GB RAM
  • CPU:6核/12线程Xeon

问题描述

Docker内Jupyter Lab中执行nvidia-smi可正常显示GPU、驱动及CUDA版本,但调用fit()训练任意规模Sequential模型时,Python内核直接崩溃。相同代码在CPU环境下可正常运行。

TensorFlow运行时警告

successful NUMA node read from SysFS had negative value (-1), but there must be at least one NUMA node, so returning NUMA node zero

已尝试的解决方法

  1. 配置GPU内存动态增长:
import tensorflow as tf
gpus = tf.config.experimental.list_physical_devices('GPU')
for gpu in gpus:
  tf.config.experimental.set_memory_growth(gpu, True)
  1. 安装tensorflow[and-cuda],此前曾出现警告:Attempting to register factory for plugin cuDNN when one has already been registered
  2. 更换过多个版本的NVIDIA驱动:530、470、370

内容的提问来源于stack exchange,提问作者Noe Guedet

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.07 08:56:13