You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

NVIDIA GPU在Python环境下显存分配极慢问题求助

NVIDIA GPU在Python环境下显存分配极慢问题求助

大家好,我最近遇到了一个非常棘手的问题——在Python环境中使用NVIDIA GPU时,显存分配速度慢得离谱,实在搞不定了,来求助各位大佬!

具体情况是这样的:每次在全新的Python会话里运行GPU计算任务时,不管是用TensorFlow还是PyTorch,它们都会以极小的增量一点点分配显存,这个过程要持续足足四分钟左右,之后才会突然分配一大块显存,然后才开始真正执行计算。但奇怪的是,第一次计算完成后,后续的所有计算都能瞬间完成,完全没有延迟。

我想问问大家,有没有人遇到过类似的问题?或者有没有办法获取显存分配过程中的详细日志,看看后台到底在执行什么操作,导致这个漫长的等待?

我已经尝试过重新安装CUDA库和NVIDIA驱动了,重装驱动后问题会暂时消失,但过一阵子又会复发,显存分配慢的情况又回来了。

下面是我的相关环境和输出信息:

Python环境及测试输出

Python 3.11.3 (main, Apr  5 2023, 14:15:06) [GCC 9.4.0] on linux

Type "help", "copyright", "credits" or "license" for more information.

>>> import timeit

>>> timeit.timeit('import tensorflow as tf;tf.random.uniform([10])', number=1)

2023-04-17 09:08:24.062130: I tensorflow/core/platform/cpu_feature_guard.cc:182] This TensorFlow binary is optimized to use available CPU instructions in performance-critical operations.

To enable the following instructions: AVX2 AVX512F FMA, in other operations, rebuild TensorFlow with the appropriate compiler flags.

2023-04-17 09:08:24.641429: W tensorflow/compiler/tf2tensorrt/utils/py_utils.cc:38] TF-TRT Warning: Could not find TensorRT

2023-04-17 09:12:12.879503: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1635] Created device /job:localhost/replica:0/task:0/device:GPU:0 with 21368 MB memory:  -> device: 0, name: GRID RTX6000-24Q, pci bus id: 0000:02:02.0, compute capability: 7.5

229.68861908599501

nvidia-smi输出

+-----------------------------------------------------------------------------+
| NVIDIA-SMI 470.182.03   Driver Version: 470.182.03   CUDA Version: 11.4     |
|-------------------------------+----------------------+----------------------+
| GPU  Name        Persistence-M| Bus-Id        Disp.A | Volatile Uncorr. ECC |
| Fan  Temp  Perf  Pwr:Usage/Cap|         Memory-Usage | GPU-Util  Compute M. |
|                               |                      |               MIG M. |
|===============================+======================+======================|
|   0  GRID RTX6000-24Q    On   | 00000000:02:02.0 Off |                  N/A |
| N/A   N/A    P8    N/A /  N/A |  23527MiB / 24576MiB |      0%      Default |
|                               |                      |                  N/A |
+-------------------------------+----------------------+----------------------+
+-----------------------------------------------------------------------------+
| Processes:                                                                  |
|  GPU   GI   CI        PID   Type   Process name                  GPU Memory |
|        ID   ID                                                   Usage      |
|=============================================================================|
|    0   N/A  N/A    122079      C   ...Model-js4zUkog/bin/python    21743MiB |
+-----------------------------------------------------------------------------+

nvcc版本输出

nvcc: NVIDIA (R) Cuda compiler driver
Copyright (c) 2005-2022 NVIDIA Corporation
Built on Wed_Sep_21_10:33:58_PDT_2022
Cuda compilation tools, release 11.8, V11.8.89
Build cuda_11.8.r11.8/compiler.31833905_0

备注:内容来源于stack exchange,提问作者petrovski

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.22 15:34:36