You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

CuPy内存管理疑问:used_bytes与total_bytes的差异及超限问题

CuPy内存池used_bytes与total_bytes的区别及内存超限问题解析

我希望理解CuPy的内存管理机制,特别是cupy.cuda.MemoryPool中used_bytes与total_bytes的区别。为此编写了测试代码,分别验证直接GPU数组分配、主机数组转设备两种场景,却发现total_bytes才是触发内存池超限的临界参数,而非直观的used_bytes,对此存在疑惑。

测试代码

import cupy as cp
import numpy as np
import argparse

mempool = cp.get_default_memory_pool()
pinned_mempool = cp.get_default_pinned_memory_pool()

# 手动设置内存池限制
with cp.cuda.Device(0):
   mempool.set_limit(size=39*1024**3)

parser = argparse.ArgumentParser(description='Array size')
parser.add_argument('-x', type=int, help='Size of x dimension')
parser.add_argument('-y', type=int, help='Size of y dimension')
parser.add_argument('-gpufirst', default = False, type=lambda x: (str(x).lower() == 'true'), help='Direct GPU alloction first')
args=parser.parse_args()

B_GB=1024**3

def direct_gpu(x,y):
   # 直接在设备上分配数组
   print("check direct device allocation")
   direct_gpu = cp.arange(x*y).reshape(x,y).astype(cp.float32) 
   print("used:",mempool.used_bytes()/(B_GB), "total:", mempool.total_bytes()/(B_GB), "limit:", mempool.get_limit()/(B_GB),"nfreeblocks:", pinned_mempool.n_free_blocks())
   direct_gpu = None   
 
def direct_cpu(x,y):
   # 在主机内存分配数组后转移到设备
   print("check moving array from host to device")
   direct_cpu = np.arange(x*y).reshape(x,y).astype(np.float32)
   copied_gpu = cp.asarray(direct_cpu)
   print("used:",mempool.used_bytes()/(B_GB), "total:", mempool.total_bytes()/(B_GB), "limit:", mempool.get_limit()/(B_GB),"nfreeblocks:", pinned_mempool.n_free_blocks())
   direct_cpu = None
   direct_gpu = None

print("gpufirst:",args.gpufirst)
if (args.gpufirst):
   direct_gpu(args.x,args.y)
   direct_cpu(args.x,args.y)
else:
   direct_cpu(args.x,args.y)
   direct_gpu(args.x,args.y)

测试结果

第一次测试(x=150000, y=20000)

运行命令:

(venv) [aditya@node01 cupy]$ python example1.py -x 150000 -y 20000 -gpufirst true

输出:

gpufirst: True
check direct device allocation
used: 11.175870895385742 total: 33.52761268615723 limit: 39.0 nfreeblocks: 0
check moving array from host to device
used: 11.175870895385742 total: 33.52761268615723 limit: 39.0 nfreeblocks: 0

可见used_bytes远小于total_bytes,且内存池未触发限制。

第二次测试(x=200000, y=20000)

运行命令:

(venv) [aditya@node01 cupy]$ python example1.py -x 200000 -y 20000 -gpufirst true

输出(触发内存不足错误):

gpufirst: True
check direct device allocation
Traceback (most recent call last):
  File "/home/aditya/Downloads/Tickets/cupy/example1.py", line 39, in <module>
    direct_gpu(args.x,args.y)
  File "/home/aditya/Downloads/Tickets/cupy/example1.py", line 24, in direct_gpu
    direct_gpu = cp.arange(x*y).reshape(x,y).astype(cp.float32) 
                 ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "cupy/_core/core.pyx", line 565, in cupy._core.core._ndarray_base.astype
  File "cupy/_core/core.pyx", line 623, in cupy._core.core._ndarray_base.astype
  File "cupy/_core/core.pyx", line 151, in cupy._core.core.ndarray.__new__
  File "cupy/_core/core.pyx", line 239, in cupy._core.core._ndarray_base._init
  File "cupy/cuda/memory.pyx", line 738, in cupy.cuda.memory.alloc
  File "cupy/cuda/memory.pyx", line 1424, in cupy.cuda.memory.MemoryPool.malloc
  File "cupy/cuda/memory.pyx", line 1445, in cupy.cuda.memory.MemoryPool.malloc
  File "cupy/cuda/memory.pyx", line 1116, in cupy.cuda.memory.SingleDeviceMemoryPool.malloc
  File "cupy/cuda/memory.pyx", line 1137, in cupy.cuda.memory.SingleDeviceMemoryPool._malloc
  File "cupy/cuda/memory.pyx", line 1344, in cupy.cuda.memory.SingleDeviceMemoryPool._try_malloc
  File "cupy/cuda/memory.pyx", line 1356, in cupy.cuda.memory.SingleDeviceMemoryPool._try_malloc
cupy.cuda.memory.OutOfMemoryError: Out of memory allocating 16,000,000,000 bytes (allocated so far: 32,000,000,000 bytes, limit set to: 41,875,931,136 bytes).

此时仅将x维度从150000增至200000(used_bytes仅增加33%),却触发了内存不足错误。

核心概念解析

used_bytes与total_bytes的区别

  • used_bytes:当前被用户代码实际占用的GPU内存字节数,即活跃的CuPy数组所占用的内存。当数组被回收(如赋值为None),这部分内存会被标记为空闲,但不会立即还给GPU驱动,而是留在内存池中复用。
  • total_bytes:内存池从GPU驱动已申请到的总内存字节数,包括当前被使用的内存,加上内存池中已被标记为空闲但尚未还给驱动的内存。这是内存池实际占用的GPU总资源。

为什么total_bytes是内存池超限的临界参数

CuPy内存池的设计目的是减少频繁向GPU驱动申请/释放内存的开销,因此会将空闲内存保留在池中。内存池的限制(set_limit设置的值)是针对池从驱动申请的总内存,而非用户实际使用的内存。当total_bytes加上新申请的内存超过限制时,就会触发OutOfMemoryError——因为内存池不会突破自己设定的总申请上限,即使池中还有空闲内存,但若新申请后总占用超过限制,也会拒绝分配。

问题原因分析

在测试中:

  1. 第一次分配时,cp.arange(...).astype(cp.float32)的过程会产生临时数组:arange创建整数数组,astype生成新的float32数组,这两个数组在内存池中都会占用空间,直到临时数组被回收。内存池为了容纳这些临时内存,向驱动申请了远大于最终used_bytes的total_bytes。
  2. 增大数组维度后,新的分配需要更大的内存块,此时内存池已申请的total_bytes加上新申请的内存超过了39GB的限制,因此触发错误。即使used_bytes增量不大,但内存池无法突破总申请上限,也无法向驱动申请更多内存,所以报错。

总结

  • used_bytes是用户当前活跃数据占用的内存,total_bytes是内存池从GPU驱动申请的总内存(含空闲)。
  • 内存池的限制针对的是total_bytes,因为它控制的是内存池向GPU驱动申请的总资源量,避免过度占用GPU内存。
  • 临时数组的创建会导致total_bytes远大于used_bytes,这是内存池为了复用内存而产生的正常现象,但也可能导致在看似内存足够的情况下触发超限错误。

内容的提问来源于stack exchange,提问作者Aditya

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.13 03:58:13