You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

TensorFlow数组操作内存泄漏求助:大尺寸网格计算内存溢出

TensorFlow大张量内存溢出问题求助

需求与实现逻辑

我需要在TensorFlow中构建一个(500,500,500)的边际网格,边际化逻辑为:先拆分生成多个(20,500,500,500)网格,在axis=0维度做求和边际化后存入列表,最后对列表再次做边际化求和。示例代码如下:

import numpy as np
import tensorflow as tf


logrho_List = np.linspace(4.6989, 9.6987, 500)
logrho_List_partition = logrho_List.reshape((25,20))

example_array =  tf.random.uniform((500, 500, 500), dtype=tf.float32)

Result = [0]*25
pos_min = [0]*25

for i in range(0,25):
    iresult = [0]*20
    for j in range(0, 20):
        iresult[j] = (10**logrho_List_partition[i,j])*example_array
    
    iresult = tf.convert_to_tensor(iresult) 
     
    iresult = tf.exp( - iresult)
    
    Result[i] = tf.reduce_sum(iresult, axis = 0)
    
    del iresult
    

Result = tf.convert_to_tensor(Result)
Result = tf.reduce_sum(Result, axis = 0)

实际场景中iresult[j]是同维度数组求和后与数值相乘得到的(500,500,500) float32数组,示例代码仅保留核心逻辑。

问题现象

循环无法执行完成,第13次迭代时TensorFlow因内存不足崩溃。通过以下代码监控内存:

process = psutil.Process(os.getpid())
memory_used = process.memory_info().rss / (1024 ** 3)  
print(f"Memory used by process: {memory_used:.2f} GB")

发现崩溃时内存占用约20GB,但我配备了120GB内存和48GB显存的RTX A6000显卡,无法理解内存占用过高的原因。

已尝试的优化措施

我尝试了常规内存优化方案,但问题仍未解决:

  • 脚本开头添加GPU内存配置:
os.environ['TF_GPU_ALLOCATOR'] = 'cuda_malloc_async'

physical_devices = tf.config.list_physical_devices('GPU')
try:
    tf.config.experimental.set_memory_growth(physical_devices[0], True)
except:
# Invalid device or cannot modify virtual devices once initialized.
    pass
  • 每次外层循环结束后显式清理会话和回收垃圾:
# Explicitly clear the session and collect garbage to free memory
tf.keras.backend.clear_session()
gc.collect()

内容的提问来源于stack exchange,提问作者Felipe Augusto

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.17 02:30:04