TensorFlow数组操作内存泄漏求助:大尺寸网格计算内存溢出
TensorFlow大张量内存溢出问题求助
需求与实现逻辑
我需要在TensorFlow中构建一个(500,500,500)的边际网格,边际化逻辑为:先拆分生成多个(20,500,500,500)网格,在axis=0维度做求和边际化后存入列表,最后对列表再次做边际化求和。示例代码如下:
import numpy as np import tensorflow as tf logrho_List = np.linspace(4.6989, 9.6987, 500) logrho_List_partition = logrho_List.reshape((25,20)) example_array = tf.random.uniform((500, 500, 500), dtype=tf.float32) Result = [0]*25 pos_min = [0]*25 for i in range(0,25): iresult = [0]*20 for j in range(0, 20): iresult[j] = (10**logrho_List_partition[i,j])*example_array iresult = tf.convert_to_tensor(iresult) iresult = tf.exp( - iresult) Result[i] = tf.reduce_sum(iresult, axis = 0) del iresult Result = tf.convert_to_tensor(Result) Result = tf.reduce_sum(Result, axis = 0)
实际场景中iresult[j]是同维度数组求和后与数值相乘得到的(500,500,500) float32数组,示例代码仅保留核心逻辑。
问题现象
循环无法执行完成,第13次迭代时TensorFlow因内存不足崩溃。通过以下代码监控内存:
process = psutil.Process(os.getpid()) memory_used = process.memory_info().rss / (1024 ** 3) print(f"Memory used by process: {memory_used:.2f} GB")
发现崩溃时内存占用约20GB,但我配备了120GB内存和48GB显存的RTX A6000显卡,无法理解内存占用过高的原因。
已尝试的优化措施
我尝试了常规内存优化方案,但问题仍未解决:
- 脚本开头添加GPU内存配置:
os.environ['TF_GPU_ALLOCATOR'] = 'cuda_malloc_async' physical_devices = tf.config.list_physical_devices('GPU') try: tf.config.experimental.set_memory_growth(physical_devices[0], True) except: # Invalid device or cannot modify virtual devices once initialized. pass
- 每次外层循环结束后显式清理会话和回收垃圾:
# Explicitly clear the session and collect garbage to free memory tf.keras.backend.clear_session() gc.collect()
内容的提问来源于stack exchange,提问作者Felipe Augusto
相关产品推荐
相关产品推荐

