You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

TensorFlow多GPU大矩阵分片运算:如何顺序执行规避内存错误

解决TensorFlow多GPU分片计算时单GPU并行调度导致的OOM问题

首先咱们明确核心问题:TensorFlow的执行引擎会自动分析张量依赖关系,对无依赖的操作尽可能并行调度。你GPU:0上的操作1、3、5之间没有显式约束,所以会被同时启动,瞬间占满单卡显存触发OOM错误。下面是几个兼顾顺序执行和低性能损耗的可行方案:

方案1:显式添加依赖控制GPU操作顺序(最优推荐)

通过tf.control_dependencies强制GPU:0上的操作按顺序执行,同时不影响两块GPU之间的并行计算(比如GPU:0的操作1和GPU:1的操作2可并行,操作3和操作4可并行)。

修改后的完整代码如下:

import matplotlib.pyplot as plt
import time
import tensorflow as tf
import numpy as np
from tensorflow.python.client import timeline

tf.set_random_seed(42)
T = tf.constant(1, tf.float32, name='Maturity')
N = tf.constant(500, tf.int32, name='Issuer')
sigma = tf.random.uniform([N], 0, 1)
riskfreeRate = 0.02
S0 = tf.random.uniform([N], 5, 10)
simulation = 10000
scenarios = 1000000  # 补充定义原代码缺失的scenarios变量

S = tf.placeholder(tf.float32, name='StartingPrice')
z = tf.random_normal([N, simulation], dtype=tf.float32, name='standardNormRandomValues')
mu = tf.placeholder(tf.float32, name='riskfreeRate')
ST = S0 * tf.exp((mu - tf.square(sigma)/2) * T + sigma * tf.transpose(z) * tf.sqrt(T))

# 创建示例500×1000000规模的矩阵A
A = np.random.rand(500, 1000000)
A2 = tf.convert_to_tensor(np.asmatrix(A), dtype=tf.float32)

# 操作1:GPU:0上的第一个计算
with tf.device('/device:GPU:0'):
    A3 = A2[:, 0:tf.cast((scenarios/5), tf.int32)]
    SC1 = tf.matmul(ST, A3)

# 操作2:GPU:1上的计算,可与操作1并行
with tf.device('/device:GPU:1'):
    A4 = A2[:, tf.cast((scenarios/5), tf.int32):tf.cast(2*scenarios/5, tf.int32)]
    SC2 = tf.matmul(ST, A4)

# 操作3:GPU:0上的第二个计算,必须等SC1完成后执行
with tf.device('/device:GPU:0'):
    with tf.control_dependencies([SC1]):
        A5 = A2[:, tf.cast((2*scenarios/5), tf.int32):tf.cast(3*scenarios/5, tf.int32)]
        SC3 = tf.matmul(ST, A5)

# 操作4:GPU:1上的计算,可与操作3并行
with tf.device('/device:GPU:1'):
    A6 = A2[:, tf.cast((3*scenarios/5), tf.int32):tf.cast(4*scenarios/5, tf.int32)]
    SC4 = tf.matmul(ST, A6)

# 操作5:GPU:0上的第三个计算,必须等SC3完成后执行
with tf.device('/device:GPU:0'):
    with tf.control_dependencies([SC3]):
        A7 = A2[:, tf.cast((4*scenarios/5), tf.int32):tf.cast(scenarios, tf.int32)]
        SC5 = tf.matmul(ST, A7)

# 修正concat的参数格式(原代码存在语法错误)
SC = tf.concat([SC1, SC2, SC3, SC4, SC5], axis=1)

# 开启显存增长模式,进一步降低OOM风险
config = tf.ConfigProto()
config.gpu_options.allow_growth = True
with tf.Session(config=config) as sess:
    priceTensor = sess.run(SC, {mu: riskfreeRate})

方案说明:

  • tf.control_dependencies([SC1])会强制上下文内的操作(SC3)等待SC1计算完成后再执行,同理SC5等待SC3完成,这样GPU:0上的三个矩阵乘法会依次执行,前一个完成释放显存后再启动下一个。
  • 两块GPU的操作仍能并行(比如操作2和操作1同时运行),不会浪费多GPU的性能。
  • allow_growth配置让TensorFlow按需分配显存,避免预占满全部GPU内存,进一步降低OOM概率。

方案2:分步执行会话调用(备选)

如果不想修改计算图依赖关系,可以将GPU:0的操作分成多次会话调用,每次只执行一个GPU:0的操作,同时搭配GPU:1的操作并行。不过这个方法会增加GPU到CPU的数据传输开销,性能损耗比方案1大。

示例代码片段:

# (计算图定义同方案1,无需添加control_dependencies)

config = tf.ConfigProto()
config.gpu_options.allow_growth = True
with tf.Session(config=config) as sess:
    # 第一步:并行执行GPU0的SC1和GPU1的SC2
    sc1, sc2 = sess.run([SC1, SC2], {mu: riskfreeRate})
    # 第二步:并行执行GPU0的SC3和GPU1的SC4
    sc3, sc4 = sess.run([SC3, SC4], {mu: riskfreeRate})
    # 第三步:执行GPU0的SC5
    sc5 = sess.run(SC5, {mu: riskfreeRate})
    # 拼接最终结果
    priceTensor = np.concatenate([sc1, sc2, sc3, sc4, sc5], axis=1)

关于你提到的“遍历设备”思路

单纯遍历设备执行操作并不能解决并行调度问题——同一个GPU上的无依赖操作,TensorFlow仍会尝试并行启动。必须配合显式依赖控制(方案1)或分步会话调用(方案2),才能让同一GPU上的操作顺序执行。

内容的提问来源于stack exchange,提问作者sount

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 08:57:27