TensorFlow多GPU大矩阵分片运算:如何顺序执行规避内存错误
解决TensorFlow多GPU分片计算时单GPU并行调度导致的OOM问题
首先咱们明确核心问题:TensorFlow的执行引擎会自动分析张量依赖关系,对无依赖的操作尽可能并行调度。你GPU:0上的操作1、3、5之间没有显式约束,所以会被同时启动,瞬间占满单卡显存触发OOM错误。下面是几个兼顾顺序执行和低性能损耗的可行方案:
方案1:显式添加依赖控制GPU操作顺序(最优推荐)
通过tf.control_dependencies强制GPU:0上的操作按顺序执行,同时不影响两块GPU之间的并行计算(比如GPU:0的操作1和GPU:1的操作2可并行,操作3和操作4可并行)。
修改后的完整代码如下:
import matplotlib.pyplot as plt import time import tensorflow as tf import numpy as np from tensorflow.python.client import timeline tf.set_random_seed(42) T = tf.constant(1, tf.float32, name='Maturity') N = tf.constant(500, tf.int32, name='Issuer') sigma = tf.random.uniform([N], 0, 1) riskfreeRate = 0.02 S0 = tf.random.uniform([N], 5, 10) simulation = 10000 scenarios = 1000000 # 补充定义原代码缺失的scenarios变量 S = tf.placeholder(tf.float32, name='StartingPrice') z = tf.random_normal([N, simulation], dtype=tf.float32, name='standardNormRandomValues') mu = tf.placeholder(tf.float32, name='riskfreeRate') ST = S0 * tf.exp((mu - tf.square(sigma)/2) * T + sigma * tf.transpose(z) * tf.sqrt(T)) # 创建示例500×1000000规模的矩阵A A = np.random.rand(500, 1000000) A2 = tf.convert_to_tensor(np.asmatrix(A), dtype=tf.float32) # 操作1:GPU:0上的第一个计算 with tf.device('/device:GPU:0'): A3 = A2[:, 0:tf.cast((scenarios/5), tf.int32)] SC1 = tf.matmul(ST, A3) # 操作2:GPU:1上的计算,可与操作1并行 with tf.device('/device:GPU:1'): A4 = A2[:, tf.cast((scenarios/5), tf.int32):tf.cast(2*scenarios/5, tf.int32)] SC2 = tf.matmul(ST, A4) # 操作3:GPU:0上的第二个计算,必须等SC1完成后执行 with tf.device('/device:GPU:0'): with tf.control_dependencies([SC1]): A5 = A2[:, tf.cast((2*scenarios/5), tf.int32):tf.cast(3*scenarios/5, tf.int32)] SC3 = tf.matmul(ST, A5) # 操作4:GPU:1上的计算,可与操作3并行 with tf.device('/device:GPU:1'): A6 = A2[:, tf.cast((3*scenarios/5), tf.int32):tf.cast(4*scenarios/5, tf.int32)] SC4 = tf.matmul(ST, A6) # 操作5:GPU:0上的第三个计算,必须等SC3完成后执行 with tf.device('/device:GPU:0'): with tf.control_dependencies([SC3]): A7 = A2[:, tf.cast((4*scenarios/5), tf.int32):tf.cast(scenarios, tf.int32)] SC5 = tf.matmul(ST, A7) # 修正concat的参数格式(原代码存在语法错误) SC = tf.concat([SC1, SC2, SC3, SC4, SC5], axis=1) # 开启显存增长模式,进一步降低OOM风险 config = tf.ConfigProto() config.gpu_options.allow_growth = True with tf.Session(config=config) as sess: priceTensor = sess.run(SC, {mu: riskfreeRate})
方案说明:
tf.control_dependencies([SC1])会强制上下文内的操作(SC3)等待SC1计算完成后再执行,同理SC5等待SC3完成,这样GPU:0上的三个矩阵乘法会依次执行,前一个完成释放显存后再启动下一个。- 两块GPU的操作仍能并行(比如操作2和操作1同时运行),不会浪费多GPU的性能。
allow_growth配置让TensorFlow按需分配显存,避免预占满全部GPU内存,进一步降低OOM概率。
方案2:分步执行会话调用(备选)
如果不想修改计算图依赖关系,可以将GPU:0的操作分成多次会话调用,每次只执行一个GPU:0的操作,同时搭配GPU:1的操作并行。不过这个方法会增加GPU到CPU的数据传输开销,性能损耗比方案1大。
示例代码片段:
# (计算图定义同方案1,无需添加control_dependencies) config = tf.ConfigProto() config.gpu_options.allow_growth = True with tf.Session(config=config) as sess: # 第一步:并行执行GPU0的SC1和GPU1的SC2 sc1, sc2 = sess.run([SC1, SC2], {mu: riskfreeRate}) # 第二步:并行执行GPU0的SC3和GPU1的SC4 sc3, sc4 = sess.run([SC3, SC4], {mu: riskfreeRate}) # 第三步:执行GPU0的SC5 sc5 = sess.run(SC5, {mu: riskfreeRate}) # 拼接最终结果 priceTensor = np.concatenate([sc1, sc2, sc3, sc4, sc5], axis=1)
关于你提到的“遍历设备”思路
单纯遍历设备执行操作并不能解决并行调度问题——同一个GPU上的无依赖操作,TensorFlow仍会尝试并行启动。必须配合显式依赖控制(方案1)或分步会话调用(方案2),才能让同一GPU上的操作顺序执行。
内容的提问来源于stack exchange,提问作者sount
相关产品推荐
相关产品推荐

