使用tf.distribute.MirroredStrategy仅单GPU工作的问题排查求助
问题
使用tensorflow-gpu 1.13.1,配备2块RTX2080-Ti显卡,尝试通过tf.distribute.MirroredStrategy()实现双GPU训练,但实际仅一块GPU工作,另一块完全闲置。
训练主循环代码:
for A in range(0, outer_loop): load_database = Load_database # <Load database> for B in range(0, inner_loop): X_train, X_valid, Y_train, Y_valid = load_data # <Load Data> self = Model(hyperparam) model = self.create() # <create model> epoch = 0 for epoch in range(max_epoch): hist = model.fit(X_train, Y_train, epochs=1, batch_size=200, verbose=1) loss, accuracy = model.evaluate(x=X_train, y=Y_train) Y_train_predict = model.predict(x=X_train)
模型创建代码:
class Model(): def create_model(self, x): # <configure layer parameters> def create(self): self.num_classes = num_classes # Create graph and session tf.compat.v1.reset_default_graph() gpu_options = tf.compat.v1.GPUOptions() config = tf.compat.v1.ConfigProto(log_device_placement=False, gpu_options=gpu_options) config.gpu_options.allow_growth = True tf.compat.v1.disable_eager_execution() self.graph = tf.compat.v1.get_default_graph() self.sess = tf.compat.v1.Session(config=config) with self.graph.as_default(): with self.sess.as_default(): with tf.compat.v1.variable_scope(tf.compat.v1.get_variable_scope()): x = Input(shape=(800, 6, 1), name='X') self.y_ = self.create_model(x) self.scope() strategy = tf.distribute.MirroredStrategy(devices=["/gpu:0","/gpu:1"], cross_device_ops=tf.contrib.distribute.AllReduceCrossDeviceOps(all_reduce_alg="hierarchical_copy")) with strategy.scope(): model = tf.keras.Model(inputs=[x], outputs=[self.y_]) model.compile(loss='categorical_crossentropy', optimizer=self.optimizer, metrics=[self.metrics]) return model
GPU状态信息:
+-----------------------------------------------------------------------------+ | NVIDIA-SMI 525.60.13 Driver Version: 525.60.13 CUDA Version: 12.0 | |-------------------------------+----------------------+----------------------+ | GPU Name Persistence-M| Bus-Id Disp.A | Volatile Uncorr. ECC | | Fan Temp Perf Pwr:Usage/Cap| Memory-Usage | GPU-Util Compute M. | | | | MIG M. | |===============================+======================+======================| | 0 NVIDIA GeForce ... On | 00000000:0A:00.0 Off | N/A | | 43% 84C P2 225W / 250W | 10915MiB / 11264MiB | 90% Default | | | | N/A | +-------------------------------+----------------------+----------------------+ | 1 NVIDIA GeForce ... On | 00000000:42:00.0 Off | N/A | | 28% 39C P8 17W / 250W | 161MiB / 11264MiB | 0% Default | | | | N/A | +-------------------------------+----------------------+----------------------+
曾尝试用strategy.scope()包裹整个主训练循环,但无效果。
原因分析与解决方案
核心问题
- 模型构建时机错误:输入层
x和模型结构self.y_ = self.create_model(x)是在strategy.scope()之外创建的,这部分计算图只会绑定到默认GPU,无法被分布式策略接管。 - 自定义Session干扰:手动创建的
tf.Session会覆盖分布式策略自动管理的设备分配逻辑,导致策略无法将计算任务分发到第二块GPU。 - TF1.x策略兼容性限制:TensorFlow 1.13的
MirroredStrategy对Keras的支持要求所有模型相关操作(输入定义、层创建、实例化、编译)必须在策略作用域内完成,否则无法触发多GPU分发。
修正步骤
1. 移除自定义Session和Graph管理
TF1.x的分布式策略会自动处理会话和计算图的生命周期,手动创建tf.Session和tf.Graph会破坏策略的设备分发逻辑,直接删除这部分代码即可。
2. 将所有模型构建逻辑移入strategy.scope()内
输入层定义、模型结构创建、模型实例化、编译操作必须全部在策略作用域中执行,确保所有计算节点被正确分发到两块GPU。
3. 调整Batch Size(可选)
使用双GPU时,设置的batch_size会被自动拆分到各个GPU(每个GPU处理batch_size/2的数据),可以适当调大batch_size以充分利用双GPU的计算能力。
修正后的模型创建代码
class Model(): def create_model(self, x): # <configure layer parameters> # 示例:替换为你的实际层定义 x = tf.keras.layers.Conv2D(32, (3,3), activation='relu')(x) x = tf.keras.layers.GlobalAveragePooling2D()(x) output = tf.keras.layers.Dense(self.num_classes, activation='softmax')(x) return output def create(self): self.num_classes = num_classes tf.compat.v1.disable_eager_execution() # 初始化分布式策略 strategy = tf.distribute.MirroredStrategy( devices=["/gpu:0","/gpu:1"], cross_device_ops=tf.contrib.distribute.AllReduceCrossDeviceOps(all_reduce_alg="hierarchical_copy") ) # 所有模型相关操作必须在strategy.scope()内执行 with strategy.scope(): x = tf.keras.Input(shape=(800, 6, 1), name='X') self.y_ = self.create_model(x) model = tf.keras.Model(inputs=[x], outputs=[self.y_]) model.compile( loss='categorical_crossentropy', optimizer=self.optimizer, metrics=[self.metrics] ) return model
额外验证建议
- 开启设备放置日志:将
tf.compat.v1.ConfigProto的log_device_placement=True,可以查看每个操作被分配到哪个GPU,验证分布式策略是否生效。 - 检查数据格式:确保训练数据是numpy数组或
tf.data.Dataset格式,TF1.13对这两种格式的分布式支持都兼容,推荐使用tf.data以获得更好的性能。
内容的提问来源于stack exchange,提问作者hson
相关产品推荐
相关产品推荐

