Keras多GPU训练报错:GPU间资源无法访问导致训练失败求助
尝试使用Keras多GPU训练Transformer模型,代码结构如下:
tf.keras.backend.clear_session() strategy = tf.distribute.MirroredStrategy() print("Number of devices: {}".format(strategy.num_replicas_in_sync)) ...... early_stop = EarlyStopping(monitor='val_accuracy', min_delta=0.001, patience=20, mode='max',restore_best_weights=True) with strategy.scope(): model = Transformer() model.compile( optimizer=optimizers.Adam(1e-4), loss= CategoricalCrossentropy(), metrics=[CategoricalAccuracy()] ) ...... train_dataset = tf.data.Dataset.from_tensor_slices( ({"encoder_input": encoder_input_train, "decoder_input": decoder_input_train}, decoder_output_train)).batch(batch_size) val_dataset = tf.data.Dataset.from_tensor_slices( ({"encoder_input": encoder_input_test, "decoder_input": decoder_input_test}, decoder_output_test)).batch(batch_size) model.fit( train_dataset, batch_size=batch_size, epochs=epochs, verbose=1, callbacks=[early_stop], validation_data=val_dataset )
模型编译正常,但运行model.fit时出现如下报错:
2023-11-27 13:36:26.033383: W tensorflow/core/framework/op_kernel.cc:1780] OP_REQUIRES failed at xla_ops.cc:289 : INVALID_ARGUMENT: Trying to access resource Resource-282-at-0x13338e62a70 (defined @ C:\anaconda3\envs\keras\lib\site-packages\tensorflow\python\ops\gen_resource_variable_ops.py:1226) located in device /job:localhost/replica:0/task:0/device:GPU:0 from device /job:localhost/replica:0/task:0/device:GPU:1 Cf. https://www.tensorflow.org/xla/known_issues#tfvariable_on_a_different_device 2023-11-27 13:36:26.034620: W tensorflow/core/framework/op_kernel.cc:1780] OP_REQUIRES failed at xla_ops.cc:289 : INVALID_ARGUMENT: Trying to access resource Resource-282-at-0x13338e62a70 (defined @ C:\anaconda3\envs\keras\lib\site-packages\tensorflow\python\ops\gen_resource_variable_ops.py:1226) located in device /job:localhost/replica:0/task:0/device:GPU:0 from device /job:localhost/replica:0/task:0/device:GPU:1 Cf. https://www.tensorflow.org/xla/known_issues#tfvariable_on_a_different_device 2023-11-27 13:36:26.036020: W tensorflow/core/framework/op_kernel.cc:1780] OP_REQUIRES failed at xla_ops.cc:289 : INVALID_ARGUMENT: Trying to access resource Resource-282-at-0x13338e62a70 (defined @ C:\anaconda3\envs\keras\lib\site-packages\tensorflow\python\ops\gen_resource_variable_ops.py:1226) located in device /job:localhost/replica:0/task:0/device:GPU:0 from device /job:localhost/replica:0/task:0/device:GPU:1 Cf. https://www.tensorflow.org/xla/known_issues#tfvariable_on_a_different_device 2023-11-27 13:36:26.037382: W tensorflow/core/framework/op_kernel.cc:1780] OP_REQUIRES failed at xla_ops.cc:289 : INVALID_ARGUMENT: Trying to access resource Resource-282-at-0x13338e62a70 (defined @ C:\anaconda3\envs\keras\lib\site-packages\tensorflow\python\ops\gen_resource_variable_ops.py:1226) located in device /job:localhost/replica:0/task:0/device:GPU:0 from device /job:localhost/replica:0/task:0/device:GPU:1 Cf. https://www.tensorflow.org/xla/known_issues#tfvariable_on_a_different_device Node: 'replica_1/StatefulPartitionedCall_64' 5 root error(s) found. (0) INVALID_ARGUMENT: Trying to access resource Resource-282-at-0x13338e62a70 (defined @ C:\anaconda3\envs\keras\lib\site-packages\tensorflow\python\ops\gen_resource_variable_ops.py:1226) located in device /job:localhost/replica:0/task:0/device:GPU:0 from device /job:localhost/replica:0/task:0/device:GPU:1 Cf. https://www.tensorflow.org/xla/known_issues#tfvariable_on_a_different_device [[{{node replica_1/StatefulPartitionedCall_64}}]] [[update_2/AssignAddVariableOp/_855]] (1) INVALID_ARGUMENT: Trying to access resource Resource-282-at-0x13338e62a70 (defined @ C:\anaconda3\envs\keras\lib\site-packages\tensorflow\python\ops\gen_resource_variable_ops.py:1226) located in device /job:localhost/replica:0/task:0/device:GPU:0 from device /job:localhost/replica:0/task:0/device:GPU:1 Cf. https://www.tensorflow.org/xla/known_issues#tfvariable_on_a_different_device [[{{node replica_1/StatefulPartitionedCall_64}}]] [[div_no_nan_1/_847]] (2) INVALID_ARGUMENT: Trying to access resource Resource-282-at-0x13338e62a70 (defined @ C:\anaconda3\envs\keras\lib\site-packages\tensorflow\python\ops\gen_resource_variable_ops.py:1226) located in device /job:localhost/replica:0/task:0/device:GPU:0 from device /job:localhost/replica:0/task:0/device:GPU:1 Cf. https://www.tensorflow.org/xla/known_issues#tfvariable_on_a_different_device [[{{node replica_1/StatefulPartitionedCall_64}}]] [[div_no_nan_1/_843]] (3) INVALID_ARGUMENT: Trying to access resource Resource-282-at-0x13338e62a70 (defined @ C:\anaconda3\envs\keras\lib\site-packages\tensorflow\python\ops\gen_resource_variable_ops.py:1226) located in device /job:localhost/replica:0/task:0/device:GPU:0 from device /job:localhost/replica:0/task:0/device:GPU:1 Cf. https://www.tensorflow.org/xla/known_issues#tfvariable_on_a_different_device [[{{node replica_1/StatefulPartitionedCall_64}}]] [[div_no_nan/ReadVariableOp_2/_796]] (4) INVALID_ARGUMENT: Trying to access resource Resource-282-at-0x13338e62a70 (defined @ C:\anaconda3\envs\keras\lib\site-packages\tensorflow\python\ops\gen_resource_variable_ops.py:1226) located in device /job:localhost/replica:0/task:0/device:GPU:0 from device /job:localhost/replica:0/task:0/device:GPU:1 Cf. https://www.tensorflow.org/xla/known_issues#tfvariable_on_a_different_device [[{{node replica_1/StatefulPartitionedCall_64}}]] 0 successful operations. 0 derived errors ignored. [Op:__inference_train_function_11278]
GPU初始化日志显示已成功识别4块NVIDIA Tesla V100:
2023-11-27 13:35:58.465328: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1616] Created device /job:localhost/replica:0/task:0/device:GPU:0 with 14779 MB memory: -> device: 0, name: NVIDIA Tesla V100-PCIE-16GB, pci bus id: 0000:06:00.0, compute capability: 7.0 2023-11-27 13:35:58.470966: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1616] Created device /job:localhost/replica:0/task:0/device:GPU:1 with 14779 MB memory: -> device: 1, name: NVIDIA Tesla V100-PCIE-16GB, pci bus id: 0000:2f:00.0, compute capability: 7.0 2023-11-27 13:35:58.475054: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1616] Created device /job:localhost/replica:0/task:0/device:GPU:2 with 14779 MB memory: -> device: 2, name: NVIDIA Tesla V100-PCIE-16GB, pci bus id: 0000:86:00.0, compute capability: 7.0 2023-11-27 13:35:58.478613: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1616] Created device /job:localhost/replica:0/task:0/device:GPU:3 with 14779 MB memory: -> device: 3, name: NVIDIA Tesla V100-PCIE-16GB, pci bus id: 0000:d8:00.0, compute capability: 7.0 4 Physical GPUs, 4 Logical GPUs Number of devices: 4
疑问:是否是服务器GPU硬件连接问题?该如何解决?
从报错信息和GPU识别情况来看,大概率是软件层面的资源分配问题,而非硬件连接故障,可按以下步骤排查解决:
确保所有模型组件在策略作用域内创建
检查Transformer类的内部实现,确保所有自定义层、变量、初始化操作都在strategy.scope()的上下文内执行。如果模型中有部分变量(比如自定义的权重、嵌入矩阵)在类初始化时就被创建,且这部分代码不在策略作用域内,就会导致变量被固定在某个GPU上,其他GPU无法访问。关闭XLA加速
报错指向XLA的资源访问问题,可在创建分布式策略前禁用XLA:tf.config.optimizer.set_jit(False) strategy = tf.distribute.MirroredStrategy()也可以通过环境变量设置:
TF_XLA_FLAGS=--tf_xla_enable_xla_devices=false使用策略分发数据集
不要直接对原始数据集做batch操作,改用策略的experimental_distribute_dataset方法处理,确保数据正确分发到各个GPU:train_dataset = tf.data.Dataset.from_tensor_slices( ({"encoder_input": encoder_input_train, "decoder_input": decoder_input_train}, decoder_output_train)).batch(batch_size * strategy.num_replicas_in_sync) val_dataset = tf.data.Dataset.from_tensor_slices( ({"encoder_input": encoder_input_test, "decoder_input": decoder_input_test}, decoder_output_test)).batch(batch_size * strategy.num_replicas_in_sync) # 用策略分发数据集 train_dataset = strategy.experimental_distribute_dataset(train_dataset) val_dataset = strategy.experimental_distribute_dataset(val_dataset)注意:这里的batch_size要乘以GPU数量,因为每个GPU会处理batch_size大小的子批次。
检查自定义层的设备绑定
如果Transformer模型中有自定义层,确保没有手动用tf.device('/GPU:0')这类代码固定设备,所有设备分配交给MirroredStrategy自动管理。验证GPU硬件连接(可选)
若上述方法无效,可通过NVIDIA工具检查GPU拓扑:nvidia-smi topo -m查看GPU之间的连接类型(如NVLink、PCIe),如果存在无法通信的GPU,可能需要检查硬件连线或服务器BIOS设置,但从当前GPU识别正常的情况看,硬件问题概率较低。
内容的提问来源于stack exchange,提问作者chengzhangpei

