You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Keras多GPU训练报错:GPU间资源无法访问导致训练失败求助

问题描述

尝试使用Keras多GPU训练Transformer模型,代码结构如下:

tf.keras.backend.clear_session()
strategy = tf.distribute.MirroredStrategy()
print("Number of devices: {}".format(strategy.num_replicas_in_sync))
......
early_stop = EarlyStopping(monitor='val_accuracy', min_delta=0.001, patience=20, mode='max',restore_best_weights=True)
with strategy.scope():
    model = Transformer()
    model.compile(
        optimizer=optimizers.Adam(1e-4),
        loss= CategoricalCrossentropy(),
        metrics=[CategoricalAccuracy()]
    )
......
train_dataset = tf.data.Dataset.from_tensor_slices(
        ({"encoder_input": encoder_input_train, "decoder_input": decoder_input_train},
         decoder_output_train)).batch(batch_size)

val_dataset = tf.data.Dataset.from_tensor_slices(
        ({"encoder_input": encoder_input_test, "decoder_input": decoder_input_test},
         decoder_output_test)).batch(batch_size)

model.fit(
    train_dataset,
    batch_size=batch_size,
    epochs=epochs,
    verbose=1,
    callbacks=[early_stop],
    validation_data=val_dataset
)

模型编译正常,但运行model.fit时出现如下报错:

2023-11-27 13:36:26.033383: W tensorflow/core/framework/op_kernel.cc:1780] OP_REQUIRES failed at xla_ops.cc:289 : INVALID_ARGUMENT: Trying to access resource Resource-282-at-0x13338e62a70 (defined @ C:\anaconda3\envs\keras\lib\site-packages\tensorflow\python\ops\gen_resource_variable_ops.py:1226) located in device /job:localhost/replica:0/task:0/device:GPU:0 from device /job:localhost/replica:0/task:0/device:GPU:1
 Cf. https://www.tensorflow.org/xla/known_issues#tfvariable_on_a_different_device
2023-11-27 13:36:26.034620: W tensorflow/core/framework/op_kernel.cc:1780] OP_REQUIRES failed at xla_ops.cc:289 : INVALID_ARGUMENT: Trying to access resource Resource-282-at-0x13338e62a70 (defined @ C:\anaconda3\envs\keras\lib\site-packages\tensorflow\python\ops\gen_resource_variable_ops.py:1226) located in device /job:localhost/replica:0/task:0/device:GPU:0 from device /job:localhost/replica:0/task:0/device:GPU:1
 Cf. https://www.tensorflow.org/xla/known_issues#tfvariable_on_a_different_device
2023-11-27 13:36:26.036020: W tensorflow/core/framework/op_kernel.cc:1780] OP_REQUIRES failed at xla_ops.cc:289 : INVALID_ARGUMENT: Trying to access resource Resource-282-at-0x13338e62a70 (defined @ C:\anaconda3\envs\keras\lib\site-packages\tensorflow\python\ops\gen_resource_variable_ops.py:1226) located in device /job:localhost/replica:0/task:0/device:GPU:0 from device /job:localhost/replica:0/task:0/device:GPU:1
 Cf. https://www.tensorflow.org/xla/known_issues#tfvariable_on_a_different_device
2023-11-27 13:36:26.037382: W tensorflow/core/framework/op_kernel.cc:1780] OP_REQUIRES failed at xla_ops.cc:289 : INVALID_ARGUMENT: Trying to access resource Resource-282-at-0x13338e62a70 (defined @ C:\anaconda3\envs\keras\lib\site-packages\tensorflow\python\ops\gen_resource_variable_ops.py:1226) located in device /job:localhost/replica:0/task:0/device:GPU:0 from device /job:localhost/replica:0/task:0/device:GPU:1
 Cf. https://www.tensorflow.org/xla/known_issues#tfvariable_on_a_different_device
Node: 'replica_1/StatefulPartitionedCall_64'
5 root error(s) found.
  (0) INVALID_ARGUMENT:  Trying to access resource Resource-282-at-0x13338e62a70 (defined @ C:\anaconda3\envs\keras\lib\site-packages\tensorflow\python\ops\gen_resource_variable_ops.py:1226) located in device /job:localhost/replica:0/task:0/device:GPU:0 from device /job:localhost/replica:0/task:0/device:GPU:1
 Cf. https://www.tensorflow.org/xla/known_issues#tfvariable_on_a_different_device
     [[{{node replica_1/StatefulPartitionedCall_64}}]]
     [[update_2/AssignAddVariableOp/_855]]
  (1) INVALID_ARGUMENT:  Trying to access resource Resource-282-at-0x13338e62a70 (defined @ C:\anaconda3\envs\keras\lib\site-packages\tensorflow\python\ops\gen_resource_variable_ops.py:1226) located in device /job:localhost/replica:0/task:0/device:GPU:0 from device /job:localhost/replica:0/task:0/device:GPU:1
 Cf. https://www.tensorflow.org/xla/known_issues#tfvariable_on_a_different_device
     [[{{node replica_1/StatefulPartitionedCall_64}}]]
     [[div_no_nan_1/_847]]
  (2) INVALID_ARGUMENT:  Trying to access resource Resource-282-at-0x13338e62a70 (defined @ C:\anaconda3\envs\keras\lib\site-packages\tensorflow\python\ops\gen_resource_variable_ops.py:1226) located in device /job:localhost/replica:0/task:0/device:GPU:0 from device /job:localhost/replica:0/task:0/device:GPU:1
 Cf. https://www.tensorflow.org/xla/known_issues#tfvariable_on_a_different_device
     [[{{node replica_1/StatefulPartitionedCall_64}}]]
     [[div_no_nan_1/_843]]
  (3) INVALID_ARGUMENT:  Trying to access resource Resource-282-at-0x13338e62a70 (defined @ C:\anaconda3\envs\keras\lib\site-packages\tensorflow\python\ops\gen_resource_variable_ops.py:1226) located in device /job:localhost/replica:0/task:0/device:GPU:0 from device /job:localhost/replica:0/task:0/device:GPU:1
 Cf. https://www.tensorflow.org/xla/known_issues#tfvariable_on_a_different_device
     [[{{node replica_1/StatefulPartitionedCall_64}}]]
     [[div_no_nan/ReadVariableOp_2/_796]]
  (4) INVALID_ARGUMENT:  Trying to access resource Resource-282-at-0x13338e62a70 (defined @ C:\anaconda3\envs\keras\lib\site-packages\tensorflow\python\ops\gen_resource_variable_ops.py:1226) located in device /job:localhost/replica:0/task:0/device:GPU:0 from device /job:localhost/replica:0/task:0/device:GPU:1
 Cf. https://www.tensorflow.org/xla/known_issues#tfvariable_on_a_different_device
     [[{{node replica_1/StatefulPartitionedCall_64}}]]
0 successful operations.
0 derived errors ignored. [Op:__inference_train_function_11278]

GPU初始化日志显示已成功识别4块NVIDIA Tesla V100:

2023-11-27 13:35:58.465328: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1616] Created device /job:localhost/replica:0/task:0/device:GPU:0 with 14779 MB memory:  -> device: 0, name: NVIDIA Tesla V100-PCIE-16GB, pci bus id: 0000:06:00.0, compute capability: 7.0
2023-11-27 13:35:58.470966: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1616] Created device /job:localhost/replica:0/task:0/device:GPU:1 with 14779 MB memory:  -> device: 1, name: NVIDIA Tesla V100-PCIE-16GB, pci bus id: 0000:2f:00.0, compute capability: 7.0
2023-11-27 13:35:58.475054: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1616] Created device /job:localhost/replica:0/task:0/device:GPU:2 with 14779 MB memory:  -> device: 2, name: NVIDIA Tesla V100-PCIE-16GB, pci bus id: 0000:86:00.0, compute capability: 7.0
2023-11-27 13:35:58.478613: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1616] Created device /job:localhost/replica:0/task:0/device:GPU:3 with 14779 MB memory:  -> device: 3, name: NVIDIA Tesla V100-PCIE-16GB, pci bus id: 0000:d8:00.0, compute capability: 7.0
4 Physical GPUs, 4 Logical GPUs
Number of devices: 4

疑问:是否是服务器GPU硬件连接问题?该如何解决?

解决方案

从报错信息和GPU识别情况来看,大概率是软件层面的资源分配问题,而非硬件连接故障,可按以下步骤排查解决:

  • 确保所有模型组件在策略作用域内创建
    检查Transformer类的内部实现,确保所有自定义层、变量、初始化操作都在strategy.scope()的上下文内执行。如果模型中有部分变量(比如自定义的权重、嵌入矩阵)在类初始化时就被创建,且这部分代码不在策略作用域内,就会导致变量被固定在某个GPU上,其他GPU无法访问。

  • 关闭XLA加速
    报错指向XLA的资源访问问题,可在创建分布式策略前禁用XLA:

    tf.config.optimizer.set_jit(False)
    strategy = tf.distribute.MirroredStrategy()
    

    也可以通过环境变量设置:TF_XLA_FLAGS=--tf_xla_enable_xla_devices=false

  • 使用策略分发数据集
    不要直接对原始数据集做batch操作,改用策略的experimental_distribute_dataset方法处理,确保数据正确分发到各个GPU:

    train_dataset = tf.data.Dataset.from_tensor_slices(
        ({"encoder_input": encoder_input_train, "decoder_input": decoder_input_train},
         decoder_output_train)).batch(batch_size * strategy.num_replicas_in_sync)
    val_dataset = tf.data.Dataset.from_tensor_slices(
        ({"encoder_input": encoder_input_test, "decoder_input": decoder_input_test},
         decoder_output_test)).batch(batch_size * strategy.num_replicas_in_sync)
    
    # 用策略分发数据集
    train_dataset = strategy.experimental_distribute_dataset(train_dataset)
    val_dataset = strategy.experimental_distribute_dataset(val_dataset)
    

    注意:这里的batch_size要乘以GPU数量,因为每个GPU会处理batch_size大小的子批次。

  • 检查自定义层的设备绑定
    如果Transformer模型中有自定义层,确保没有手动用tf.device('/GPU:0')这类代码固定设备,所有设备分配交给MirroredStrategy自动管理。

  • 验证GPU硬件连接(可选)
    若上述方法无效,可通过NVIDIA工具检查GPU拓扑:

    nvidia-smi topo -m
    

    查看GPU之间的连接类型(如NVLink、PCIe),如果存在无法通信的GPU,可能需要检查硬件连线或服务器BIOS设置,但从当前GPU识别正常的情况看,硬件问题概率较低。

内容的提问来源于stack exchange,提问作者chengzhangpei

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.05 13:10:53