You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

M1 Pro上TensorFlow Metal加速训练平台未找到错误求助

解决Mac上TensorFlow Metal训练时"could not find registered platform with id"错误

问题背景

在MacOS 13.1系统中,使用TensorFlow 2.11.0 + TensorFlow Metal 0.7.0加速训练模型时,调用.fit()方法持续抛出NOT_FOUND: could not find registered platform with id错误,通过with tf.device('/CPU:0'):指定CPU训练则可正常运行,已尝试重建conda环境但问题未解决。环境配置如下:

  • Python: 3.10.8
  • TensorFlow: 2.11.0
  • TensorFlow Metal: 0.7.0
  • conda环境配置:
name: tf-metal
channels:
  - apple
  - conda-forge
dependencies:
  - python
  - pip
  - tensorflow-deps
  - ipykernel

  - pip:
    - tensorflow-macos
    - tensorflow-metal  

解决方案

1. 匹配TensorFlow与TensorFlow Metal版本兼容性

TensorFlow与TensorFlow Metal存在严格的版本对应关系,TensorFlow 2.11.0需要搭配TensorFlow Metal 0.8.0而非0.7.0。执行以下命令升级:

pip install --upgrade tensorflow-metal==0.8.0

2. 禁用XLA编译

错误堆栈指向XLA相关操作,临时禁用XLA编译可规避该问题,在训练代码开头添加:

import tensorflow as tf
tf.config.optimizer.set_jit(False)

3. 替换实验性优化器

错误源于keras.optimizers.optimizer_experimental.optimizer的XLA更新步骤,实验性优化器与Metal兼容性不佳,将代码中的实验性优化器替换为稳定版:

  • 替换前:tf.keras.optimizers.experimental.Adam(learning_rate=0.001)
  • 替换后:tf.keras.optimizers.Adam(learning_rate=0.001)

4. 强制指定GPU设备

确保TensorFlow正确识别Metal设备,在代码开头添加设备检测与配置:

import tensorflow as tf
# 打印所有可用设备,确认GPU被识别
print(tf.config.list_physical_devices())
# 强制设置GPU为可见设备
gpus = tf.config.list_physical_devices('GPU')
if gpus:
    tf.config.set_visible_devices(gpus[0], 'GPU')

内容的提问来源于stack exchange,提问作者qtaro

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 04:55:44