Apple M2 Pro运行TensorFlow测试脚本出现总线错误求助
问题描述
在Apple M2 Pro设备上运行TensorFlow测试脚本时触发总线错误(zsh: bus error)。
测试脚本代码:
import tensorflow as tf cifar = tf.keras.datasets.cifar100 (x_train, y_train), (x_test, y_test) = cifar.load_data() model = tf.keras.applications.ResNet50( include_top=True, weights=None, input_shape=(32, 32, 3), classes=100,) loss_fn = tf.keras.losses.SparseCategoricalCrossentropy(from_logits=True) model.compile(optimizer="adam", loss=loss_fn, metrics=["accuracy"]) model.fit(x_train, y_train, epochs=5, batch_size=4)
终端输出信息:
Metal device set to: Apple M2 Pro systemMemory: 16.00 GB maxCacheSize: 5.33 GB 2023-03-23 00:26:32.203361: I tensorflow/core/common_runtime/pluggable_device/pluggable_device_factory.cc:305] Could not identify NUMA node of platform GPU ID 0, defaulting to 0. Your kernel may not have been built with NUMA support. 2023-03-23 00:26:32.203521: I tensorflow/core/common_runtime/pluggable_device/pluggable_device_factory.cc:271] Created TensorFlow device (/job:localhost/replica:0/task:0/device:GPU:0 with 0 MB memory) -> physical PluggableDevice (device: 0, name: METAL, pci bus id: <undefined>) zsh: bus error python3 app/model/tf_verify.py
故障原因分析
- TensorFlow Metal插件兼容性bug:Apple Silicon设备上的TensorFlow依赖
tensorflow-metal实现GPU加速,早期版本的插件在处理ResNet50这类模型时,存在底层内存访问或指令集适配问题,直接触发总线错误。 - 输入尺寸不匹配:ResNet50原生设计的输入尺寸为224×224,强行传入32×32的CIFAR-100数据,会导致模型内部部分层的张量计算出现内存越界,引发总线错误。
- GPU内存识别异常:终端输出显示GPU设备内存为0MB,说明TensorFlow对Metal设备的内存识别存在问题,进而导致内存分配失败触发总线错误。
解决办法
- 升级TensorFlow及Metal插件:安装最新版的
tensorflow-macos和tensorflow-metal,新版本会修复Apple Silicon平台的兼容性问题。执行以下命令更新:pip install --upgrade tensorflow-macos tensorflow-metal - 调整输入尺寸至模型适配规格:将CIFAR-100的32×32图片resize为ResNet50适配的224×224,修改脚本如下:
import tensorflow as tf cifar = tf.keras.datasets.cifar100 (x_train, y_train), (x_test, y_test) = cifar.load_data() # 调整输入尺寸到224×224 x_train = tf.image.resize(x_train, (224, 224)) x_test = tf.image.resize(x_test, (224, 224)) model = tf.keras.applications.ResNet50( include_top=True, weights=None, input_shape=(224, 224, 3), classes=100,) loss_fn = tf.keras.losses.SparseCategoricalCrossentropy(from_logits=True) model.compile(optimizer="adam", loss=loss_fn, metrics=["accuracy"]) model.fit(x_train, y_train, epochs=5, batch_size=4) - 强制使用CPU运行:如果升级插件后问题仍存在,可临时禁用GPU,让TensorFlow使用CPU执行,规避Metal插件的适配问题。在脚本开头添加:
import os os.environ["CUDA_VISIBLE_DEVICES"] = "-1" - 进一步降低batch size:将batch size从4调整为2,减少单步训练的内存占用,避免内存分配异常。
内容的提问来源于stack exchange,提问作者Josh Purtell
相关产品推荐
相关产品推荐

