You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

M2 Max运行DistilBert推理时GPU利用率持续下降求助排查

问题描述

我使用带96GB显存的Apple M2 Max,在Pandas数据框文本列上运行推理,采用循环单条处理(未做批处理),模型是微调后的Hugging Face distilbert-base-cased。初始GPU利用率约50%,但随后缓慢降至1%甚至更低。怀疑是热节流问题,外接风扇后无改善,当前推理速度极慢,寻求排查方向与解决建议。

代码示例
from transformers import AutoTokenizer, TFDistilBertForSequenceClassification
from datasets import load_dataset
import tqdm
import numpy as np

imdb = load_dataset('imdb')
sentences = imdb['train']['text'][:500]

tokenizer = AutoTokenizer.from_pretrained("distilbert-base-cased")
model = TFDistilBertForSequenceClassification.from_pretrained('distilbert-base-cased')

for i, sentence in tqdm.tqdm(enumerate(sentences)):
  inputs = tokenizer(sentence, truncation=True, return_tensors='tf')
  output = model(inputs).logits
  pred = np.argmax(output.numpy(), axis=1)

  if i % 100 == 0:
    print(f"len(input_ids): {inputs['input_ids'].shape[-1]}")
运行输出
Metal device set to: Apple M2 Max

systemMemory: 96.00 GB
maxCacheSize: 36.00 GB

3it [00:00, 10.87it/s]
len(input_ids): 391
101it [00:13,  6.38it/s]
len(input_ids): 215
201it [00:34,  4.78it/s]
len(input_ids): 237
301it [00:55,  4.26it/s]
len(input_ids): 256
401it [01:54,  1.12it/s]
len(input_ids): 55
500it [03:40,  2.27it/s]
排查方向与解决建议
  • 改用批处理推理:单条循环会让GPU频繁启停,无法发挥并行计算能力。收集所有句子后统一tokenize(开启padding),分批次输入模型,能直接拉满GPU利用率,提升推理速度。
  • 验证TensorFlow-Metal兼容性:M系列芯片需安装tensorflow-metal插件,且版本要与TensorFlow匹配。旧版本可能存在内存泄漏或调度bug,导致GPU利用率持续下降。
  • 减少CPU-GPU数据交互:每次循环调用output.numpy()会触发跨设备数据拷贝,累计拖慢速度。建议先将所有预测结果存在GPU张量中,最后一次性转成numpy数组。
  • 清理GPU内存碎片:循环中生成的张量若未及时释放,会产生内存碎片,降低调度效率。可在循环末尾调用tf.keras.backend.clear_session(),或用上下文管理器限制张量生命周期。
  • 确认模型加载到GPU:检查模型是否运行在MPS设备上,可通过model.summary()查看,或显式执行model = model.to('mps')确保模型加载到GPU。
  • 监控系统内存占用:用Activity Monitor查看统一内存使用情况,若其他程序占用过多内存,会触发交换内存,导致整体运行变慢。即使96GB显存充足,系统内存不足也会影响模型性能。

内容的提问来源于stack exchange,提问作者kawingkelvin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 16:22:09