You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在Huggingface Pipeline中使用TPU触发PyTorch RuntimeError问题排查

TPU运行Hugging Face Pipeline报错RuntimeError: Cannot set version_counter for inference tensor的修复方案

错误含义

这个错误的核心原因是PyTorch XLA(TPU的PyTorch适配框架)在处理推理模式下的张量时,不支持设置version_counter属性。该属性是PyTorch原生用于跟踪张量版本变化(比如训练时参数更新)的,但TPU推理时的张量是只读、无需梯度计算的,不需要版本跟踪,因此框架触发了兼容性冲突。

修复步骤及代码修改

问题出在Hugging Face Pipeline自动处理TPU设备时的兼容缺陷,手动控制模型加载和设备分配即可解决:

  • 正确获取TPU设备:通过xm.xla_device()获取TPU设备对象,而非直接使用未初始化的device变量
  • 手动加载模型与分词器:避免依赖Pipeline自动加载逻辑,手动将模型移至TPU设备
  • 调整Pipeline的device参数:传入设备索引而非设备对象,同时推理时禁用梯度计算

修改后的完整代码:

from transformers import pipeline, AutoModelForSequenceClassification, AutoTokenizer
import torch
import torch_xla.core.xla_model as xm

# 获取TPU设备
device = xm.xla_device()

# 手动加载模型和分词器,并将模型移至TPU
model = AutoModelForSequenceClassification.from_pretrained('bhadresh-savani/distilbert-base-uncased-emotion').to(device)
tokenizer = AutoTokenizer.from_pretrained('bhadresh-savani/distilbert-base-uncased-emotion')

# 构建pipeline时传入已加载的模型、分词器,设备参数传索引
classifier = pipeline(
    "text-classification",
    model=model,
    tokenizer=tokenizer,
    return_all_scores=True,
    device=device.index
)

def detect_emotions(emotion_input):
    """模型推理部分"""
    with torch.no_grad():  # 禁用梯度计算,避免生成需要版本跟踪的张量
        prediction = classifier(emotion_input)
    output = {}
    for emotion in prediction[0]:
        output[emotion["label"]] = emotion["score"]   
    return output

# 执行推理
print(detect_emotions('Rest in Power: The Trayvon Martin Story’ takes an emotional look back at the shooting that divided a nation'))

额外说明

  • torch.no_grad()是关键:它会阻止PyTorch生成梯度相关的张量属性,从根源上避免version_counter的设置尝试
  • 直接传入设备索引而非对象:Pipeline的device参数对TPU设备的兼容逻辑更适配整数索引,而非PyTorch设备对象

内容的提问来源于stack exchange,提问作者DarknessPlusPlus

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 03:31:03