You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在GCP TPU VM上运行PyTorch时遇设备匹配错误求助

解决GCP TPU VM运行PyTorch代码时的TPU设备匹配错误

问题场景

在GCP上创建TPU VM后,按文档配置环境并设置环境变量:

export XRT_TPU_CONFIG="localservice;0;localhost:51011"

运行测试Python代码:

import torch
import torch_xla.core.xla_model as xm

dev = xm.xla_device()
t1 = torch.randn(3,3,device=dev)
t2 = torch.randn(3,3,device=dev)
print(t1 + t2)

执行python3 tpu-test.py时触发设备匹配错误:

RuntimeError: tensorflow/compiler/xla/xla_client/xrt_computation_client.cc:1374 : Check failed: session.Run({tensorflow::Output(result, 0)}, &outputs) == ::tensorflow::Status::OK() (INVALID_ARGUMENT: No matching devices found for '/job:localservice/replica:0/task:0/device:TPU_SYSTEM:0' vs. OK)

解决方案

1. 检查并重启TPU代理服务

TPU VM依赖tpu-vm-agent服务管理设备连接,先确认服务状态:

sudo systemctl status tpu-vm-agent

若服务未运行,启动并设置开机自启:

sudo systemctl start tpu-vm-agent
sudo systemctl enable tpu-vm-agent

2. 调整XRT_TPU_CONFIG环境变量格式

部分TorchXLA版本需用tpu_worker替代localservice,重新设置变量:

export XRT_TPU_CONFIG="tpu_worker;0;localhost:51011"

执行echo $XRT_TPU_CONFIG确认生效后,重新运行代码。

3. 确保PyTorch与TorchXLA版本兼容

版本不匹配是常见诱因,用官方推荐命令重新安装对应依赖:

# 以PyTorch 2.0为例,可根据TPU VM操作系统调整版本
pip install torch==2.0.0 torchvision==0.15.1 torchaudio==2.0.1 torch-xla[tpu]==2.0.0 -f https://storage.googleapis.com/libtpu-releases/index.html

安装完成后验证版本:

import torch
import torch_xla
print(torch.__version__)
print(torch_xla.__version__)

4. 重启TPU VM

若上述操作无效,尝试重启整个VM:

sudo reboot

重启后重新配置环境变量并运行代码。

5. 确认GCP TPU节点状态

登录GCP控制台,检查对应TPU节点是否处于READY状态,且与TPU VM在同一区域、同一VPC网络下。若节点状态异常,可删除后重新创建。

内容的提问来源于stack exchange,提问作者BioGeek

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 22:10:30