You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Huggingface Trainer触发vCenter中Ubuntu VM无预警无日志关机求助

Trainer()模块触发VM立即关机(GPU环境下出现异常,CPU运行正常)

我已排查该问题一周有余,此问题未在任何日志中留下任何错误痕迹,特来询问是否有其他用户遇到过类似情况。无论使用何种笔记本、安装/升级/卸载何种模块,Trainer()模块都会导致VM立即关机。我怀疑该问题与GPU相关,因为在CPU上运行无异常。我已设置可见设备为(0,1),也尝试过启用/禁用wandb并设置report_to="none"。

相关环境信息

Is cuda available? True
Cuda torch version? 12.1
Is cuDNN version: 8902
cuDNN enabled?  True
Device count? 1
Current device? 0
Device name?  NVIDIA A30
tensor([[0.4543, 0.0545, 0.9293],
        [0.7722, 0.6535, 0.1276],
        [0.9957, 0.5621, 0.1621],
        [0.3164, 0.2845, 0.6874],
        [0.5489, 0.7582, 0.7139]])

相关代码

# setting device on GPU if available, else CPU

device = torch.device('cuda' if torch.cuda.is_available() else 'cpu')
print('Using device:', device)
print()

#Additional Info when using cuda
if device.type == 'cuda':
    print(torch.cuda.get_device_name(0))
    print('Memory Usage:')
    print('Allocated:', round(torch.cuda.memory_allocated(0)/1024**3,1), 'GB')
    print('Cached:   ', round(torch.cuda.memory_reserved(0)/1024**3,1), 'GB')

请问是否有用户遇到过此类问题?

内容的提问来源于stack exchange,提问作者texasdave

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.08 01:38:34