Huggingface Trainer触发vCenter中Ubuntu VM无预警无日志关机求助
Trainer()模块触发VM立即关机(GPU环境下出现异常,CPU运行正常)
我已排查该问题一周有余,此问题未在任何日志中留下任何错误痕迹,特来询问是否有其他用户遇到过类似情况。无论使用何种笔记本、安装/升级/卸载何种模块,Trainer()模块都会导致VM立即关机。我怀疑该问题与GPU相关,因为在CPU上运行无异常。我已设置可见设备为(0,1),也尝试过启用/禁用wandb并设置report_to="none"。
相关环境信息
Is cuda available? True Cuda torch version? 12.1 Is cuDNN version: 8902 cuDNN enabled? True Device count? 1 Current device? 0 Device name? NVIDIA A30 tensor([[0.4543, 0.0545, 0.9293], [0.7722, 0.6535, 0.1276], [0.9957, 0.5621, 0.1621], [0.3164, 0.2845, 0.6874], [0.5489, 0.7582, 0.7139]])
相关代码
# setting device on GPU if available, else CPU device = torch.device('cuda' if torch.cuda.is_available() else 'cpu') print('Using device:', device) print() #Additional Info when using cuda if device.type == 'cuda': print(torch.cuda.get_device_name(0)) print('Memory Usage:') print('Allocated:', round(torch.cuda.memory_allocated(0)/1024**3,1), 'GB') print('Cached: ', round(torch.cuda.memory_reserved(0)/1024**3,1), 'GB')
请问是否有用户遇到过此类问题?
内容的提问来源于stack exchange,提问作者texasdave
相关产品推荐
相关产品推荐

