You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

仅使用CPU运行PyTorch/Torchtune时遇RuntimeError问题求助

解决仅用CPU运行Torchtune微调时的RuntimeError问题

错误原因

默认的llama3/8B_lora_single_device配置是为GPU环境设计的,Torchtune会自动尝试调用GPU,当系统无可用GPU时就会触发RuntimeError: The local rank is larger than the number of available GPUs.错误。

解决方案

方法1:修改配置文件强制使用CPU

  • 先把官方配置导出到本地文件:
    tune run lora_finetune_single_device --config llama3/8B_lora_single_device --export_config ./cpu_config.yaml
    
  • 打开cpu_config.yaml,修改两处关键参数:
    • 将device: cuda替换为device: cpu
    • 降低batch_size(比如改为1),避免CPU内存不足
  • 用修改后的配置运行微调:
    tune run lora_finetune_single_device --config ./cpu_config.yaml
    

方法2:命令行直接覆盖配置参数

无需修改配置文件,直接在运行命令里指定设备和调整batch size:

tune run lora_finetune_single_device --config llama3/8B_lora_single_device device=cpu batch_size=1

注意事项

  • Llama3 8B模型体积较大,CPU运行微调速度会极慢,且需要至少32GB以上的内存,内存不足会触发OOM错误
  • 如果内存不够,建议换用更小的模型(比如Llama3 7B量化版),或者调整gradient_accumulation_steps参数来间接提升有效batch size

内容的提问来源于stack exchange,提问作者graph User

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.23 10:09:53