仅使用CPU运行PyTorch/Torchtune时遇RuntimeError问题求助
解决仅用CPU运行Torchtune微调时的RuntimeError问题
错误原因
默认的llama3/8B_lora_single_device配置是为GPU环境设计的,Torchtune会自动尝试调用GPU,当系统无可用GPU时就会触发RuntimeError: The local rank is larger than the number of available GPUs.错误。
解决方案
方法1:修改配置文件强制使用CPU
- 先把官方配置导出到本地文件:
tune run lora_finetune_single_device --config llama3/8B_lora_single_device --export_config ./cpu_config.yaml - 打开
cpu_config.yaml,修改两处关键参数:- 将
device: cuda替换为device: cpu - 降低
batch_size(比如改为1),避免CPU内存不足
- 将
- 用修改后的配置运行微调:
tune run lora_finetune_single_device --config ./cpu_config.yaml
方法2:命令行直接覆盖配置参数
无需修改配置文件,直接在运行命令里指定设备和调整batch size:
tune run lora_finetune_single_device --config llama3/8B_lora_single_device device=cpu batch_size=1
注意事项
- Llama3 8B模型体积较大,CPU运行微调速度会极慢,且需要至少32GB以上的内存,内存不足会触发OOM错误
- 如果内存不够,建议换用更小的模型(比如Llama3 7B量化版),或者调整
gradient_accumulation_steps参数来间接提升有效batch size
内容的提问来源于stack exchange,提问作者graph User
相关产品推荐
相关产品推荐

