Colab Pro+中微调BioMedLM做摘要时PyTorch报CUDA设备序号无效错误
问题:BioMedLM摘要微调时触发CUDA无效设备序号错误
环境与执行命令
- 运行环境:Google Colab Pro+(启用GPU加速器、Premium级GPU、高RAM运行时)
- 执行的微调命令:
python -m torch.distributed.launch --nproc_per_node=8 --nnodes=1 --node_rank=0 finetune_for_summarization.py --output_dir {run_dir} --model_name_or_path stanford-crfm/BioMedLM --tokenizer_name stanford-crfm/pubmed_gpt_tokenizer --per_device_train_batch_size 1 --per_device_eval_batch_size 1 --save_strategy no --do_eval --train_data_file data/meqsum/train.source --eval_data_file data/meqsum/val.source --save_total_limit 2 --overwrite_output_dir --gradient_accumulation_steps 1 --learning_rate 1e-5 --warmup_ratio 0.5 --weight_decay 0.0 --seed 7 --evaluation_strategy steps --eval_steps 200 --bf16 --num_train_epochs 3 --logging_steps 100 --logging_first_step
错误日志
Traceback (most recent call last): File "/content/BioMedLM/finetune/textgen/gpt2/finetune_for_summarization.py", line 151, in <module> return func(*args, **kwargs) File "/usr/local/lib/python3.9/dist-packages/transformers/training_args.py", line 1468, in _setup_devices return func(*args, **kwargs)finetune() File "/usr/local/lib/python3.9/dist-packages/transformers/training_args.py", line 1468, in _setup_devices File "/content/BioMedLM/finetune/textgen/gpt2/finetune_for_summarization.py", line 105, in finetune model_args, data_args, training_args = parser.parse_args_into_dataclasses() File "/usr/local/lib/python3.9/dist-packages/transformers/hf_argparser.py", line 226, in parse_args_into_dataclasses torch.cuda.set_device(device) File "/usr/local/lib/python3.9/dist-packages/torch/cuda/__init__.py", line 350, in set_device obj = dtype(**inputs) File "<string>", line 104, in __init__ File "/usr/local/lib/python3.9/dist-packages/transformers/training_args.py", line 1118, in __post_init__ torch._C._cuda_setDevice(device) RuntimeError: CUDA error: invalid device ordinal CUDA kernel errors might be asynchronously reported at some other API call, so the stacktrace below might be incorrect. For debugging consider passing CUDA_LAUNCH_BLOCKING=1. Compile with `TORCH_USE_CUDA_DSA` to enable device-side assertions.
不修改源码的解决方案
错误根源是命令中指定了--nproc_per_node=8,要求启动8个GPU进程,但Google Colab Pro+的Premium GPU实例仅提供单张GPU(如A100或T4),进程尝试访问不存在的GPU设备序号(序号从0开始,单卡仅0有效),导致报错。
解决步骤:
- 先确认当前Colab实例的GPU数量,执行以下命令查看:
import torch print(torch.cuda.device_count())
- 修改分布式启动参数,将
--nproc_per_node的值改为实际可用的GPU数量(单卡则设为1),修改后的命令示例:
python -m torch.distributed.launch --nproc_per_node=1 --nnodes=1 --node_rank=0 finetune_for_summarization.py --output_dir {run_dir} --model_name_or_path stanford-crfm/BioMedLM --tokenizer_name stanford-crfm/pubmed_gpt_tokenizer --per_device_train_batch_size 1 --per_device_eval_batch_size 1 --save_strategy no --do_eval --train_data_file data/meqsum/train.source --eval_data_file data/meqsum/val.source --save_total_limit 2 --overwrite_output_dir --gradient_accumulation_steps 1 --learning_rate 1e-5 --warmup_ratio 0.5 --weight_decay 0.0 --seed 7 --evaluation_strategy steps --eval_steps 200 --bf16 --num_train_epochs 3 --logging_steps 100 --logging_first_step
- 若不需要分布式训练,也可以直接去掉分布式启动前缀,用普通方式运行:
python finetune_for_summarization.py --output_dir {run_dir} --model_name_or_path stanford-crfm/BioMedLM --tokenizer_name stanford-crfm/pubmed_gpt_tokenizer --per_device_train_batch_size 1 --per_device_eval_batch_size 1 --save_strategy no --do_eval --train_data_file data/meqsum/train.source --eval_data_file data/meqsum/val.source --save_total_limit 2 --overwrite_output_dir --gradient_accumulation_steps 1 --learning_rate 1e-5 --warmup_ratio 0.5 --weight_decay 0.0 --seed 7 --evaluation_strategy steps --eval_steps 200 --bf16 --num_train_epochs 3 --logging_steps 100 --logging_first_step
内容的提问来源于stack exchange,提问作者Chance
相关产品推荐
相关产品推荐

