You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Colab Pro+中微调BioMedLM做摘要时PyTorch报CUDA设备序号无效错误

问题:BioMedLM摘要微调时触发CUDA无效设备序号错误

环境与执行命令

  • 运行环境:Google Colab Pro+(启用GPU加速器、Premium级GPU、高RAM运行时)
  • 执行的微调命令:
python -m torch.distributed.launch --nproc_per_node=8 --nnodes=1 --node_rank=0 finetune_for_summarization.py --output_dir {run_dir} --model_name_or_path stanford-crfm/BioMedLM --tokenizer_name stanford-crfm/pubmed_gpt_tokenizer --per_device_train_batch_size 1 --per_device_eval_batch_size 1 --save_strategy no --do_eval --train_data_file data/meqsum/train.source --eval_data_file data/meqsum/val.source --save_total_limit 2 --overwrite_output_dir --gradient_accumulation_steps 1 --learning_rate 1e-5 --warmup_ratio 0.5 --weight_decay 0.0 --seed 7 --evaluation_strategy steps --eval_steps 200 --bf16 --num_train_epochs 3 --logging_steps 100 --logging_first_step

错误日志

Traceback (most recent call last):
  File "/content/BioMedLM/finetune/textgen/gpt2/finetune_for_summarization.py", line 151, in <module>
    return func(*args, **kwargs)
  File "/usr/local/lib/python3.9/dist-packages/transformers/training_args.py", line 1468, in _setup_devices
        return func(*args, **kwargs)finetune()

  File "/usr/local/lib/python3.9/dist-packages/transformers/training_args.py", line 1468, in _setup_devices
  File "/content/BioMedLM/finetune/textgen/gpt2/finetune_for_summarization.py", line 105, in finetune
    model_args, data_args, training_args = parser.parse_args_into_dataclasses()
  File "/usr/local/lib/python3.9/dist-packages/transformers/hf_argparser.py", line 226, in parse_args_into_dataclasses
    torch.cuda.set_device(device)
  File "/usr/local/lib/python3.9/dist-packages/torch/cuda/__init__.py", line 350, in set_device
    obj = dtype(**inputs)
  File "<string>", line 104, in __init__
  File "/usr/local/lib/python3.9/dist-packages/transformers/training_args.py", line 1118, in __post_init__
    torch._C._cuda_setDevice(device)
RuntimeError: CUDA error: invalid device ordinal
CUDA kernel errors might be asynchronously reported at some other API call, so the stacktrace below might be incorrect.
For debugging consider passing CUDA_LAUNCH_BLOCKING=1.
Compile with `TORCH_USE_CUDA_DSA` to enable device-side assertions.

不修改源码的解决方案

错误根源是命令中指定了--nproc_per_node=8,要求启动8个GPU进程,但Google Colab Pro+的Premium GPU实例仅提供单张GPU(如A100或T4),进程尝试访问不存在的GPU设备序号(序号从0开始,单卡仅0有效),导致报错。

解决步骤:

  1. 先确认当前Colab实例的GPU数量,执行以下命令查看:
import torch
print(torch.cuda.device_count())
  1. 修改分布式启动参数,将--nproc_per_node的值改为实际可用的GPU数量(单卡则设为1),修改后的命令示例:
python -m torch.distributed.launch --nproc_per_node=1 --nnodes=1 --node_rank=0 finetune_for_summarization.py --output_dir {run_dir} --model_name_or_path stanford-crfm/BioMedLM --tokenizer_name stanford-crfm/pubmed_gpt_tokenizer --per_device_train_batch_size 1 --per_device_eval_batch_size 1 --save_strategy no --do_eval --train_data_file data/meqsum/train.source --eval_data_file data/meqsum/val.source --save_total_limit 2 --overwrite_output_dir --gradient_accumulation_steps 1 --learning_rate 1e-5 --warmup_ratio 0.5 --weight_decay 0.0 --seed 7 --evaluation_strategy steps --eval_steps 200 --bf16 --num_train_epochs 3 --logging_steps 100 --logging_first_step
  1. 若不需要分布式训练,也可以直接去掉分布式启动前缀,用普通方式运行:
python finetune_for_summarization.py --output_dir {run_dir} --model_name_or_path stanford-crfm/BioMedLM --tokenizer_name stanford-crfm/pubmed_gpt_tokenizer --per_device_train_batch_size 1 --per_device_eval_batch_size 1 --save_strategy no --do_eval --train_data_file data/meqsum/train.source --eval_data_file data/meqsum/val.source --save_total_limit 2 --overwrite_output_dir --gradient_accumulation_steps 1 --learning_rate 1e-5 --warmup_ratio 0.5 --weight_decay 0.0 --seed 7 --evaluation_strategy steps --eval_steps 200 --bf16 --num_train_epochs 3 --logging_steps 100 --logging_first_step

内容的提问来源于stack exchange,提问作者Chance

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 11:47:43