You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用HuggingFace Trainer时如何指定训练所用GPU?

指定HuggingFace Trainer运行的GPU

有多种方法可以指定Trainer使用的GPU设备,以下是具体实现方式:

1. 通过TrainingArguments直接指定单GPU

在TrainingArguments中添加device参数,明确指定要使用的GPU索引(索引从0开始),比如指定使用第2块GPU:

training_args = TrainingArguments(
    output_dir='./results',          # output directory
    num_train_epochs=3,              # total # of training epochs
    per_device_train_batch_size=16,  # batch size per device during training
    per_device_eval_batch_size=64,   # batch size for evaluation
    warmup_steps=500,                # number of warmup steps for learning rate scheduler
    weight_decay=0.01,               # strength of weight decay
    logging_dir='./logs',            # directory for storing logs
    device="cuda:1"  # 指定使用cuda:1设备
)

2. 通过环境变量限制可见GPU

运行训练脚本前,设置CUDA_VISIBLE_DEVICES环境变量,限定脚本能访问的GPU范围:

  • Linux/macOS终端:
CUDA_VISIBLE_DEVICES=1,2 python your_training_script.py
  • Windows命令提示符:
set CUDA_VISIBLE_DEVICES=1,2
python your_training_script.py

设置后,Trainer默认会使用可见列表中的第一块GPU(实际硬件的cuda:1会被脚本识别为cuda:0),如果需要多GPU训练,可结合后续的多GPU配置参数。

3. 指定多GPU训练

如果要使用多块GPU,除了通过环境变量限定设备列表,还可以在TrainingArguments中设置n_gpu参数指定GPU数量:

training_args = TrainingArguments(
    output_dir='./results',
    num_train_epochs=3,
    per_device_train_batch_size=16,
    per_device_eval_batch_size=64,
    warmup_steps=500,
    weight_decay=0.01,
    logging_dir='./logs',
    n_gpu=2  # 指定使用2块GPU
)

注意:此方式需要确保CUDA_VISIBLE_DEVICES已设置好目标GPU,或者系统能直接识别到这些设备。

内容的提问来源于stack exchange,提问作者Arist12

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 21:22:09