使用HuggingFace Trainer时如何指定训练所用GPU?
指定HuggingFace Trainer运行的GPU
有多种方法可以指定Trainer使用的GPU设备,以下是具体实现方式:
1. 通过TrainingArguments直接指定单GPU
在TrainingArguments中添加device参数,明确指定要使用的GPU索引(索引从0开始),比如指定使用第2块GPU:
training_args = TrainingArguments( output_dir='./results', # output directory num_train_epochs=3, # total # of training epochs per_device_train_batch_size=16, # batch size per device during training per_device_eval_batch_size=64, # batch size for evaluation warmup_steps=500, # number of warmup steps for learning rate scheduler weight_decay=0.01, # strength of weight decay logging_dir='./logs', # directory for storing logs device="cuda:1" # 指定使用cuda:1设备 )
2. 通过环境变量限制可见GPU
运行训练脚本前,设置CUDA_VISIBLE_DEVICES环境变量,限定脚本能访问的GPU范围:
- Linux/macOS终端:
CUDA_VISIBLE_DEVICES=1,2 python your_training_script.py
- Windows命令提示符:
set CUDA_VISIBLE_DEVICES=1,2 python your_training_script.py
设置后,Trainer默认会使用可见列表中的第一块GPU(实际硬件的cuda:1会被脚本识别为cuda:0),如果需要多GPU训练,可结合后续的多GPU配置参数。
3. 指定多GPU训练
如果要使用多块GPU,除了通过环境变量限定设备列表,还可以在TrainingArguments中设置n_gpu参数指定GPU数量:
training_args = TrainingArguments( output_dir='./results', num_train_epochs=3, per_device_train_batch_size=16, per_device_eval_batch_size=64, warmup_steps=500, weight_decay=0.01, logging_dir='./logs', n_gpu=2 # 指定使用2块GPU )
注意:此方式需要确保CUDA_VISIBLE_DEVICES已设置好目标GPU,或者系统能直接识别到这些设备。
内容的提问来源于stack exchange,提问作者Arist12
相关产品推荐
相关产品推荐

