You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Hugging Face Trainer计算ROUGE时CUDA内存不足求助

问题描述

使用Hugging Face Trainer进行模型评估时,调用compute_metrics函数计算ROUGE分数遭遇CUDA out of memory错误,GPU内存耗尽。环境配置与详情:

  • 验证数据集val_dataset包含352个样本
  • 使用模型为GPT-2
  • 运行环境为Google Colab

疑问:

  1. 如何在不耗尽GPU内存的前提下高效计算ROUGE或BLEU指标?
  2. 针对有限GPU内存的大规模评估,有哪些推荐的策略或配置方案?

相关代码

tokenized_train_dataset = train_dataset.map(tokenize_function, batched=True)
tokenized_validation_dataset = validation_dataset.map(tokenize_function, batched=True)
config = AutoConfig.from_pretrained(
    "gpt2",
    vocab_size=len(tokenizer),
    n_ctx=MAX,
    bos_token_id=tokenizer.bos_token_id,
    eos_token_id=tokenizer.eos_token_id,
)
model = GPT2LMHeadModel(config).to('cuda' if torch.cuda.is_available() else 'cpu')
mode_size = sum(t.numel() for t in model.parameters())
print(f"Model size: {mode_size/1000**2:.1f}M parameters")

from transformers import DataCollatorForLanguageModeling
data_collator = DataCollatorForLanguageModeling(
    tokenizer=tokenizer, mlm=False,
)

class EarlyStoppingCallback(TrainerCallback):
    def __init__(self, patience=3):
        super().__init__()  
        self.patience = patience
        self.best_loss = np.inf
        self.epochs_no_improve = 0

    def on_evaluate(self, args, state, control, metrics=None, **kwargs):
        eval_loss = metrics.get("eval_loss", None)
        if eval_loss is not None:
            if eval_loss < self.best_loss:
                self.best_loss = eval_loss
                self.epochs_no_improve = 0
            else:
                self.epochs_no_improve += 1
                if self.epochs_no_improve >= self.patience:
                    control.should_training_stop = True  


early_stopping_callback = EarlyStoppingCallback(patience=3)

training_args = TrainingArguments(
    output_dir="./model",
    hub_model_id="profile/model",
    eval_strategy="epoch",
    gradient_accumulation_steps=4,
    per_device_train_batch_size=4,
    per_device_eval_batch_size=4,
    num_train_epochs=10,
    weight_decay=0.01,
    logging_dir='./logs',
    save_steps=500,
    save_total_limit=3,
    learning_rate=1e-4,
    fp16=True,
    push_to_hub=True,
    logging_steps=100,
)


trainer = Trainer(
    model=model,
    tokenizer=tokenizer,
    args=training_args,
    data_collator=data_collator,
    train_dataset=tokenized_train_dataset,
    eval_dataset=tokenized_validation_dataset,
    callbacks=[early_stopping_callback]
)


trainer.train()
解决方案

1. 低内存下高效计算ROUGE/BLEU指标

  • 将指标计算移至CPU:ROUGE、BLEU属于文本统计类指标,完全不需要GPU加速。在compute_metrics函数中先把模型输出的预测和标签从GPU转移到CPU,再解码为文本计算指标,避免占用GPU内存:
    def compute_metrics(eval_pred):
        predictions, labels = eval_pred
        # 转移到CPU并转换为numpy数组
        predictions = predictions.cpu().numpy()
        labels = labels.cpu().numpy()
        # 解码为可读文本
        decoded_preds = tokenizer.batch_decode(predictions, skip_special_tokens=True)
        decoded_labels = tokenizer.batch_decode(labels, skip_special_tokens=True)
        # 计算ROUGE
        rouge = evaluate.load("rouge")
        result = rouge.compute(predictions=decoded_preds, references=decoded_labels)
        return {k: round(v, 4) for k, v in result.items()}
    
  • 缩小评估批次:在TrainingArguments中降低per_device_eval_batch_size,比如设为1或2,减少单次评估时GPU的内存负载。
  • 离线计算指标:如果已保存模型生成的预测结果,直接在CPU上离线计算指标,无需在Trainer的评估流程中实时运行,节省GPU资源。

2. 有限GPU内存下大规模评估的策略

  • 小批次/逐样本评估:手动实现评估循环,将验证集拆分为极小批次甚至逐样本处理,每处理完一批就清空GPU缓存,避免内存累积。
  • 启用梯度检查点:给模型开启梯度检查点,减少评估时的内存占用,初始化模型后添加:
    model.gradient_checkpointing_enable()
    
  • 保持混合精度:继续使用fp16=True配置,半精度计算能大幅降低内存消耗,Colab的GPU完全支持该模式。
  • 手动清空GPU缓存:在评估前执行torch.cuda.empty_cache(),释放训练阶段遗留的临时内存。
  • 模型量化:使用bitsandbytes库进行4-bit或8-bit量化,可将模型内存占用降至原有的1/4或1/2,且性能损失极小:
    from transformers import GPT2LMHeadModel, BitsAndBytesConfig
    
    bnb_config = BitsAndBytesConfig(
        load_in_4bit=True,
        bnb_4bit_use_double_quant=True,
        bnb_4bit_quant_type="nf4",
        bnb_4bit_compute_dtype=torch.bfloat16
    )
    # 初始化量化后的模型
    model = GPT2LMHeadModel.from_pretrained("gpt2", quantization_config=bnb_config)
    
  • 分阶段评估:将大验证集拆分为多个子集,分批次评估后合并指标结果,避免一次性加载所有数据到GPU。

内容的提问来源于stack exchange,提问作者KainnT

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 10:37:02