如何用Transformers的Trainer保存每次评估的预测结果?
解决每个Epoch保存预测结果的问题
1. 定位保存预测结果的代码块
先在run_summarization.py里找到负责保存预测结果的代码段,通常在评估逻辑中,原代码大致如下:
if args.predict_with_generate: predictions = tokenizer.batch_decode(preds, skip_special_tokens=True, clean_up_tokenization_spaces=True) predictions = [pred.strip() for pred in predictions] with open(args.predictions_file, "w") as writer: writer.write("\n".join(predictions))
2. 修改文件名,加入Epoch编号
原问题出在固定文件名会覆盖之前的结果,只需给文件名加上当前Epoch的编号,就能生成唯一的结果文件:
- 获取当前Epoch数:如果使用Trainer API,直接通过
trainer.state.epoch获取,注意该值可能是浮点数,转成整数即可。 - 生成带Epoch后缀的文件名,比如原文件是
predictions.txt,修改后变为predictions_epoch_1.txt、predictions_epoch_2.txt等。
修改后的代码示例:
# 在评估代码段中先获取当前epoch current_epoch = int(trainer.state.epoch) # 生成带epoch标识的新文件名 predictions_file = f"{args.predictions_file}_epoch_{current_epoch}" if args.predict_with_generate: predictions = tokenizer.batch_decode(preds, skip_special_tokens=True, clean_up_tokenization_spaces=True) predictions = [pred.strip() for pred in predictions] with open(predictions_file, "w") as writer: writer.write("\n".join(predictions))
3. 确保每个Epoch触发评估
如果使用Trainer,需要在TrainingArguments中设置evaluation_strategy="epoch",让每个Epoch结束后自动执行评估;如果是自定义训练循环,要手动在每个Epoch末尾调用评估函数,并传入当前Epoch编号。
4. 可选:同步保存评估指标
如果需要同时记录每个Epoch的评估指标(如ROUGE分数),可以新增指标文件保存逻辑:
import json metrics_file = f"{args.predictions_file}_metrics_epoch_{current_epoch}" with open(metrics_file, "w") as writer: writer.write(json.dumps(metrics, indent=4))
内容的提问来源于stack exchange,提问作者Benjamin Lynch
相关产品推荐
相关产品推荐

