基于distilbart-cnn-12-6微调DistilBART:数据集问题与实现咨询
Hey there, let's break down your questions one by one with practical details:
1. 如何以distilbart-cnn-12-6为初始检查点,在Gigaword和CNN DailyMail数据集上微调DistilBART实现抽象文本摘要?
You're already on the right track with the finetune.py script from the transformers repo—here's how to adapt it for both datasets:
步骤1:准备数据集
- 对于Gigaword和CNN DailyMail,先确保你的数据集格式符合
finetune.py的要求:在data_dir下需要有train.source、train.target、val.source、val.target、test.source、test.target文件,每个文件中每行对应一个样本的原文或摘要。如果用TensorFlow Datasets的版本,你需要先把数据集导出成这种文本格式(提取每个样本的document和summary字段,分别写入对应的source/target文件)。
步骤2:调整微调命令
修改你提供的代码,把model_name_or_path指向官方的distilbart-cnn-12-6检查点(或者你本地下载好的路径),然后针对每个数据集单独运行:
针对CNN DailyMail的示例命令:
import os os.environ['PYTHONPATH'] += ":/content/transformers/examples" %cd "/content/transformers/examples" !python /content/transformers/examples/seq2seq/finetune.py \ --learning_rate=3e-5 \ --fp16 \ --gpus 1 \ --do_train \ --do_predict \ --n_val 1000 \ --val_check_interval 0.1 \ --sortish_sampler \ --data_dir '/content/cnn_dailymail_dataset' # 替换为你的CNN DailyMail数据集路径 --train_batch_size=4 \ --eval_batch_size=4 \ --output_dir=distilbart_cnn_dm_finetuned \ --num_train_epochs 3 \ --model_name_or_path facebook/distilbart-cnn-12-6
针对Gigaword的示例命令:
只需要把data_dir换成Gigaword的数据集路径,output_dir改成对应的保存目录即可,其他参数可以根据你的硬件情况调整(比如batch size、epochs)。
步骤3:验证与预测
训练完成后,脚本会自动在测试集上生成预测结果,你可以对比test.target和生成的摘要文件,评估模型效果。
2. 使用TensorFlow提供的Gigaword数据集时,数字被替换为“#”,生成的摘要中也出现#,该现象是否正常?
Absolutely normal! The Gigaword dataset comes with built-in preprocessing that replaces all numeric values with # as a form of normalization. This is done to reduce the size of the vocabulary and help the model focus on semantic content rather than specific numbers.
Since your model is trained on this preprocessed data, it learns to map numeric contexts to # in its outputs. If you want to retain actual numbers in your summaries, you'll need to reprocess the raw Gigaword data without this number-replacement step, then retrain the model—just note that this will increase your vocabulary size, so you might need to adjust training parameters like learning rate or batch size accordingly.
3. 是否可以用distilbart-cnn-12-6检查点对CNN DailyMail数据集进行微调?
Definitely! In fact, distilbart-cnn-12-6 is already distilled from BART-base which was fine-tuned on CNN DailyMail—so using this checkpoint as a starting point for further fine-tuning on CNN DailyMail is a great choice. It will let the model refine its abstractive summarization skills specifically for this dataset, leading to better performance than training from scratch.
You can use the example command I shared in question 1 for CNN DailyMail—just make sure your dataset is correctly formatted and the paths are set properly.
内容的提问来源于stack exchange,提问作者Moon Days

