微调Persian-T5复述模型遇两类ValueError问题求助
问题解决:微调波斯语T5模型做复述检测
错误1:使用AutoModelForSequenceClassification加载T5模型失败
原因:AutoModelForSequenceClassification的模型映射中不包含T5Config,T5有专门的序列分类模型实现类。
解决方法:直接使用T5ForSequenceClassification加载预训练模型,指定分类标签数:
from transformers import T5ForSequenceClassification model = T5ForSequenceClassification.from_pretrained("erfan226/persian-t5-paraphraser", num_labels=2)
错误2:trainer.train()报ValueError: not enough values to unpack (expected 2, got 1)
原因:
- 误用了
AutoModelForSeq2SeqLM(生成式模型)处理分类任务,输入格式不符合Seq2Seq模型要求,且任务逻辑不匹配。 - 分词后的数据集未正确适配分类模型的字段要求。
修正步骤:
- 修改分词函数:T5处理分类任务时,需将两个句子用模型默认分隔符
</s>拼接,同时保留label字段:
def tokenize_function(examples): # 拼接两个句子,使用T5标准分隔符 texts = [f"{s1} </s> {s2}" for s1, s2 in zip(examples["sentence1"], examples["sentence2"])] tokenized = tokenizer(texts, truncation=True, padding="max_length", max_length=128) # 绑定标签字段 tokenized["labels"] = examples["label"] return tokenized
替换为分类专用模型:使用
T5ForSequenceClassification替代AutoModelForSeq2SeqLM(参考错误1的解决代码)。完善训练配置(可选):添加分类指标计算,提升训练监控效果:
# 加载MRPC评估指标 import evaluate metric = evaluate.load("glue", "mrpc") def compute_metrics(eval_pred): predictions, labels = eval_pred predictions = predictions.argmax(axis=1) return metric.compute(predictions=predictions, references=labels) # 初始化Trainer training_args = TrainingArguments( "farsi_paraphraser", num_train_epochs=5, evaluation_strategy="epoch", per_device_train_batch_size=8, per_device_eval_batch_size=8 ) trainer = Trainer( model=model, args=training_args, train_dataset=tokenized_dataset["train"], eval_dataset=tokenized_dataset["validation"], data_collator=data_collator, tokenizer=tokenizer, compute_metrics=compute_metrics ) trainer.train()
关键说明
- T5是Encoder-Decoder架构,用于分类任务必须使用
T5ForSequenceClassification,该类会在模型尾部添加分类头,而非沿用生成任务的Decoder结构。 - 句子拼接时保留
</s>分隔符是T5模型的标准输入格式,也可添加任务前缀(如"paraphrase detection: ")引导模型理解任务目标。
内容的提问来源于stack exchange,提问作者Ali Ghasemi
相关产品推荐
相关产品推荐

