You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

微调mT5用于摘要任务后,无掩码输入推理仍生成<extra_id_n>标记的问题求助

微调mT5用于摘要任务后,无掩码输入推理仍生成<extra_id_n>标记的问题求助

大家好,我最近用新数据集对mT5模型做了监督式微调,用来完成摘要生成任务。但现在推理阶段遇到了一个棘手的问题:就算输入文本没有包含掩码标记,模型生成的输出里还是会出现<extra_id_1>这类特殊标记,这和我预期的正常摘要输出不符。

下面是我用来编码输入的代码:

tokenized_inputs = self.tokenizer.batch_encode_plus(
    [line],
    max_length=self.max_len,
    padding="max_length",
    return_tensors="pt"
).to(self.args.device)

以及生成输出的代码:

Summary_input_ids= model.generate(
    input_ids=input_ids,
    attention_mask=input_mask,
    do_sample=True,
    temperature=0.8,
    top_k=45,
    top_p=0.9,
    max_length=_max_length,
    min_length=_min_length,
    num_beams=_num_beams,
    repetition_penalty=2.5,
    no_repeat_ngram_size = _no_repeat_ngram_size,
    length_penalty=2.5,
    early_stopping=False,
    use_cache=True,
    num_return_sequences=1
)
Summary = tokenizer.batch_decode(Summary_input_ids,
    skip_special_tokens=True,clean_up_tokenization_spaces=False)

因为我的模型是通过监督式微调训练的,原本应该是给每个输入生成对应的正常摘要才对,现在却出现了这些额外的特殊标记,实在搞不清楚问题出在哪,希望有经验的朋友能帮忙分析一下原因,或者给我一些排查的方向,谢谢大家!

备注:内容来源于stack exchange,提问作者Saeed Farzi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.22 07:43:10