You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

微调Hugging Face T5-base模型后generate输出过短问题求助

T5-base微调后生成输出过短问题的解决方案

问题背景

用Hugging Face的T5-base在新任务上微调,任务输入和目标均为256词的句子,模型损失已收敛至较低值,但调用generate方法生成的输出始终过短。尝试设置min_length与max_length参数无效,怀疑问题源于分词前句子长度固定为256、分词后长度不恒定(训练时靠padding保证输入尺寸一致)。

现有代码

Generate方法代码

model = transformers.T5ForConditionalGeneration.from_pretrained('t5-base')
tokenizer = T5Tokenizer.from_pretrained('t5-base')
generated_ids = model.generate(
    input_ids=ids,
    attention_mask=attn_mask,
    max_length=1024,
    min_length=256,
    num_beams=2,
    early_stopping=False,
    repetition_penalty=10.0
)
preds = [tokenizer.decode(g, skip_special_tokens=True, clean_up_tokenization_spaces=True) for g in generated_ids][0]
preds = preds.replace("<pad>", "").replace("</s>", "").strip().replace("  ", " ")
target = [tokenizer.decode(t, skip_special_tokens=True, clean_up_tokenization_spaces=True) for t in reference][0]
target = target.replace("<pad>", "").replace("</s>", "").strip().replace("  ", " ")

输入数据创建代码

tokens = tokenizer([f"task: {text}"], return_tensors="pt", max_length=1024, padding='max_length')
inputs_ids = tokens.input_ids.squeeze().to(dtype=torch.long)
attention_mask = tokens.attention_mask.squeeze().to(dtype=torch.long)
labels = self.tokenizer([target_text], return_tensors="pt", max_length=1024, padding='max_length')
label_ids = labels.input_ids.squeeze().to(dtype=torch.long)
label_attention = labels.attention_mask.squeeze().to(dtype=torch.long)

核心问题分析

  1. 训练标签的padding干扰:训练时目标序列padding到1024,但未将padding部分的标签设为-100,导致模型学习到"padding是有效输出"的错误逻辑,容易提前终止生成。
  2. min_length参数的误解:min_length是基于token数而非原始词数,256词对应的token数通常在300-400之间,直接设为256会导致长度不足。
  3. 解码逻辑冗余:skip_special_tokens=True已经会自动过滤<pad>和</s>,额外替换属于重复操作,不影响结果但增加冗余。

具体解决步骤

1. 修正训练时的标签处理

将padding部分的标签设为-100,让模型忽略padding内容:

labels = tokenizer([target_text], return_tensors="pt", max_length=1024, padding='max_length')
label_ids = labels.input_ids.squeeze().to(dtype=torch.long)
# 把padding对应的token_id替换为-100,避免模型学习padding
label_ids[label_ids == tokenizer.pad_token_id] = -100

2. 优化generate参数配置

  • 用max_new_tokens替代max_length,更直观控制生成的新token数量
  • 根据分词后的实际长度调整min_length(先打印目标文本分词长度确认)
  • 添加length_penalty鼓励模型生成更长输出

调整后的generate代码:

# 先确认目标文本分词后的实际token数
target_token_len = len(tokenizer(target_text)['input_ids'])
print(f"目标文本分词后长度: {target_token_len}")

generated_ids = model.generate(
    input_ids=ids,
    attention_mask=attn_mask,
    max_new_tokens=400,  # 对应256词的token数,可根据实际调整
    min_length=target_token_len,  # 设为训练时目标的实际分词长度
    num_beams=2,
    early_stopping=False,
    repetition_penalty=10.0,
    length_penalty=2.0  # 正数,数值越大越鼓励长输出
)

3. 简化解码逻辑

移除冗余的特殊token替换操作:

preds = tokenizer.decode(generated_ids[0], skip_special_tokens=True, clean_up_tokenization_spaces=True).strip()
target = tokenizer.decode(reference[0], skip_special_tokens=True, clean_up_tokenization_spaces=True).strip()

内容的提问来源于stack exchange,提问作者Tamir

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 08:55:12