You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

微调T5-Base模型生成合规问题时输出不完整的解决咨询

合规条款疑问句生成模型微调问题解答

问题背景

我正在微调t5-base模型,用于针对合规条款段落生成过去式的疑问句。已完成训练、保存并重新加载用于推理,但对未知语句生成问题时,输出内容不完整。

训练与推理代码

import pandas as pd
import torch
from datasets import Dataset
import transformers
from transformers import AutoTokenizer, AutoModelForSeq2SeqLM, Seq2SeqTrainingArguments, Seq2SeqTrainer, T5Tokenizer

df = pd.read_csv(r'/content/questionsgenerator.csv', encoding='unicode_escape')
df.head()

# Load pre-trained model and tokenizer
model_name = "t5-base"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForSeq2SeqLM.from_pretrained(model_name)

# Define the training arguments
training_args = Seq2SeqTrainingArguments(
    output_dir="./output_dir",
    per_device_train_batch_size=8,
    per_device_eval_batch_size=8,
    predict_with_generate=True,
    logging_steps=100,
    save_steps=5000,
    eval_steps=5000,
    num_train_epochs=3,
    learning_rate=1e-4,
    warmup_steps=1000,
    save_total_limit=3,
)

# Define the training dataset
train_dataset = Dataset.from_pandas(df.rename(columns={"Compliance Item": "input_text", "Question": "target_text"}))

# Define the function to preprocess the dataset
def preprocess_function(examples):
    inputs = [f"compliance item: {ci}" for ci in examples["input_text"]]
    targets = [f"{question} </s>" for question in examples["target_text"]]
    model_inputs = tokenizer(inputs, max_length=512, padding="max_length", truncation=True)
    with tokenizer.as_target_tokenizer():
        labels = tokenizer(targets, max_length=32, padding="max_length", truncation=True)
    model_inputs["labels"] = labels["input_ids"]
    return model_inputs

# Preprocess the dataset
train_dataset = train_dataset.map(preprocess_function, batched=True)

# Define the trainer
trainer = Seq2SeqTrainer(
    model=model,
    args=training_args,
    train_dataset=train_dataset,
)

# Fine-tune the model on the dataset
trainer.train()

model.save_pretrained("./fine_tuned_model_question_generation")

tokenizer = T5Tokenizer.from_pretrained("t5-large")
model = transformers.AutoModelForSeq2SeqLM.from_pretrained("./fine_tuned_model_question_generation")

context = 'When the Installment Due Date falls on a non-business day, the Mortgagee must consider a Borrower’s Notice of Intent to Prepay or the receipt of the prepayment amount for a Mortgage closed before January 21, 2015 timely if received on the next business day.'

encoding = tokenizer.encode_plus(context, return_tensors="pt")

input_ids = encoding["input_ids"]
attention_mask = encoding["attention_mask"]

output = model.generate(input_ids=input_ids, attention_mask=attention_mask, max_length=1000)
decoded_output = tokenizer.decode(output[0], skip_special_tokens=True)

decoded_output

模型输出问题

模型生成的不完整输出:
When the Installment Due Date fell on a non-business day, was the Borrower’s Notice of Intent to Prepay or the receipt of the prepayment amount for

疑问解答

1. 是否需要增加训练轮数?

是否增加轮数要结合训练过程的指标判断:

  • 如果训练时损失仍在持续下降,验证集的生成质量也在提升,说明模型还没学透任务模式,可以尝试把num_train_epochs调到5-8轮,继续训练。
  • 如果训练损失已经平稳,甚至验证集出现过拟合(比如生成内容重复、偏离任务要求),那增加轮数反而会恶化效果,此时要换其他优化方向。
    另外,你的预处理代码中,目标文本的max_length设为32,这会限制模型学习生成更长的问题,而推理时却设置了max_length=1000,这种不匹配也是输出不完整的重要原因——建议把训练时labels的max_length调高到128或256,匹配推理时的预期输出长度。

2. 是否有更适合该任务的模型?

针对合规条款的疑问句生成任务,这些模型表现更优:

  • T5-large/xl:更大参数规模的T5模型,对复杂正式文本的理解和生成能力更强,微调后能更稳定地处理长文本问题生成。注意你当前推理时错误加载了t5-large的tokenizer,应该和训练时保持一致(用t5-base的tokenizer),避免token不匹配导致的生成异常。
  • facebook/bart-large:BART模型擅长文本生成与理解,在问答、摘要类任务上表现突出,适配合规条款这类正式文本的生成需求。
  • FLAN-T5系列:经过指令微调的T5模型,已经适配了多种文本生成任务,针对问题生成这类指令式任务,微调效率更高,生成效果更稳定。

内容的提问来源于stack exchange,提问作者Daremitsu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 13:58:37