You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

开源LLM生成时持续重复令牌至最大令牌数,如何解决?

解决方案

针对你遇到的葡萄牙语LLM生成重复令牌或无意义内容直至达到最大token限制的问题,可尝试以下关键调整:

  • 启用采样模式:temperature和top_p参数仅在do_sample=True时生效,默认的贪婪搜索极易导致重复生成,必须显式开启采样才能让随机性参数发挥作用。
  • 正确配置令牌:确保pad_token被正确设置,很多开源模型默认未定义pad_token,需将其指定为eos_token,避免生成过程中无法正确终止。
  • 优化重复惩罚与采样参数:在启用repetition_penalty的同时调整参数值,配合top_k等采样策略,平衡随机性与生成质量,避免出现无意义表情符号。
  • 严格遵循模型Prompt模板:确认输入格式完全符合目标模型的要求,该模型基于Mistral Chat,需保证<s>[INST]...[/INST]的结构正确。

修改后的代码示例

# Import necessary libraries
from transformers import AutoModelForCausalLM, AutoTokenizer, TextStreamer, BitsAndBytesConfig
import torch

# Set up 4-bit quantization configuration
bnb_4bit_config = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_quant_type="nf4",
    bnb_4bit_compute_dtype=torch.bfloat16,
    bnb_4bit_use_double_quant=True
)

# Load the model with 4-bit quantization
model = AutoModelForCausalLM.from_pretrained(
    "rhaymison/Mistral-portuguese-luana-7b-chat",
    quantization_config=bnb_4bit_config,
    device_map={"": 0}
)

# Load the tokenizer and set pad_token
tokenizer = AutoTokenizer.from_pretrained("rhaymison/Mistral-portuguese-luana-7b-chat")
# 显式设置pad_token为eos_token,避免生成时无法终止
tokenizer.pad_token = tokenizer.eos_token

# Set the model to evaluation mode
model.eval()

# Define the input prompt (严格遵循模型的Chat模板)
prompt = """<s>[INST] Qual o seu nome? [/INST]"""

# Tokenize the input
inputs = tokenizer([prompt], return_tensors="pt").to(model.device)

# Set up the text streamer
streamer = TextStreamer(tokenizer, skip_prompt=False, skip_special_tokens=False)

# Generate the response with adjusted parameters
generation_config = {
    'max_new_tokens': 128,
    'do_sample': True,  # 必须开启采样,否则temperature/top_p不生效
    'temperature': 0.6,  # 适度降低温度,减少过度随机的无意义内容
    'top_p': 0.85,
    'top_k': 50,  # 限制候选令牌数量,提升生成质量
    'repetition_penalty': 1.15,  # 调整惩罚值,避免重复同时减少乱码
    'eos_token_id': tokenizer.eos_token_id,
    'pad_token_id': tokenizer.pad_token_id
}

# Generate the response
_ = model.generate(**inputs, streamer=streamer, **generation_config)

额外建议

  • 可尝试调整repetition_penalty的数值(范围1.0-1.5),找到最适合该模型的平衡点。
  • 如果仍出现无意义内容,可尝试去掉4-bit量化,用完整精度加载模型,排除量化带来的潜在问题。
  • 查看模型官方文档,确认是否有特定的生成参数要求或Prompt格式细节。

内容的提问来源于stack exchange,提问作者Miguel Casagrande

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.23 03:43:30