开源LLM生成时持续重复令牌至最大令牌数,如何解决?
解决方案
针对你遇到的葡萄牙语LLM生成重复令牌或无意义内容直至达到最大token限制的问题,可尝试以下关键调整:
- 启用采样模式:
temperature和top_p参数仅在do_sample=True时生效,默认的贪婪搜索极易导致重复生成,必须显式开启采样才能让随机性参数发挥作用。 - 正确配置令牌:确保
pad_token被正确设置,很多开源模型默认未定义pad_token,需将其指定为eos_token,避免生成过程中无法正确终止。 - 优化重复惩罚与采样参数:在启用
repetition_penalty的同时调整参数值,配合top_k等采样策略,平衡随机性与生成质量,避免出现无意义表情符号。 - 严格遵循模型Prompt模板:确认输入格式完全符合目标模型的要求,该模型基于Mistral Chat,需保证
<s>[INST]...[/INST]的结构正确。
修改后的代码示例
# Import necessary libraries from transformers import AutoModelForCausalLM, AutoTokenizer, TextStreamer, BitsAndBytesConfig import torch # Set up 4-bit quantization configuration bnb_4bit_config = BitsAndBytesConfig( load_in_4bit=True, bnb_4bit_quant_type="nf4", bnb_4bit_compute_dtype=torch.bfloat16, bnb_4bit_use_double_quant=True ) # Load the model with 4-bit quantization model = AutoModelForCausalLM.from_pretrained( "rhaymison/Mistral-portuguese-luana-7b-chat", quantization_config=bnb_4bit_config, device_map={"": 0} ) # Load the tokenizer and set pad_token tokenizer = AutoTokenizer.from_pretrained("rhaymison/Mistral-portuguese-luana-7b-chat") # 显式设置pad_token为eos_token,避免生成时无法终止 tokenizer.pad_token = tokenizer.eos_token # Set the model to evaluation mode model.eval() # Define the input prompt (严格遵循模型的Chat模板) prompt = """<s>[INST] Qual o seu nome? [/INST]""" # Tokenize the input inputs = tokenizer([prompt], return_tensors="pt").to(model.device) # Set up the text streamer streamer = TextStreamer(tokenizer, skip_prompt=False, skip_special_tokens=False) # Generate the response with adjusted parameters generation_config = { 'max_new_tokens': 128, 'do_sample': True, # 必须开启采样,否则temperature/top_p不生效 'temperature': 0.6, # 适度降低温度,减少过度随机的无意义内容 'top_p': 0.85, 'top_k': 50, # 限制候选令牌数量,提升生成质量 'repetition_penalty': 1.15, # 调整惩罚值,避免重复同时减少乱码 'eos_token_id': tokenizer.eos_token_id, 'pad_token_id': tokenizer.pad_token_id } # Generate the response _ = model.generate(**inputs, streamer=streamer, **generation_config)
额外建议
- 可尝试调整
repetition_penalty的数值(范围1.0-1.5),找到最适合该模型的平衡点。 - 如果仍出现无意义内容,可尝试去掉4-bit量化,用完整精度加载模型,排除量化带来的潜在问题。
- 查看模型官方文档,确认是否有特定的生成参数要求或Prompt格式细节。
内容的提问来源于stack exchange,提问作者Miguel Casagrande
相关产品推荐
相关产品推荐

