如何替换Hugging Face Transformers中废弃的conversational pipeline实现聊天机器人?
用最新版Transformers构建对话式聊天机器人的方案
Conversational pipeline已被官方移除,当前推荐采用text-generation pipeline或直接调用模型的generate方法实现对话功能,核心是利用tokenizer的apply_chat_template方法统一处理对话历史,生成符合模型要求的输入格式。
方式一:使用text-generation pipeline
该方式封装度高,快速上手,适合快速搭建对话原型:
from transformers import AutoModelForCausalLM, AutoTokenizer, pipeline # 加载模型与tokenizer(替换为你使用的模型ID) model_id = "meta-llama/Llama-2-7b-chat-hf" tokenizer = AutoTokenizer.from_pretrained(model_id) model = AutoModelForCausalLM.from_pretrained(model_id) # 初始化text-generation pipeline,设置生成参数 chat_pipeline = pipeline( "text-generation", model=model, tokenizer=tokenizer, max_new_tokens=512, temperature=0.7, top_p=0.95, return_full_text=False # 只返回新生成的回复,不包含输入的对话历史 ) # 定义初始对话历史 chat_history = [ {"role": "system", "content": "You are a smart chatbot"}, {"role": "user", "content": "What is the capital of France?"}, ] # 用tokenizer处理对话历史,生成模型可识别的输入格式 prompt = tokenizer.apply_chat_template(chat_history, tokenize=False, add_generation_prompt=True) # 生成回复 response = chat_pipeline(prompt)[0]['generated_text'] # 将机器人回复加入对话历史,用于后续多轮对话 chat_history.append({"role": "assistant", "content": response}) print(f"Assistant: {response}")
方式二:直接调用model.generate方法
该方式更灵活,可自定义更多生成参数,适合需要精细控制的场景:
from transformers import AutoModelForCausalLM, AutoTokenizer import torch # 加载模型与tokenizer(替换为你使用的模型ID) model_id = "meta-llama/Llama-2-7b-chat-hf" tokenizer = AutoTokenizer.from_pretrained(model_id) model = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype=torch.float16).to("cuda") # 若有GPU则用cuda # 定义初始对话历史 chat_history = [ {"role": "system", "content": "You are a smart chatbot"}, {"role": "user", "content": "What is the capital of France?"}, ] # 处理对话历史,生成模型输入的张量 inputs = tokenizer.apply_chat_template( chat_history, tokenize=True, return_tensors="pt", add_generation_prompt=True ).to("cuda") # 与模型设备一致 # 调用generate方法生成回复 with torch.no_grad(): outputs = model.generate( inputs, max_new_tokens=512, temperature=0.7, top_p=0.95, do_sample=True ) # 解码输出,提取机器人回复 response = tokenizer.decode(outputs[0][len(inputs[0]):], skip_special_tokens=True) # 更新对话历史 chat_history.append({"role": "assistant", "content": response}) print(f"Assistant: {response}")
关键说明
apply_chat_template是核心方法:它会根据模型的预设对话格式(如Llama的<s>[INST]...[/INST]、GPT的<|system|>...<|user|>...<|assistant|>)自动格式化对话历史,无需手动拼接字符串。add_generation_prompt=True会在对话历史末尾添加模型所需的生成触发标记(如assistant角色的前缀),确保模型知道要生成助手的回复。- 多轮对话只需重复将新的用户提问和助手回复加入
chat_history,再用apply_chat_template处理即可。
内容的提问来源于stack exchange,提问作者celsowm
相关产品推荐
相关产品推荐

