You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何替换Hugging Face Transformers中废弃的conversational pipeline实现聊天机器人?

用最新版Transformers构建对话式聊天机器人的方案

Conversational pipeline已被官方移除,当前推荐采用text-generation pipeline或直接调用模型的generate方法实现对话功能,核心是利用tokenizer的apply_chat_template方法统一处理对话历史,生成符合模型要求的输入格式。

方式一:使用text-generation pipeline

该方式封装度高,快速上手,适合快速搭建对话原型:

from transformers import AutoModelForCausalLM, AutoTokenizer, pipeline

# 加载模型与tokenizer(替换为你使用的模型ID)
model_id = "meta-llama/Llama-2-7b-chat-hf"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id)

# 初始化text-generation pipeline,设置生成参数
chat_pipeline = pipeline(
    "text-generation",
    model=model,
    tokenizer=tokenizer,
    max_new_tokens=512,
    temperature=0.7,
    top_p=0.95,
    return_full_text=False  # 只返回新生成的回复,不包含输入的对话历史
)

# 定义初始对话历史
chat_history = [
    {"role": "system", "content": "You are a smart chatbot"},
    {"role": "user", "content": "What is the capital of France?"},
]

# 用tokenizer处理对话历史,生成模型可识别的输入格式
prompt = tokenizer.apply_chat_template(chat_history, tokenize=False, add_generation_prompt=True)

# 生成回复
response = chat_pipeline(prompt)[0]['generated_text']

# 将机器人回复加入对话历史,用于后续多轮对话
chat_history.append({"role": "assistant", "content": response})

print(f"Assistant: {response}")

方式二:直接调用model.generate方法

该方式更灵活,可自定义更多生成参数,适合需要精细控制的场景:

from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

# 加载模型与tokenizer(替换为你使用的模型ID)
model_id = "meta-llama/Llama-2-7b-chat-hf"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype=torch.float16).to("cuda")  # 若有GPU则用cuda

# 定义初始对话历史
chat_history = [
    {"role": "system", "content": "You are a smart chatbot"},
    {"role": "user", "content": "What is the capital of France?"},
]

# 处理对话历史,生成模型输入的张量
inputs = tokenizer.apply_chat_template(
    chat_history,
    tokenize=True,
    return_tensors="pt",
    add_generation_prompt=True
).to("cuda")  # 与模型设备一致

# 调用generate方法生成回复
with torch.no_grad():
    outputs = model.generate(
        inputs,
        max_new_tokens=512,
        temperature=0.7,
        top_p=0.95,
        do_sample=True
    )

# 解码输出,提取机器人回复
response = tokenizer.decode(outputs[0][len(inputs[0]):], skip_special_tokens=True)

# 更新对话历史
chat_history.append({"role": "assistant", "content": response})

print(f"Assistant: {response}")

关键说明

  • apply_chat_template是核心方法:它会根据模型的预设对话格式(如Llama的<s>[INST]...[/INST]、GPT的<|system|>...<|user|>...<|assistant|>)自动格式化对话历史,无需手动拼接字符串。
  • add_generation_prompt=True会在对话历史末尾添加模型所需的生成触发标记(如assistant角色的前缀),确保模型知道要生成助手的回复。
  • 多轮对话只需重复将新的用户提问和助手回复加入chat_history,再用apply_chat_template处理即可。

内容的提问来源于stack exchange,提问作者celsowm

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.16 22:40:02