如何通过llama.cpp Web Server传递Llama 3格式完整对话Prompt
实现llama.cpp Web Server支持Llama 3完整对话历史的方法
1. 明确Llama 3的官方Prompt格式
Llama 3的对话格式需严格遵循以下结构(包含指定特殊token):
<|begin_of_text|> <|start_header_id|>system<|end_header_id|> {系统提示内容} <|eot_id|> <|start_header_id|>user<|end_header_id|> {用户提问1} <|eot_id|> <|start_header_id|>assistant<|end_header_id|> {AI回复1} <|eot_id|> <|start_header_id|>user<|end_header_id|> {用户当前提问} <|eot_id|> <|start_header_id|>assistant<|end_header_id|>
2. 本地维护对话历史
在客户端代码中,用数组维护对话历史,每个元素包含role(user/assistant)和content字段:
conversation_history = [ {"role": "user", "content": "你好,介绍下你自己"}, {"role": "assistant", "content": "我是基于Llama 3的AI助手,很高兴为你服务!"} ]
3. 动态拼接符合格式的完整Prompt
用户发起新提问时,按Llama 3格式拼接系统提示、历史对话和当前提问:
- 先拼接系统提示块(如有)
- 遍历历史对话,交替拼接用户、助手的对话块
- 最后拼接当前用户提问块和助手角色起始标记(供模型生成回复)
示例Python代码:
def build_llama3_prompt(system_prompt, conversation_history, current_user_input): prompt_parts = ["<|begin_of_text|>"] # 添加系统提示 if system_prompt: prompt_parts.extend([ "<|start_header_id|>system<|end_header_id|>\n\n", system_prompt, "<|eot_id|>\n" ]) # 添加历史对话 for msg in conversation_history: role = msg["role"] content = msg["content"] prompt_parts.extend([ f"<|start_header_id|>{role}<|end_header_id|>\n\n", content, "<|eot_id|>\n" ]) # 添加当前用户输入 prompt_parts.extend([ "<|start_header_id|>user<|end_header_id|>\n\n", current_user_input, "<|eot_id|>\n", "<|start_header_id|>assistant<|end_header_id|>\n\n" ]) return "".join(prompt_parts)
4. 调用llama.cpp Web Server接口
将拼接好的完整Prompt传给/completion接口的prompt参数,同时将system_prompt设为空字符串(已手动拼接系统提示),并设置停止token为<|eot_id|>避免生成多余内容。
示例curl请求:
curl http://localhost:8080/completion \ -H "Content-Type: application/json" \ -d '{ "prompt": "<|begin_of_text|><|start_header_id|>system<|end_header_id|>\n\n你是一个专业助手<|eot_id|>\n<|start_header_id|>user<|end_header_id|>\n\n你好<|eot_id|>\n<|start_header_id|>assistant<|end_header_id|>\n\n你好!有什么可以帮你的?<|eot_id|>\n<|start_header_id|>user<|end_header_id|>\n\n介绍下Llama 3<|eot_id|>\n<|start_header_id|>assistant<|end_header_id|>\n\n", "system_prompt": "", "stop": ["<|eot_id|>"], "temperature": 0.7 }'
5. 更新对话历史
收到模型回复后,将当前用户提问和模型回复追加到对话历史数组,供下一次请求使用。
内容的提问来源于stack exchange,提问作者user2741831
相关产品推荐
相关产品推荐

