You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过llama.cpp Web Server传递Llama 3格式完整对话Prompt

实现llama.cpp Web Server支持Llama 3完整对话历史的方法

1. 明确Llama 3的官方Prompt格式

Llama 3的对话格式需严格遵循以下结构(包含指定特殊token):

<|begin_of_text|>
<|start_header_id|>system<|end_header_id|>

{系统提示内容}
<|eot_id|>
<|start_header_id|>user<|end_header_id|>

{用户提问1}
<|eot_id|>
<|start_header_id|>assistant<|end_header_id|>

{AI回复1}
<|eot_id|>
<|start_header_id|>user<|end_header_id|>

{用户当前提问}
<|eot_id|>
<|start_header_id|>assistant<|end_header_id|>

2. 本地维护对话历史

在客户端代码中,用数组维护对话历史,每个元素包含role(user/assistant)和content字段:

conversation_history = [
    {"role": "user", "content": "你好,介绍下你自己"},
    {"role": "assistant", "content": "我是基于Llama 3的AI助手,很高兴为你服务!"}
]

3. 动态拼接符合格式的完整Prompt

用户发起新提问时,按Llama 3格式拼接系统提示、历史对话和当前提问:

  • 先拼接系统提示块(如有)
  • 遍历历史对话,交替拼接用户、助手的对话块
  • 最后拼接当前用户提问块和助手角色起始标记(供模型生成回复)

示例Python代码:

def build_llama3_prompt(system_prompt, conversation_history, current_user_input):
    prompt_parts = ["<|begin_of_text|>"]
    # 添加系统提示
    if system_prompt:
        prompt_parts.extend([
            "<|start_header_id|>system<|end_header_id|>\n\n",
            system_prompt,
            "<|eot_id|>\n"
        ])
    # 添加历史对话
    for msg in conversation_history:
        role = msg["role"]
        content = msg["content"]
        prompt_parts.extend([
            f"<|start_header_id|>{role}<|end_header_id|>\n\n",
            content,
            "<|eot_id|>\n"
        ])
    # 添加当前用户输入
    prompt_parts.extend([
        "<|start_header_id|>user<|end_header_id|>\n\n",
        current_user_input,
        "<|eot_id|>\n",
        "<|start_header_id|>assistant<|end_header_id|>\n\n"
    ])
    return "".join(prompt_parts)

4. 调用llama.cpp Web Server接口

将拼接好的完整Prompt传给/completion接口的prompt参数,同时将system_prompt设为空字符串(已手动拼接系统提示),并设置停止token为<|eot_id|>避免生成多余内容。

示例curl请求:

curl http://localhost:8080/completion \
  -H "Content-Type: application/json" \
  -d '{
    "prompt": "<|begin_of_text|><|start_header_id|>system<|end_header_id|>\n\n你是一个专业助手<|eot_id|>\n<|start_header_id|>user<|end_header_id|>\n\n你好<|eot_id|>\n<|start_header_id|>assistant<|end_header_id|>\n\n你好!有什么可以帮你的?<|eot_id|>\n<|start_header_id|>user<|end_header_id|>\n\n介绍下Llama 3<|eot_id|>\n<|start_header_id|>assistant<|end_header_id|>\n\n",
    "system_prompt": "",
    "stop": ["<|eot_id|>"],
    "temperature": 0.7
  }'

5. 更新对话历史

收到模型回复后,将当前用户提问和模型回复追加到对话历史数组,供下一次请求使用。

内容的提问来源于stack exchange,提问作者user2741831

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.23 07:00:03