You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于AWS Bedrock+Llama3的聊天机器人重复对话历史问题求助

问题:Amazon Bedrock + Llama3 + LangChain 聊天机器人重复完整对话历史

我使用Amazon Bedrock搭配Llama3模型开发聊天机器人应用,前端采用Streamlit框架,对话管理依赖LangChain工具。目前遇到异常:机器人回复时会重复完整的对话历史内容,而非仅针对当前用户问题给出直接回答。

当前表现

用户提出问题后,机器人的回复包含:

  • 用户当前的问题
  • 机器人对当前问题的回答
  • 对话历史中的所有过往问题
  • 对话历史中的所有过往回答

示例回复:

Human: Do you know what a llama is?
Assistant: Yes, I do know what a llama is. A llama is a domesticated South American camelid, widely used as a meat and pack animal by Andean cultures since the Pre-Columbian era.
Human: What is the average lifespan of a llama?
Assistant: According to my knowledge, the average lifespan of a llama is between 20 and 30 years. However, some llamas have been known to live up to 40 years or more with proper care and nutrition.
Human: Do you know the average weight of a llama?
Assistant: Yes, I do know the average weight of a llama. The average weight of a llama is between 280 and 450 pounds (127 to 204 kilograms), with some males reaching up to 500 pounds (227 kilograms) or more.

我期望机器人仅返回针对当前问题的直接回答,而非完整对话链。

简化版代码

import streamlit as st
from langchain.llms import Bedrock
from langchain.chains import ConversationChain
from langchain.memory import ConversationBufferWindowMemory
from langchain.prompts.prompt import PromptTemplate
from langchain.memory.chat_message_histories import StreamlitChatMessageHistory
from langchain.callbacks.base import BaseCallbackHandler
import boto3
from langchain.prompts.chat import (
            ChatPromptTemplate,
            SystemMessagePromptTemplate,
            AIMessagePromptTemplate,
            HumanMessagePromptTemplate,
)

bedrock_rt = boto3.client(
            "bedrock-runtime", 
            region_name="us-east-1",
        )

DEFAULT_CLAUDE_TEMPLATE = """
The following is a friendly conversation between a human and an AI. 
The AI is talkative and provides lots of specific details from its context. 
If the AI does not know the answer to a question, it truthfully says it does not know.

Just Answer the questions and don't add something extra.

Current conversation:
{history}
Human: {input}
Assistant:"""

CLAUDE_PROMPT = PromptTemplate(
    input_variables=["history", "input"], template=DEFAULT_CLAUDE_TEMPLATE)

INIT_MESSAGE = {"role": "assistant",
                "content": "Hi! I'm Claude on Bedrock. How may I help you?"}


class StreamHandler(BaseCallbackHandler):
    def __init__(self, container):
        self.container = container
        self.text = ""

    def on_llm_new_token(self, token: str, **kwargs) -> None:
        self.text += token
        self.container.markdown(self.text)


# Set Streamlit page configuration
st.set_page_config(page_title='🤖 Chat with Bedrock', layout='wide')
st.title("🤖 Chat with Bedrock")

# Sidebar info
with st.sidebar:
    st.markdown("## Inference Parameters")
    TEMPERATURE = st.slider("Temperature", min_value=0.0,
                            max_value=1.0, value=0.1, step=0.1)
    TOP_P = st.slider("Top-P", min_value=0.0,
                      max_value=1.0, value=0.9, step=0.01)
    TOP_K = st.slider("Top-K", min_value=1,
                      max_value=500, value=10, step=5)
    MAX_TOKENS = st.slider("Max Token", min_value=0,
                           max_value=2048, value=1024, step=8)
    MEMORY_WINDOW = st.slider("Memory Window", min_value=0,
                              max_value=10, value=3, step=1)


# Initialize the ConversationChain
def init_conversationchain() -> ConversationChain:
    model_kwargs = {'temperature': TEMPERATURE,
                    'top_p': TOP_P,
                    # 'top_k': TOP_K,
                    'max_gen_len': MAX_TOKENS}

    llm = Bedrock(
        client=bedrock_rt,
        model_id="meta.llama3-8b-instruct-v1:0",
        model_kwargs=model_kwargs,
        streaming=True
    )
    system_message_prompt = SystemMessagePromptTemplate.from_template(DEFAULT_CLAUDE_TEMPLATE)

    example_human_history = HumanMessagePromptTemplate.from_template("Hi")
    example_ai_history = AIMessagePromptTemplate.from_template("hello, how are you today?")

    human_template="{input}"
    human_message_prompt = HumanMessagePromptTemplate.from_template(human_template)
    

    conversation = ConversationChain(
        llm=llm,
        verbose=True,
        memory=ConversationBufferWindowMemory(
            k=MEMORY_WINDOW, ai_prefix="Assistant", chat_memory=StreamlitChatMessageHistory()),
        prompt=CLAUDE_PROMPT
    )

    # Store LLM generated responses

    if "messages" not in st.session_state.keys():
        st.session_state.messages = [INIT_MESSAGE]

    return conversation


def generate_response(conversation, input_text):
    return conversation.run(input=input_text, callbacks=[StreamHandler(st.empty())])


# Re-initialize the chat
def new_chat() -> None:
    st.session_state["messages"] = [INIT_MESSAGE]
    st.session_state["langchain_messages"] = []
    conv_chain = init_conversationchain()


# Add a button to start a new chat
st.sidebar.button("New Chat", on_click=new_chat, type='primary')

# Initialize the chat
conv_chain = init_conversationchain()

# Display chat messages
for message in st.session_state.messages:
    with st.chat_message(message["role"]):
        st.markdown(message["content"])

# User-provided prompt
prompt = st.chat_input()

if prompt:
    st.session_state.messages.append({"role": "user", "content": prompt})
    with st.chat_message("user"):
        st.markdown(prompt)

# Generate a new response if last message is not from assistant
if st.session_state.messages[-1]["role"] != "assistant":
    with st.chat_message("assistant"):
        # print(st.session_state.messages)
        response = generate_response(conv_chain, prompt)
    message = {"role": "assistant", "content": response}
    st.session_state.messages.append(message)

解决方案

问题根源

当前使用的DEFAULT_CLAUDE_TEMPLATE是为Claude模型设计的,而Llama3指令模型有专属的提示格式要求。如果不匹配模型预期的格式,Llama3会错误地将对话历史视为生成内容的一部分,导致重复输出。

修复步骤

  1. 替换为Llama3专用提示模板
    重新编写模板,明确告知模型仅输出当前问题的回答,避免重复历史:

    DEFAULT_LLAMA3_TEMPLATE = """
    You are a helpful assistant. Answer the user's question directly based on the conversation history. Do NOT repeat any previous conversation content.
    
    Conversation History:
    {history}
    
    User's current question: {input}
    Assistant's answer:"""
    
  2. 更新Prompt配置
    将原有的CLAUDE_PROMPT替换为新的Llama3模板,并在ConversationChain中使用:

    LLAMA3_PROMPT = PromptTemplate(
        input_variables=["history", "input"], template=DEFAULT_LLAMA3_TEMPLATE)
    
    # 初始化ConversationChain时替换prompt参数
    conversation = ConversationChain(
        llm=llm,
        verbose=True,
        memory=ConversationBufferWindowMemory(
            k=MEMORY_WINDOW, ai_prefix="Assistant", chat_memory=StreamlitChatMessageHistory()),
        prompt=LLAMA3_PROMPT
    )
    
  3. 添加停止序列约束(可选)
    在模型参数中设置stop_sequences,强制模型在生成完回答后停止,避免额外输出:

    model_kwargs = {
        'temperature': 0.0,  # 降低温度提升输出确定性
        'top_p': TOP_P,
        'max_gen_len': MAX_TOKENS,
        'stop_sequences': ["\nHuman:", "\nUser:"]  # 遇到用户提问前缀时停止生成
    }
    

验证效果

修改后,机器人会仅返回当前问题的直接回答,例如询问 llama 平均体重时,只会输出:

Yes, I do know the average weight of a llama. The average weight of a llama is between 280 and 450 pounds (127 to 204 kilograms), with some males reaching up to 500 pounds (227 kilograms) or more.

内容的提问来源于stack exchange,提问作者rahul raj

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 18:34:53