You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

SageMakerEndpoint结合Langchain调用HuggingFace模型输出异常排查

问题排查:Langchain调用SageMaker部署的HuggingFace模型无输出

问题描述

将HuggingFace模型部署在Amazon SageMaker终端节点后,直接使用predictor.predict()调用可得到符合预期的输出,但通过Langchain的SagemakerEndpoint类结合RetrievalQA调用时,模型仅返回空内容或少量字符,核对Langchain官方文档及示例配置后仍未解决问题。

复现代码

初始化代码

endpoint = "xxxxxx-2023-07-14-05-34-901"

parameters = {
    "do_sample": True,
    "top_p": 0.95,
    "temperature": 0.1,
    "max_new_tokens": 256,
    "num_return_sequences": 4,
}
    
class ContentHandler(LLMContentHandler):
        content_type = "application/json"
        accepts = "application/json"
        
        def transform_input(self, prompt: str, model_kwargs: Dict) -> bytes:
            input_str = json.dumps({"inputs": prompt, **model_kwargs})
            return input_str.encode('utf-8')
        
        def transform_output(self, output: bytes) -> str:
            response_json = json.loads(output.read().decode("utf-8"))
            return response_json[0]['generated_text']
        
content_handler = ContentHandler()
        
sm_llm=SagemakerEndpoint(
        endpoint_name=endpoint,
        region_name="us-west-2",
        model_kwargs= parameters,
        content_handler=content_handler,
    )

vectordb = Chroma(persist_directory="db", embedding_function=embedding, collection_name="docs")
retriever = vectordb.as_retriever(search_kwargs={'k':3})

qa_chain = RetrievalQA.from_chain_type(llm=sm_llm, 
                                  chain_type="stuff", 
                                  retriever=retriever,
                                  return_source_documents=True)

调用代码

system_prompt = """<|SYSTEM|># Your are a helpful and harmless assistant for providing clear and succint answers to questions."""

question = "What is your purpose?"
query = (system_prompt + "<|USER|>" + question + "<|ASSISTANT|>")
llm_response = qa_chain(query)
print(llm_response['result'])

实际输出

Use the following pieces of context to answer the question at the end. If you don't know the answer, just say that you don't know, don't try to make up an answer.

<some doc context from vectordb goes here - removed due to info sensitivity>

Question: # Your are a helpful and harmless assistant for providing clear and succint answers to questions.
What is your purpose?
Helpful Answer:

排查方向与解决方案

1. 提示词格式冲突

RetrievalQA的stuff链会自动生成固定结构的提示模板(包含上下文、Question、Helpful Answer字段),但你将带特殊token(<|SYSTEM|>等)的完整提示词作为query传入,导致最终传给模型的prompt是链模板与自定义prompt的混合,模型无法正确识别指令逻辑。

解决方法:自定义适配模型的提示模板,替换RetrievalQA的默认模板:

from langchain.prompts import PromptTemplate

# 自定义符合模型要求的提示模板,整合上下文与特殊token
prompt_template = """<|SYSTEM|># Your are a helpful and harmless assistant for providing clear and succint answers to questions.
Use the following pieces of context to answer the question at the end. If you don't know the answer, just say that you don't know, don't try to make up an answer.

{context}

<|USER|>{question}<|ASSISTANT|>"""

PROMPT = PromptTemplate(
    template=prompt_template, input_variables=["context", "question"]
)

# 创建QA链时指定自定义模板
qa_chain = RetrievalQA.from_chain_type(
    llm=sm_llm,
    chain_type="stuff",
    retriever=retriever,
    return_source_documents=True,
    chain_type_kwargs={"prompt": PROMPT}
)

# 调用时直接传入问题,无需拼接完整prompt
question = "What is your purpose?"
llm_response = qa_chain({"query": question})
print(llm_response['result'])

2. 模型参数配置问题

你设置了num_return_sequences:4,但RetrievalQA仅需要单个结果,虽然ContentHandler中取了第一个结果,但部分模型可能因多序列生成的逻辑影响单序列输出的完整性。

解决方法:将parameters中的num_return_sequences改为1,测试输出是否正常:

parameters = {
    "do_sample": True,
    "top_p": 0.95,
    "temperature": 0.1,
    "max_new_tokens": 256,
    "num_return_sequences": 1,  # 修改为1
}

3. 输出处理逻辑验证

在ContentHandler.transform_output中添加日志,确认SageMaker返回的原始数据结构是否符合预期,避免因解析错误导致空输出:

def transform_output(self, output: bytes) -> str:
    response_json = json.loads(output.read().decode("utf-8"))
    # 添加日志查看原始返回结构
    print("SageMaker原始返回:", response_json)
    return response_json[0]['generated_text']

内容的提问来源于stack exchange,提问作者AxeCap

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.16 14:52:47