You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用LlamaIndex与Llama3-70B-Instruct做序列匹配的结果偏差问题

解决Llama3-70B-Instruct返回非指定术语的问题

你遇到的核心问题是模型对“匹配术语”的理解偏差,加上prompt和输入格式的模糊性,导致返回了包含目标术语的长短语而非指定术语本身。以下是具体解决办法:

  • 严格约束Prompt的匹配规则
    修改Prompt模板,明确要求精确匹配给定术语本身,禁止返回包含术语的更长短语。调整后的模板示例:

    template = (
    "We have provided context information below. \n"
    "---------------------
    "
    "{context_str}"
    "\n---------------------
    "
    "Given this information, please find all **exact terms from the context list** that appear in the input sentence. You must only return the exact terms provided, NOT any extended phrases containing them. \n"
    "\n---------------------
    "
    "{query_str}"
    "\n---------------------
    "
    "Identify and return the exact terms from the context list found in the input sentence as list_cancer_terms. Additionally, return list_of_sentences where each term is preceded by the seven words before it and followed by the seven words after it in the input sentence. list_cancer_terms must only include terms from the context list, no extra words allowed. Provide the response in JSON format only."
    )
    
  • 优化术语列表的呈现格式
    原代码直接传入列表data_cancer,模型可能无法正确识别为待匹配术语。将列表转为清晰的条目格式:

    # 替换原context_str生成方式
    context_str = "\n".join([f"- {term}" for term in data_cancer])
    prompt = qa_template.format(context_str=context_str, query_str=result_sentence, pydantic_schema=pydantic_schema)
    
  • 强化系统提示的约束性
    更新system prompt,明确告知模型必须严格遵守规则:

    messages = [
        {"role": "system", "content": "You are a precise assistant that strictly follows instructions. You must only return the exact terms provided in the context list. Do NOT return any extended phrases containing these terms. All responses must adhere to the specified JSON schema:\n<schema>\n{{" + pydantic_schema + "}}\n<schema>\n"},
        {"role": "user", "content": prompt}
    ]
    
  • 用Pydantic Schema强制约束返回值
    定义Schema时明确限定术语范围,让模型更清晰返回要求:

    from pydantic import BaseModel, Literal
    
    class CancerMatchResponse(BaseModel):
        list_cancer_terms: list[Literal["MYELOMA", "carcinoma"]]
        list_of_sentences: list[str]
    
    # 生成schema字符串传入prompt
    pydantic_schema = CancerMatchResponse.model_json_schema()
    
  • 代码层面增加后处理过滤
    作为兜底方案,在模型返回结果后手动过滤不符合要求的术语:

    # 假设解析后的模型响应为response_data
    response_data = ... # 解析JSON响应的代码
    response_data["list_cancer_terms"] = [term for term in response_data["list_cancer_terms"] if term in data_cancer]
    

内容的提问来源于stack exchange,提问作者joshpopelka20

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 12:25:22