You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在LCEL中基于链前序步骤动态设置search_kwargs?

解决LCEL中动态设置检索过滤条件的问题

问题出在你直接将itemgetter('page_nums')传入了检索器的search_kwargs,ChromaDB无法识别这个操作符对象,因为它需要的是实际的页码列表。要在LCEL中实现动态过滤,需要用运行时动态生成检索参数的方式,而不是提前固定search_kwargs。

修改后的代码实现

user_input = {"question": "What is SomeKeyword?"}

# 1. 关键词检索步骤不变
keyword_retrieval = RunnableParallel(
  keywords=itemgetter("question") | keywords.as_retriever(
    search_type="similarity_score_threshold",
    search_kwargs={"score_threshold": 0.4}),
  question=itemgetter("question"))

# 2. 提取页码步骤不变
page_nums = RunnableParallel(
  page_nums=itemgetter("keywords") | RunnableLambda(get_page_nums),
  question=itemgetter("question"))

# 3. 动态生成文档检索步骤:用RunnableLambda包装,运行时获取页码列表并设置过滤条件
def retrieve_docs_with_pages(inputs):
    question = inputs["question"]
    pages = inputs["page_nums"]
    # 动态创建带过滤条件的检索器
    retriever = docs.as_retriever(
        search_kwargs={"filter": {"page": {"$in": pages}}}
    )
    return {"context": retriever.invoke(question), "question": question}

docs_retrieval = RunnableLambda(retrieve_docs_with_pages)

# 组装完整链
chain = keyword_retrieval | page_nums | docs_retrieval | prompt | llm

关键原理

  • 不能在初始化检索器时直接引用链中后续才会生成的变量(比如page_nums),因为as_retriever()是在代码加载时执行的,此时还没有运行时的上下文数据。
  • 用RunnableLambda包装检索逻辑,可以在链运行到这一步时,拿到前序步骤生成的page_nums,再动态创建带过滤条件的检索器并执行检索。

额外优化提示

如果你的get_page_nums函数返回的是字符串格式的页码(比如"23,29"),记得先转换成整数列表:

def get_page_nums(keyword_docs):
    pages = []
    for doc in keyword_docs:
        # 拆分字符串并转成整数
        pages.extend(int(p.strip()) for p in doc.metadata['pages'].split(','))
    # 去重避免重复检索同一页码
    return list(set(pages))

内容的提问来源于stack exchange,提问作者btonasse

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.25 07:31:19