You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Langchain+DeepLake: SelfQueryRetriever含文件名查询触发TQL错误

问题:DeepLake结合SelfQueryRetriever检索含文件名的查询时触发错误

问题描述

使用DeepLake向量数据库存储项目代码片段,通过SelfQueryRetriever根据需求问题检索对应代码块时,当查询中包含train.py script这类文件名表达式会触发错误,移除该类表达式则运行正常。由于需适配所有问题场景,无法规避此类表达式。

自定义检索器代码

def CustomRetriever(files, dataset_path,issue):
    metadata_field_info = [
        AttributeInfo(
            name="source",
            description="The soruce file the chunk was extracted from",
            type="string",
        ),
        AttributeInfo(
            name="file_name",
            description="The name of the file the chunk was extracted from",
            type="string",
        ),
        AttributeInfo(
            name="chunk_id",
            description="the id of the chunk",
            type="string",
        ),
    ]
    document_content_description = "The sourcecode of a project"
    model = ChatOpenAI(model="gpt-4")
    embeddings = OpenAIEmbeddings(disallowed_special=())
    db = DeepLake(dataset_path=dataset_path, read_only=True, embedding=embeddings, exec_option='python')
    docs = (db.similarity_search(query=" ", k=10000000))
    retriever = SelfQueryRetriever.from_llm(
        model, db, document_content_description, metadata_field_info, verbose=True
    )
    try:
        print('TEST', retriever.get_relevant_documents(
            f"Which documents contain code to resolve the following issue? -> {issue}"))
    except ValueError as e:
        print(traceback.format_exc())

报错信息

Traceback (most recent call last):
  File "/Users/kaanerbay/GitHub/Github_Issue_Solver/langchainLogic/retriever2.py", line 93, in CustomRetriever
    print('TEST', retriever.get_relevant_documents(
  File "/Users/kaanerbay/miniconda3/envs/main/lib/python3.10/site-packages/langchain/schema/retriever.py", line 208, in get_relevant_documents
    raise e
  File "/Users/kaanerbay/miniconda3/envs/main/lib/python3.10/site-packages/langchain/schema/retriever.py", line 201, in get_relevant_documents
    result = self._get_relevant_documents(
  File "/Users/kaanerbay/miniconda3/envs/main/lib/python3.10/site-packages/langchain/retrievers/self_query/base.py", line 135, in _get_relevant_documents
    docs = self.vectorstore.search(new_query, self.search_type, **search_kwargs)
  File "/Users/kaanerbay/miniconda3/envs/main/lib/python3.10/site-packages/langchain/vectorstores/base.py", line 121, in search
    return self.similarity_search(query, **kwargs)
  File "/Users/kaanerbay/miniconda3/envs/main/lib/python3.10/site-packages/langchain/vectorstores/deeplake.py", line 475, in similarity_search
    return self._search(
  File "/Users/kaanerbay/miniconda3/envs/main/lib/python3.10/site-packages/langchain/vectorstores/deeplake.py", line 348, in _search
    return self._search_tql(
  File "/Users/kaanerbay/miniconda3/envs/main/lib/python3.10/site-packages/langchain/vectorstores/deeplake.py", line 267, in _search_tql
    result = self.vectorstore.search(
  File "/Users/kaanerbay/miniconda3/envs/main/lib/python3.10/site-packages/deeplake/core/vectorstore/deeplake_vectorstore.py", line 429, in search
    utils.parse_search_args(
  File "/Users/kaanerbay/miniconda3/envs/main/lib/python3.10/site-packages/deeplake/core/vectorstore/vector_search/utils.py", line 229, in parse_search_args
    raise ValueError(
ValueError: User-specified TQL queries are not support for exec_option=python.

错误核心原因:当查询包含文件名时,SelfQueryRetriever会自动生成基于元数据的TQL过滤查询,但当前DeepLake实例使用exec_option='python'模式,该模式不支持TQL查询。

测试用问题

应该在train.py脚本中使用CNN替代BERT模型,因为它更适合处理这类数据。
CNN不能太复杂也不能太简单,需用TensorFlow生成。
要将CNN集成到现有逻辑中,并根据所用的词向量进行适配,尽可能优化代码。

解决方案

方案1:修改DeepLake的执行模式(推荐)

移除exec_option='python'参数(默认使用tql模式),或者显式设置为tql,这样就能支持SelfQueryRetriever生成的TQL元数据过滤查询:

# 修改后的DeepLake初始化代码
db = DeepLake(dataset_path=dataset_path, read_only=True, embedding=embeddings)
# 或者显式指定exec_option='tql'
# db = DeepLake(dataset_path=dataset_path, read_only=True, embedding=embeddings, exec_option='tql')

方案2:强制SelfQueryRetriever仅执行向量检索

如果必须使用exec_option='python',可以通过配置强制SelfQueryRetriever不生成元数据过滤条件,但这种方式会丢失基于文件名的精准过滤能力,仅能依赖向量相似性检索:

retriever = SelfQueryRetriever.from_llm(
    model, db, document_content_description, metadata_field_info, 
    verbose=True, enable_limit=False
)

内容的提问来源于stack exchange,提问作者alpa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.10 16:05:54