You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从OpenSearch Langchain的similarity_search_with_score获取文档_id?

问题描述

我用similarity_search_with_score做相似度搜索时,返回结果里没有文档的_id,返回格式为(Document(page_content="", metadata={embedding=[], content=""}), )。我明确每个文档都有_id——往OpenSearch添加文档时它会返回这个ID,我也已经把ID存在PostgreSQL里了。

试过两个方案但都有阻碍:

  • 想把_id放进metadata,但这个ID是添加文档后才生成的,不知道怎么实现
  • 想把PostgreSQL的行ID放进metadata,但业务逻辑要求先加OpenSearch再存PostgreSQL,不想调整顺序

现在需要拿到相似文档的_id,求其他可行方案。

当前代码

添加文档函数

def add_document(opensearch_client, index_name, embedding, content):
    document_data = {
        "embedding": embedding,
        "content": content
    }
    response = opensearch_client.index(index=index_name, body=document_data)
    logger.info("Document added to opensearch")
    return response['_id']

搜索函数

def search_vector_db(query, _is_aoss=False):
    session = boto3.Session()
    credentials = session.get_credentials()
    aws_auth = AWS4Auth(credentials.access_key, credentials.secret_key, "ap-southeast-1", 'es', session_token=credentials.token)

    opensearch_endpoint = get_opensearch_endpoint("vector-kb", "ap-southeast-1")

    docsearch = OpenSearchVectorSearch(
        index_name="vector-kb-index",
        embedding_function=get_openai_embedding_client(),
        opensearch_url=f"https://{opensearch_endpoint}",
        http_auth=aws_auth,
        timeout=30,
        is_aoss=_is_aoss,
        connection_class=RequestsHttpConnection,
        use_ssl=True,
        verify_certs=True,
    )

    docs = docsearch.similarity_search_with_score(
        query,
        search_type="script_scoring",
        space_type="cosinesimil",
        vector_field="embedding",
        text_field="content",
        score_threshold=1.5
    )

    contexts = []
    for doc in docs:
        logger.info(doc)
        contexts.append(doc[0].page_content)
    
    logger.info("Similar documents retrieved from Opensearch for context")
    return contexts

可行解决方案

方案1:添加文档后立即更新,把_id写入文档字段

拿到OpenSearch返回的_id后,立刻调用更新接口把_id写入文档的metadata或独立字段中。修改add_document函数:

def add_document(opensearch_client, index_name, embedding, content):
    document_data = {
        "embedding": embedding,
        "content": content
    }
    # 先索引文档获取_id
    response = opensearch_client.index(index=index_name, body=document_data)
    doc_id = response['_id']
    # 更新文档,将_id加入metadata
    update_body = {
        "doc": {
            "metadata": {
                "opensearch_id": doc_id
            }
        }
    }
    opensearch_client.update(index=index_name, id=doc_id, body=update_body)
    logger.info("Document added and updated with opensearch_id to opensearch")
    return doc_id

之后搜索返回的Document对象的metadata里就会包含opensearch_id字段,直接提取即可。

方案2:改用OpenSearch原生查询,直接获取_id

OpenSearchVectorSearch的封装方法可能没包含_id返回,直接用原生OpenSearch客户端构造查询,能直接拿到_id。修改搜索函数:

def search_vector_db(query, _is_aoss=False):
    session = boto3.Session()
    credentials = session.get_credentials()
    aws_auth = AWS4Auth(credentials.access_key, credentials.secret_key, "ap-southeast-1", 'es', session_token=credentials.token)

    opensearch_endpoint = get_opensearch_endpoint("vector-kb", "ap-southeast-1")
    # 初始化原生OpenSearch客户端
    opensearch_client = OpenSearch(
        hosts=[{"host": opensearch_endpoint, "port": 443}],
        http_auth=aws_auth,
        use_ssl=True,
        verify_certs=True,
        connection_class=RequestsHttpConnection,
        timeout=30
    )

    # 生成查询向量
    query_embedding = get_openai_embedding_client().embed_query(query)
    # 构造向量搜索查询,指定返回字段(_id会自动返回)
    search_query = {
        "query": {
            "script_score": {
                "query": {"match_all": {}},
                "script": {
                    "source": "cosineSimilarity(params.query_vector, 'embedding') + 1.0",
                    "params": {"query_vector": query_embedding}
                }
            }
        },
        "_source": ["content", "embedding"],
        "size": 10,
        "min_score": 1.5
    }

    response = opensearch_client.search(index="vector-kb-index", body=search_query)
    contexts = []
    for hit in response['hits']['hits']:
        doc_id = hit['_id']
        content = hit['_source']['content']
        score = hit['_score']
        logger.info(f"Found doc: id={doc_id}, score={score}, content={content[:50]}...")
        contexts.append({"id": doc_id, "content": content, "score": score})
    
    logger.info("Similar documents retrieved from Opensearch for context")
    return contexts

这种方式直接控制查询逻辑,能拿到完整的搜索结果,包括_id。

方案3:通过PostgreSQL关联查询_id

如果不想修改OpenSearch的文档结构,可以在拿到搜索返回的内容后,去PostgreSQL中通过内容关联查询对应的OpenSearch _id。为了避免长内容查询效率低,建议存内容的哈希值到PostgreSQL:

  1. 存储PostgreSQL时添加哈希字段:
import hashlib
# 生成内容哈希
content_hash = hashlib.md5(content.encode('utf-8')).hexdigest()
# 存入PostgreSQL:opensearch_id, content, content_hash
  1. 搜索时通过哈希查_id:
# 在search_vector_db函数中处理
import hashlib
contexts = []
for doc in docs:
    content = doc[0].page_content
    content_hash = hashlib.md5(content.encode('utf-8')).hexdigest()
    # 执行PostgreSQL查询:SELECT opensearch_id FROM your_table WHERE content_hash = %s
    # 假设拿到opensearch_id后存入结果
    contexts.append({"id": opensearch_id, "content": content})

这个方案适合无法修改OpenSearch操作的场景,但需要确保内容哈希的唯一性。

内容的提问来源于stack exchange,提问作者chulin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.21 01:41:07