如何从OpenSearch Langchain的similarity_search_with_score获取文档_id?
问题描述
我用similarity_search_with_score做相似度搜索时,返回结果里没有文档的_id,返回格式为(Document(page_content="", metadata={embedding=[], content=""}), )。我明确每个文档都有_id——往OpenSearch添加文档时它会返回这个ID,我也已经把ID存在PostgreSQL里了。
试过两个方案但都有阻碍:
- 想把
_id放进metadata,但这个ID是添加文档后才生成的,不知道怎么实现 - 想把PostgreSQL的行ID放进metadata,但业务逻辑要求先加OpenSearch再存PostgreSQL,不想调整顺序
现在需要拿到相似文档的_id,求其他可行方案。
当前代码
添加文档函数
def add_document(opensearch_client, index_name, embedding, content): document_data = { "embedding": embedding, "content": content } response = opensearch_client.index(index=index_name, body=document_data) logger.info("Document added to opensearch") return response['_id']
搜索函数
def search_vector_db(query, _is_aoss=False): session = boto3.Session() credentials = session.get_credentials() aws_auth = AWS4Auth(credentials.access_key, credentials.secret_key, "ap-southeast-1", 'es', session_token=credentials.token) opensearch_endpoint = get_opensearch_endpoint("vector-kb", "ap-southeast-1") docsearch = OpenSearchVectorSearch( index_name="vector-kb-index", embedding_function=get_openai_embedding_client(), opensearch_url=f"https://{opensearch_endpoint}", http_auth=aws_auth, timeout=30, is_aoss=_is_aoss, connection_class=RequestsHttpConnection, use_ssl=True, verify_certs=True, ) docs = docsearch.similarity_search_with_score( query, search_type="script_scoring", space_type="cosinesimil", vector_field="embedding", text_field="content", score_threshold=1.5 ) contexts = [] for doc in docs: logger.info(doc) contexts.append(doc[0].page_content) logger.info("Similar documents retrieved from Opensearch for context") return contexts
可行解决方案
方案1:添加文档后立即更新,把_id写入文档字段
拿到OpenSearch返回的_id后,立刻调用更新接口把_id写入文档的metadata或独立字段中。修改add_document函数:
def add_document(opensearch_client, index_name, embedding, content): document_data = { "embedding": embedding, "content": content } # 先索引文档获取_id response = opensearch_client.index(index=index_name, body=document_data) doc_id = response['_id'] # 更新文档,将_id加入metadata update_body = { "doc": { "metadata": { "opensearch_id": doc_id } } } opensearch_client.update(index=index_name, id=doc_id, body=update_body) logger.info("Document added and updated with opensearch_id to opensearch") return doc_id
之后搜索返回的Document对象的metadata里就会包含opensearch_id字段,直接提取即可。
方案2:改用OpenSearch原生查询,直接获取_id
OpenSearchVectorSearch的封装方法可能没包含_id返回,直接用原生OpenSearch客户端构造查询,能直接拿到_id。修改搜索函数:
def search_vector_db(query, _is_aoss=False): session = boto3.Session() credentials = session.get_credentials() aws_auth = AWS4Auth(credentials.access_key, credentials.secret_key, "ap-southeast-1", 'es', session_token=credentials.token) opensearch_endpoint = get_opensearch_endpoint("vector-kb", "ap-southeast-1") # 初始化原生OpenSearch客户端 opensearch_client = OpenSearch( hosts=[{"host": opensearch_endpoint, "port": 443}], http_auth=aws_auth, use_ssl=True, verify_certs=True, connection_class=RequestsHttpConnection, timeout=30 ) # 生成查询向量 query_embedding = get_openai_embedding_client().embed_query(query) # 构造向量搜索查询,指定返回字段(_id会自动返回) search_query = { "query": { "script_score": { "query": {"match_all": {}}, "script": { "source": "cosineSimilarity(params.query_vector, 'embedding') + 1.0", "params": {"query_vector": query_embedding} } } }, "_source": ["content", "embedding"], "size": 10, "min_score": 1.5 } response = opensearch_client.search(index="vector-kb-index", body=search_query) contexts = [] for hit in response['hits']['hits']: doc_id = hit['_id'] content = hit['_source']['content'] score = hit['_score'] logger.info(f"Found doc: id={doc_id}, score={score}, content={content[:50]}...") contexts.append({"id": doc_id, "content": content, "score": score}) logger.info("Similar documents retrieved from Opensearch for context") return contexts
这种方式直接控制查询逻辑,能拿到完整的搜索结果,包括_id。
方案3:通过PostgreSQL关联查询_id
如果不想修改OpenSearch的文档结构,可以在拿到搜索返回的内容后,去PostgreSQL中通过内容关联查询对应的OpenSearch _id。为了避免长内容查询效率低,建议存内容的哈希值到PostgreSQL:
- 存储PostgreSQL时添加哈希字段:
import hashlib # 生成内容哈希 content_hash = hashlib.md5(content.encode('utf-8')).hexdigest() # 存入PostgreSQL:opensearch_id, content, content_hash
- 搜索时通过哈希查
_id:
# 在search_vector_db函数中处理 import hashlib contexts = [] for doc in docs: content = doc[0].page_content content_hash = hashlib.md5(content.encode('utf-8')).hexdigest() # 执行PostgreSQL查询:SELECT opensearch_id FROM your_table WHERE content_hash = %s # 假设拿到opensearch_id后存入结果 contexts.append({"id": opensearch_id, "content": content})
这个方案适合无法修改OpenSearch操作的场景,但需要确保内容哈希的唯一性。
内容的提问来源于stack exchange,提问作者chulin
相关产品推荐
相关产品推荐

