嵌套字段过滤下Elasticsearch KNN查询结果异常问题
Elasticsearch向量查询添加用户过滤后分数异常偏高问题
索引结构
"properties": { "id": {"type": "text"}, "date": {"type": "text"}, "title": {"type": "text"}, "users": {"type": "nested"}, "ideas": { "type": "nested", "properties": { "vector": { "type": "knn_vector", "dimension": 384 }, "timestamp": {"type": "integer"}, "content": { "type": "keyword", }, } } }
初始向量查询(无用户过滤)
该查询用于搜索与给定向量相似度高的ideas:
body={ "_source": { "includes": [ "title", "users", "ideas._score" ] }, "size": 1, "query": { "bool": { "must": [ { "nested": { "path": "ideas", "query": { "bool": { "must": [ { "script_score": { "query": { "match_all": {} }, "script": { "source": "knn_score", "lang": "knn", "params": { "field": "ideas.vector", "query_value": query_vector, "space_type": "cosinesimil" } } } } ] } } } } ] } } }
返回结果:
{"hits": {"_id": "kHCOaIkBUHe9xr-u09Xr", "_score": 1.1901431, "_source": {"title": "doc1", "users": [{"id": "user1"}]}}}
添加用户过滤后的查询及异常结果
为过滤特定用户可访问的文档,修改查询添加users嵌套过滤:
body={ "_source": { "includes": [ "title", "users", "ideas._score" ] }, "size": 1, "query": { "bool": { "must": [ { "nested": { "path": "users", "query": { "match_phrase": { "users.id": user_id } } } }, { "nested": { "path": "ideas", "query": { "bool": { "must": [ { "script_score": { "query": { "match_all": {} }, "script": { "source": "knn_score", "lang": "knn", "params": { "field": "ideas.vector", "query_value": query_vector, "space_type": "cosinesimil" } } } } ] } } } } ] } } }
当传入user_id="user1"时,预期返回与初始查询相同的结果,但实际返回:
{"hits": {"_id": "kXCOaIkBUHe9xr-u2tV1", "_score": 2.866434, "_source": {"title": "doc2", "users": [{"id": "user1"}]}}}
异常表现:分数远超1-2的合理范围,且返回了无关文档。
问题原因及解决办法
原因
Elasticsearch的bool查询中must子句的分数默认是相加的:
- 第一个
users嵌套的match_phrase查询会生成一个匹配分数(通常大于1) - 第二个
ideas嵌套的script_score返回余弦相似度分数(1-2之间)
两者累加导致总分数超出合理范围,同时排序逻辑被打乱,原本相似度最高的文档可能因用户匹配分数的叠加被挤下去。
另外,users过滤属于文档级权限控制,仅需过滤不符合条件的文档,不应参与分数计算。
解决办法
将users过滤从bool.must移到bool.filter中,filter子句仅过滤文档,不贡献分数:
body={ "_source": { "includes": [ "title", "users", "ideas._score" ] }, "size": 1, "query": { "bool": { "filter": [ { "nested": { "path": "users", "query": { "match_phrase": { "users.id": user_id } } } } ], "must": [ { "nested": { "path": "ideas", "query": { "bool": { "must": [ { "script_score": { "query": { "match_all": {} }, "script": { "source": "knn_score", "lang": "knn", "params": { "field": "ideas.vector", "query_value": query_vector, "space_type": "cosinesimil" } } } } ] } } } } ] } } }
修改后,总分数仅由ideas的向量相似度分数决定,既实现用户权限过滤,又保证排序逻辑与初始查询一致。
内容的提问来源于stack exchange,提问作者AlwaysLearning
相关产品推荐
相关产品推荐

