如何在Elasticsearch中实现Term Centric+短语检索?
问题:如何以Term Centric方式实现跨字段短语检索?
Elasticsearch的「multi-match + cross_fields」模式实现的是Term Centric检索,这个特性很实用,因为不需要per-field-IDF。但我想以Term Centric的方式执行短语检索,发现cross_fields和combined_fields都不支持短语检索,有没有其他实现方式?
我试过「multi-match + phrase」模式,但它采用的是per-field-IDF,不符合需求。
详细示例
1. 创建索引
PUT /testindex { "settings": { "number_of_shards": 1, "number_of_replicas": 0 }, "mappings": { "properties": { "fn": { "type": "text" }, "ln": { "type": "text" } } } }
2. 添加10条相同文档
POST /testindex/_doc/{from 1 to 10} { "fn": "hello world", "ln": "bye" }
3. 添加另一条文档,将「hello world」放入ln字段
POST /testindex/_doc/100 { "fn": "bye", "ln": "hello world" }
4. 执行短语检索
我希望对「hello world」执行短语检索,使用如下查询:
POST /testindex/_search { "query": { "bool": { "must": { "multi_match": { "query": "hello world", "fields": ["fn", "ln"], "type": "phrase" } } } } }
但「multi-match+phrase」是字段中心式的,「hello world」在ln字段中的IDF值极高,导致ID为100的文档得分远高于其他文档,这不是我想要的结果:
{ "took": 5, "timed_out": false, "_shards": { "total": 1, "successful": 1, "skipped": 0, "failed": 0 }, "hits": { "total": { "value": 11, "relation": "eq" }, "max_score": 3.1015396, "hits": [ { "_index": "testindex", "_type": "_doc", "_id": "100", "_score": 3.1015396, "_source": { "fn": "bye", "ln": "hello world" } }, { "_index": "testindex", "_type": "_doc", "_id": "1", "_score": 0.26195967, "_source": { "fn": "hello world", "ln": "bye" } }, ... ] } }
内容的提问来源于stack exchange,提问作者Michael Liu
相关产品推荐
相关产品推荐

