Elasticsearch中如何扁平化嵌套子字段并实现关联匹配与时间返回?
嘿,这个需求我刚好碰到过,给你一步步拆解解决方案:
第一步:设计合适的索引结构
因为你需要保留txt和time的对应关系,同时能匹配连续的文本短语,所以我们要把content设为nested类型(这样每个txt和time的配对会被当作独立的子文档处理,避免关联混乱),同时给txt字段添加keyword子字段用来精确匹配完整字符串,另外新增一个full_content字段用来存储所有txt按时间顺序拼接的文本,方便短语匹配。
创建索引的命令:
PUT /my_documents { "mappings": { "properties": { "id": { "type": "integer" }, "content": { "type": "nested", "properties": { "txt": { "type": "text", "analyzer": "standard", "fields": { "keyword": { "type": "keyword" } } }, "time": { "type": "integer" } } }, "full_content": { "type": "text", "analyzer": "standard" } } } }
第二步:插入文档数据
插入时要手动把content里的txt按time顺序拼接成full_content的值(也可以用Elasticsearch的copy_to+脚本自动生成,但手动拼接更直观):
POST /my_documents/_doc/1 { "id": 1, "content": [ { "txt": "I", "time": 0 }, { "txt": "have", "time": 1 }, { "txt": "a book", "time": 2 }, { "txt": "do not match this block", "time": 3 } ], "full_content": "I have a book do not match this block" }
第三步:构建查询语句
我们需要先通过match_phrase匹配full_content里的目标短语,确保文档确实包含这个连续文本,再用nested查询结合inner_hits提取对应time值:
POST /my_documents/_search { "query": { "bool": { "must": [ // 匹配完整的连续短语 { "match_phrase": { "full_content": "I have a book" } }, // 匹配组成短语的三个txt,并返回对应的time { "nested": { "path": "content", "query": { "bool": { "should": [ { "match": { "content.txt": "I" } }, { "match": { "content.txt": "have" } }, // 用keyword子字段精确匹配"a book"整个字符串 { "match": { "content.txt.keyword": "a book" } } ], "minimum_should_match": 1 } }, "inner_hits": {} } } ] } } }
查询结果优化
如果只想返回time字段而不需要其他冗余内容,可以给inner_hits添加_source过滤:
"inner_hits": { "_source": ["content.time"] }
修改后返回的inner_hits会只包含匹配到的time值(0、1、2),完全符合你的需求。
内容的提问来源于stack exchange,提问作者David Zhao
相关产品推荐
相关产品推荐

