ElasticSearch索引如何关联字段实现跨字段高亮匹配并返回对应时间
实现方案
你当前的索引结构无法直接实现该需求,核心原因是文本字段text的分词结果和start字段里的时间戳没有绑定对应关系,ES无法感知哪个词对应哪个时间节点。可以通过调整索引结构+配合查询实现需求,首选方案如下:
1. 调整索引Mapping设计
将逐词和对应的时间戳绑定为嵌套对象存储,Mapping示例:
{ "mappings": { "properties": { "id": { "type": "keyword" }, "full_text": { "type": "text", "store": true }, "word_timestamps": { "type": "nested", "properties": { "word": { "type": "text" }, "start": { "type": "keyword" } } } } } }
2. 写入数据时做预处理
把原始的text按空格拆分得到单个词,和start按空格拆分后的时间戳一一对应,存入word_timestamps数组。你给出的示例数据写入后结构如下:
{ "id": "xxx-xxx", "full_text": "ALICE was beginning to get very tired of sitting by her sister on the bank, and of having nothing to do", "word_timestamps": [ {"word": "ALICE", "start": "0:00:01.94"}, {"word": "was", "start": "0:00:02.30"}, // 中间其他词省略 {"word": "very", "start": "0:00:03.30"}, {"word": "tired", "start": "0:00:03.42"}, // 后续其他词省略 ] }
3. 构造查询语句
使用nested短语查询+inner_hits返回匹配到的词和对应时间,查询语句示例:
{ "query": { "nested": { "path": "word_timestamps", "query": { "match_phrase": { "word_timestamps.word": "very tired" } }, "inner_hits": { "_source": ["word_timestamps.word", "word_timestamps.start"] } } } }
查询返回结果的inner_hits字段中会包含匹配到的very和tired两个对象,直接拼接即可得到你期望的输出结果。
备选方案
如果不想调整索引结构,也可以使用ES的payload分词插件,在自定义分词阶段把每个词的时间戳作为附加 payload 存入词条,查询时提取匹配词条的payload即可,该方案需要自定义分词器,开发成本略高于上述首选方案。
内容的提问来源于stack exchange,提问作者Regitzu
相关产品推荐
相关产品推荐

