如何在Elasticsearch中统计索引文档内指定关键词的出现次数?
统计Elasticsearch文档内指定关键词的出现次数
首先,先看你提供的示例文档:
{ "_index": "testdata", "_type": "tweet", "_id": "Dbo5qmMBSUBLqBARJmBG", "_version": 1, "_score": 1, "_source": { "fileName": "alibaba.pdf", "chapter": "chapter1", "page": 1, "timeDate": "2018-05-24T11:06:48+00:00", "text": "So why do we need machine learning, why do we want a machine to learn as a human? There are many problems involving huge datasets, or complex calculations for instance, where it makes sense to let computers do all the work. In general, of course, computers and robots dont get tired, dont have to sleep, and may be cheaper. There is also an emerging school of thought called active learning or human-in-the-loop, which advocates combining the efforts of machine learners and humans. The idea is that there are routine boring tasks more suitable for computers, and creative tasks more suitable for humans.According to this philosophy, machines are able to learn, by following rules or algorithms designed by humans and to do repetitive and logic tasks desired by a human" } }
要统计指定关键词(比如machine、humans)在文档的text字段里的出现次数,有几种实用的方法:
方法1:查询时使用脚本字段实时计算
如果只需要针对单个或少量文档统计,可以在查询中添加script_fields,用Painless脚本直接统计关键词出现次数:
GET testdata/_search { "query": { "match": { "_id": "Dbo5qmMBSUBLqBARJmBG" } }, "script_fields": { "machine_count": { "script": { "source": "def text = params._source.text; return text.split('machine').length - 1;" } }, "humans_count": { "script": { "source": "def text = params._source.text; return text.split('humans').length - 1;" } } } }
这个脚本的逻辑很简单:用关键词分割文本,分割后的数组长度减1就是关键词出现的次数。注意这里是精确匹配,如果需要忽略大小写,可以先把文本转小写再处理:
def text = params._source.text.toLowerCase(); return text.split('machine').length - 1;
方法2:利用分词和术语聚合(适合批量文档)
如果你的text字段已经被正确分词(比如用标准分词器),可以用terms聚合来统计关键词在多个文档中的总出现次数,或者针对单个文档:
首先确认你的text字段的分词设置,假设是标准分词,那么可以这样查询:
GET testdata/_search { "query": { "match": { "_id": "Dbo5qmMBSUBLqBARJmBG" } }, "aggs": { "keyword_counts": { "terms": { "field": "text", "include": ["machine", "humans"], "size": 10 } } } }
不过要注意:这种方法依赖分词结果,如果关键词被分词器拆分(比如复合词),可能得不到准确结果。如果需要精确匹配,建议给text字段添加一个不分词的子字段(比如text.keyword),然后聚合这个子字段:
GET testdata/_search { "query": { "match": { "_id": "Dbo5qmMBSUBLqBARJmBG" } }, "aggs": { "keyword_counts": { "terms": { "field": "text.keyword", "include": ".*machine.*|.*humans.*", "size": 10 } } } }
不过这种正则匹配的方式效率不如脚本字段,适合批量统计场景。
方法3:预计算并存储次数(适合频繁查询)
如果需要频繁统计这些关键词,最好的方式是在索引文档时就预计算好次数,存储到一个新字段里,比如keyword_counts:
{ "_source": { "fileName": "alibaba.pdf", // ... 其他字段 "text": "...", "keyword_counts": { "machine": 3, "humans": 2 } } }
这样查询时直接读取这个字段即可,性能最优。你可以在客户端索引前计算,或者用Elasticsearch的 ingest pipeline 自动计算。
内容的提问来源于stack exchange,提问作者Prashant Patel
相关产品推荐
相关产品推荐

