Elasticsearch如何统计索引text字段中指定单词的全局总出现次数
解决方案
前置说明
你之前方案失败的核心原因有两个:
- 你的索引映射里没有配置
content.keyword子字段,所以访问该字段会直接报错 - text类型默认关闭了
doc_values,无法通过doc['content']直接读取字段值,只能从_source中获取
以下两个方案都可以满足需求,可根据作业数据量选择:
方案1:Python侧统计(小数据量首选,实现最简单)
如果作业的索引数据量不大,直接滚动遍历所有匹配文档,在Python侧累加词频即可:
实现步骤
- 用滚动查询拉取所有包含目标词的文档的
content字段,避免深分页限制 - 对每个文档的
content内容统计目标词的出现次数,累加得到总和
示例代码
from elasticsearch import Elasticsearch # 初始化ES连接 es = Elasticsearch(["http://127.0.0.1:9200"]) INDEX_NAME = "你的索引名" TARGET_WORD = "ustawa" total = 0 # 初始化滚动查询,保留1分钟上下文,每次拉取100条 resp = es.search( index=INDEX_NAME, query={"match": {"content": TARGET_WORD}}, scroll="1m", size=100, _source=["content"] ) scroll_id = resp["_scroll_id"] # 遍历所有结果 while True: # 处理当前批次的文档 for hit in resp["hits"]["hits"]: content = hit["_source"].get("content", "") # 如需和ES自定义分词器结果完全匹配,可替换为调用_analyze接口分词后统计 total += content.count(TARGET_WORD) # 拉取下一批 if len(resp["hits"]["hits"]) < 100: break resp = es.scroll(scroll_id=scroll_id, scroll="1m") # 清理滚动上下文 es.clear_scroll(scroll_id=scroll_id) print(f"总出现次数:{total}")
方案2:ES侧直接统计(无需遍历数据,一次查询返回结果)
如果不想在Python侧做计算,可以用ES的runtime字段+sum聚合直接得到结果,查询语句如下:
{ "size": 0, "query": { "match": { "content": "ustawa" } }, "runtime_mappings": { "target_word_count": { "type": "long", "script": { "source": """ String content = ctx['_source'].content; if (content == null) {emit(0); return;} int count = 0; int idx = 0; while((idx = content.indexOf(params.word, idx)) != -1) { count++; idx += params.word.length(); } emit(count); """, "params": { "word": "ustawa" } } } }, "aggs": { "total_count": { "sum": { "field": "target_word_count" } } } }
查询返回结果中aggregations.total_count.value就是总出现次数,Python侧直接读取该值即可。
内容的提问来源于stack exchange,提问作者qalis
相关产品推荐
相关产品推荐

