如何通过Python的elasticsearch库获取ElasticSearch索引字段的全部索引词?
获取Elasticsearch文本字段的所有已索引词(Python实现)
有两种简便的方法可以实现这个需求,具体选择取决于你的数据量和ES版本:
方法一:使用Terms聚合(适合小数据量)
如果字段数据量不大(不超过ES默认的max_result_window,通常为10000),可以用Terms聚合直接获取所有词。注意:如果操作的是text字段,需要先开启fielddata(仅推荐小字段使用,否则会占用大量内存);如果要获取不分词的完整值,建议使用对应的keyword子字段。
Python代码示例
from elasticsearch import Elasticsearch # 连接ES实例 es = Elasticsearch(["http://your-es-host:9200"]) # 索引名和目标字段 index_name = "your_index" target_field = "your_text_field" # 若要获取不分词的值,替换为"your_text_field.keyword" # 构造聚合查询 query = { "size": 0, # 无需返回文档,只获取聚合结果 "aggs": { "all_terms": { "terms": { "field": target_field, "size": 10000 # 设置足够大的数值,不超过max_result_window即可 } } } } # 执行查询 response = es.search(index=index_name, body=query) # 提取所有已索引词 all_terms = [bucket["key"] for bucket in response["aggregations"]["all_terms"]["buckets"]] print(all_terms)
方法二:使用_terms_enum API(适合大数据量,ES 7.10+)
如果数据量很大,超过max_result_window限制,推荐用_terms_enum API,它支持分页获取所有词,不会受窗口限制,性能更优。
Python代码示例
from elasticsearch import Elasticsearch es = Elasticsearch(["http://your-es-host:9200"]) index_name = "your_index" target_field = "your_text_field" # 同样可替换为keyword子字段 all_terms = [] current_after = None while True: # 构造请求参数 params = { "index": index_name, "field": target_field, "size": 1000 # 每次批量获取的词数量 } if current_after: params["after"] = current_after # 调用_terms_enum API response = es.termvectors.terms_enum(**params) # 添加当前页的词到列表 all_terms.extend(response["terms"]) # 判断是否还有下一页 if "after_term" not in response: break current_after = response["after_term"] print(all_terms)
注意事项
- 若操作text字段,开启
fielddata会占用堆内存,大字段不建议这么做,最好提前在mapping中设置keyword子字段(示例:"your_text_field": {"type": "text", "fields": {"keyword": {"type": "keyword"}}}),然后通过your_text_field.keyword获取不分词的完整值。 _terms_enumAPI仅支持ES 7.10及以上版本,低版本可使用Terms聚合结合scroll的方式实现,但效率不如前者。
内容的提问来源于stack exchange,提问作者Miguel
相关产品推荐
相关产品推荐

