Elasticsearch使用ngram实现simple_query_string子词匹配无结果排查
问题核心原因
一共有4处配置/调用错误,按影响优先级排序:
- 索引配置结构层级错误
你定义的settings变量结构不符合Elasticsearch规范,analysis属于settings的子属性,你当前代码里把analysis和settings、mappings平级,导致自定义的ngram分词器完全没有注册生效,filename字段实际用了默认的standard分词器,只会拆分完整单词,不会生成子词索引。 analysis节点下的属性命名错误
Elasticsearch的analysis节点下的分词器配置统一放在analyzer属性下,你代码里写的index_analyzer、search_analyzer属于非法顶层属性,就算结构层级对了也不会生效。- 未指定查询的目标字段
simple_query_string没有配置fields参数,默认不会定向到filename字段查询,自然匹配不到对应内容。 - 搜索分词器配置错误
你给搜索分词器也加了ngram过滤器属于冗余配置,搜索时不需要对搜索词做ngram拆分,直接用小写分词即可,否则会生成多余子词影响匹配结果。
修复后的可运行代码
from elasticsearch import Elasticsearch # 自行替换为你的ES连接信息 es = Elasticsearch(["http://localhost:9200"]) test_mapping = { "properties": { "filename": { "type": "text", # 索引时用带ngram的分词器生成子词 "analyzer": "my_index_analyzer", # 搜索时用普通小写分词器,不拆分搜索词 "search_analyzer": "my_search_analyzer" }, } } def create_index(index_name, mapping): created = False # 修正后的完整配置结构 settings = { "settings": { "number_of_shards": 1, "number_of_replicas": 0, "analysis": { "analyzer": { "my_index_analyzer": { "type": "custom", "tokenizer": "standard", "filter": [ "lowercase", "mynGram" ] }, "my_search_analyzer": { "type": "custom", "tokenizer": "standard", "filter": [ "lowercase" ] } }, "filter": { "mynGram": { "type": "nGram", "min_gram": 2, # 原配置50过大,调整为20足够覆盖常用搜索词长度,减少冗余索引 "max_gram": 20 } } } }, "mappings": mapping } try: # 先删除已存在的错误索引,避免旧配置干扰 if es.indices.exists(index_name): es.indices.delete(index=index_name) es.indices.create(index=index_name, ignore=400, body=settings) print(f'Created Index: {index_name}') created = True except Exception as ex: print(str(ex)) finally: return created create_index("test", test_mapping) doc = { 'filename': r"C:\Users\Sven Onderbeke\Documents\Arduino", } # 写入后加refresh参数,强制刷新索引避免实时查询不到 es.index(index="test", document=doc, refresh=True) needle = "ocumen" q = { "simple_query_string": { "query": needle, # 明确指定搜索filename字段 "fields": ["filename"], "default_operator": "and" } } res = es.search(index="test", query=q) print(res) for hit in res['hits']['hits']: print(hit)
效果验证
你可以调用ES的分词测试接口确认配置是否生效:
POST test/_analyze { "analyzer": "my_index_analyzer", "text": "Documents" }
返回的tokens列表中如果包含ocumen子词,就说明索引分词配置正常,搜索即可命中目标文档。
内容的提问来源于stack exchange,提问作者Niya
相关产品推荐
相关产品推荐

