Elasticsearch:如何仅在非严格查询中使用asciifolding过滤器?
问题:Elasticsearch实现非严格/严格字符匹配需求
需要实现两种查询效果:
- 非严格查询(输入
Stephane):可匹配Stephane和Stéphane(当前已正常工作) - 严格查询(输入
"Stephane"):仅匹配Stephane(当前无法实现,会同时匹配两者)
疑问:是否可以设置asciifolding过滤器仅用于非严格查询?或者是否需要为严格查询使用不同的分析器?
当前索引设置
protected function createIndex(string $indexName, array $properties): array { return [ 'index' => $indexName, 'body' => [ 'settings' => [ 'number_of_shards' => 1, 'number_of_replicas' => 0, 'analysis' => [ 'analyzer' => [ 'special_chars' => [ 'type' => 'custom', 'tokenizer' => 'standard', 'filter' => [ 'preserve_asciifolding', ], ], ], 'filter' => [ 'preserve_asciifolding' => [ 'type' => 'asciifolding', 'preserve_original' => TRUE, ], ], ], ], 'mappings' => [ '_source' => [ 'enabled' => TRUE, ], 'properties' => $properties, ], ], ]; }
当前查询代码
{ "_source": [ "first_name", "last_name", "title", "lead", ... ], "from": 0, "size": 20, "query": { "bool": { "must": [ { "query_string": { "query": "\"Stephane\"", "default_operator": "AND" } }, { "nested": { "path": "target", "query": { "bool": { "should": [ { "match": { "target.id": 1 } } ] } } } }, { "nested": { "path": "location", "query": { "bool": { "should": [ { "match": { "location.id": 1 } } ] } } } } ], "filter": [], "should": [] } }, "aggs": { ..... } } }
解决方案
要实现严格/非严格两种匹配,最佳方案是为目标字段设置多字段映射:一个字段使用带asciifolding的分析器(用于非严格查询),另一个字段使用不包含该过滤器的分析器(用于严格查询)。具体步骤如下:
1. 修改索引设置,添加无asciifolding的分析器
在现有的analysis配置中,新增一个仅用标准分词器的分析器:
protected function createIndex(string $indexName, array $properties): array { return [ 'index' => $indexName, 'body' => [ 'settings' => [ 'number_of_shards' => 1, 'number_of_replicas' => 0, 'analysis' => [ 'analyzer' => [ 'special_chars' => [ 'type' => 'custom', 'tokenizer' => 'standard', 'filter' => [ 'preserve_asciifolding', ], ], // 新增:无asciifolding的分析器,用于严格匹配 'standard_strict' => [ 'type' => 'custom', 'tokenizer' => 'standard', 'filter' => [ 'lowercase' // 若需要严格区分大小写可去掉此过滤器 ], ], ], 'filter' => [ 'preserve_asciifolding' => [ 'type' => 'asciifolding', 'preserve_original' => TRUE, ], ], ], ], 'mappings' => [ '_source' => [ 'enabled' => TRUE, ], 'properties' => $properties, ], ], ]; }
2. 修改字段映射,添加多字段
假设目标字段是first_name,在properties中配置多字段:
// 示例$properties中的字段配置 'first_name' => [ 'type' => 'text', 'analyzer' => 'special_chars', // 默认使用非严格分析器 'fields' => [ 'strict' => [ 'type' => 'text', 'analyzer' => 'standard_strict' // 严格匹配字段使用无asciifolding的分析器 ] ] ]
3. 调整查询逻辑
- 非严格查询:直接查询原字段(如
first_name),保持现有逻辑即可。 - 严格查询:查询新增的
.strict字段,比如使用query_string时指定字段:
{ "query": { "bool": { "must": [ { "query_string": { "query": "\"Stephane\"", "default_operator": "AND", "fields": ["first_name.strict", "last_name.strict"] // 指定严格匹配字段 } }, // ... 其他嵌套查询保持不变 ] } } }
关键说明
- 无法直接让
asciifolding过滤器仅作用于非严格查询,因为分析逻辑是在索引阶段就确定的,查询时必须匹配索引的分词结果。 - 多字段映射相当于为同一内容生成两种不同的分词版本,分别服务不同的查询需求,是Elasticsearch处理这类场景的标准方案。
内容的提问来源于stack exchange,提问作者Adam Bodera
相关产品推荐
相关产品推荐

