You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Elasticsearch:如何仅在非严格查询中使用asciifolding过滤器?

问题:Elasticsearch实现非严格/严格字符匹配需求

需要实现两种查询效果:

  • 非严格查询(输入Stephane):可匹配Stephane和Stéphane(当前已正常工作)
  • 严格查询(输入"Stephane"):仅匹配Stephane(当前无法实现,会同时匹配两者)

疑问:是否可以设置asciifolding过滤器仅用于非严格查询?或者是否需要为严格查询使用不同的分析器?

当前索引设置

protected function createIndex(string $indexName, array $properties): array {
    return [
      'index' => $indexName,
      'body' => [
        'settings' => [
          'number_of_shards' => 1,
          'number_of_replicas' => 0,
          'analysis' => [
            'analyzer' => [
              'special_chars' => [
                'type' => 'custom',
                'tokenizer' => 'standard',
                'filter' => [
                  'preserve_asciifolding',
                ],
              ],
            ],
            'filter' => [
              'preserve_asciifolding' => [
                'type' => 'asciifolding',
                'preserve_original' => TRUE,
              ],
            ],
          ],
        ],
        'mappings' => [
          '_source' => [
            'enabled' => TRUE,
          ],
          'properties' => $properties,
        ],
      ],
    ];
  }

当前查询代码

{
    "_source": [
        "first_name",
        "last_name",
        "title",
        "lead",
         ...
    ],
    "from": 0,
    "size": 20,
    "query": {
        "bool": {
            "must": [
                {
                    "query_string": {
                        "query": "\"Stephane\"",
                        "default_operator": "AND"
                    }
                },
                {
                    "nested": {
                        "path": "target",
                        "query": {
                            "bool": {
                                "should": [
                                    {
                                        "match": {
                                            "target.id": 1
                                        }
                                    }
                                ]
                            }
                        }
                    }
                },
                {
                    "nested": {
                        "path": "location",
                        "query": {
                            "bool": {
                                "should": [
                                    {
                                        "match": {
                                            "location.id": 1
                                        }
                                    }
                                ]
                            }
                        }
                    }
                }
            ],
            "filter": [],
            "should": []
        }
    },
    "aggs": {
      .....  
        }
    }
}

解决方案

要实现严格/非严格两种匹配,最佳方案是为目标字段设置多字段映射:一个字段使用带asciifolding的分析器(用于非严格查询),另一个字段使用不包含该过滤器的分析器(用于严格查询)。具体步骤如下:

1. 修改索引设置,添加无asciifolding的分析器

在现有的analysis配置中,新增一个仅用标准分词器的分析器:

protected function createIndex(string $indexName, array $properties): array {
    return [
      'index' => $indexName,
      'body' => [
        'settings' => [
          'number_of_shards' => 1,
          'number_of_replicas' => 0,
          'analysis' => [
            'analyzer' => [
              'special_chars' => [
                'type' => 'custom',
                'tokenizer' => 'standard',
                'filter' => [
                  'preserve_asciifolding',
                ],
              ],
              // 新增:无asciifolding的分析器,用于严格匹配
              'standard_strict' => [
                'type' => 'custom',
                'tokenizer' => 'standard',
                'filter' => [
                  'lowercase' // 若需要严格区分大小写可去掉此过滤器
                ],
              ],
            ],
            'filter' => [
              'preserve_asciifolding' => [
                'type' => 'asciifolding',
                'preserve_original' => TRUE,
              ],
            ],
          ],
        ],
        'mappings' => [
          '_source' => [
            'enabled' => TRUE,
          ],
          'properties' => $properties,
        ],
      ],
    ];
  }

2. 修改字段映射,添加多字段

假设目标字段是first_name,在properties中配置多字段:

// 示例$properties中的字段配置
'first_name' => [
    'type' => 'text',
    'analyzer' => 'special_chars', // 默认使用非严格分析器
    'fields' => [
        'strict' => [
            'type' => 'text',
            'analyzer' => 'standard_strict' // 严格匹配字段使用无asciifolding的分析器
        ]
    ]
]

3. 调整查询逻辑

  • 非严格查询:直接查询原字段(如first_name),保持现有逻辑即可。
  • 严格查询:查询新增的.strict字段,比如使用query_string时指定字段:
{
    "query": {
        "bool": {
            "must": [
                {
                    "query_string": {
                        "query": "\"Stephane\"",
                        "default_operator": "AND",
                        "fields": ["first_name.strict", "last_name.strict"] // 指定严格匹配字段
                    }
                },
                // ... 其他嵌套查询保持不变
            ]
        }
    }
}

关键说明

  • 无法直接让asciifolding过滤器仅作用于非严格查询,因为分析逻辑是在索引阶段就确定的,查询时必须匹配索引的分词结果。
  • 多字段映射相当于为同一内容生成两种不同的分词版本,分别服务不同的查询需求,是Elasticsearch处理这类场景的标准方案。

内容的提问来源于stack exchange,提问作者Adam Bodera

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.29 23:55:54