Elasticsearch 5.3.2如何实现大小写不敏感搜索及聚合优化?
Hey there! Let's tackle your two questions about Elasticsearch 5.3.2—since this is an older version, we'll focus on solutions that are fully compatible with it.
1. How to Perform Case-Insensitive Search
The approach depends on whether you're working with text or keyword fields:
For Text Fields
By default, Elasticsearch's standard analyzer (applied to text fields out of the box) converts all text to lowercase during indexing. That means a simple match query will automatically match case variants. For example:
GET /your_index/_search { "query": { "match": { "your_text_field": "Windows" } } }
This will return documents with "windows", "WINDOWS", "Windows", etc., in your_text_field.
For Keyword Fields
keyword fields store exact values by default, so case-sensitive matches are the norm. To make them case-insensitive, you need to define a lowercase normalizer in your index settings and apply it to the field:
First, create the index with the normalizer and mapping:
PUT /your_index { "settings": { "analysis": { "normalizer": { "lowercase_norm": { "type": "custom", "filter": ["lowercase"] } } } }, "mappings": { "your_doc_type": { "properties": { "your_keyword_field": { "type": "keyword", "normalizer": "lowercase_norm" } } } } }
If you already have an existing index, you can't modify the normalizer of an existing keyword field directly. Instead, you'll need to:
- Create a new index with the updated mapping
- Reindex your data into the new index using the
_reindexAPI
Once set up, queries like term or match on this field will ignore case.
2. Fixing Case-Sensitive Aggregations for 'Windows' Results
Aggregations use the exact terms stored in the index, so case variants (like "Windows" vs "windows") will show up as separate buckets. Here are two reliable fixes:
Option 1: Use a Normalized Keyword Sub-Field (Recommended)
Add a sub-field to your target field that uses the lowercase normalizer we defined earlier. This ensures all values are indexed in lowercase, so aggregations group them correctly:
Update your mapping (or create a new index with this mapping):
PUT /your_index { "settings": { "analysis": { "normalizer": { "lowercase_norm": { "type": "custom", "filter": ["lowercase"] } } } }, "mappings": { "your_doc_type": { "properties": { "os_name": { "type": "text", "fields": { "lowercase_keyword": { "type": "keyword", "normalizer": "lowercase_norm" } } } } } } }
Then run your aggregation against this sub-field:
GET /your_index/_search { "size": 0, // Skip returning individual docs "aggs": { "windows_agg": { "terms": { "field": "os_name.lowercase_keyword" } } } }
This will group all case variants of "windows" into a single bucket (displayed as "windows").
Option 2: Use a Script for On-the-Fly Conversion (Quick Fix)
If you can't modify your mapping right now, use a Painless script to convert values to lowercase during aggregation. Note: This is less performant than the normalized field approach, especially with large datasets.
GET /your_index/_search { "size": 0, "aggs": { "windows_agg": { "terms": { "script": { "lang": "painless", "source": "doc['os_name.keyword'].value.toLowerCase()" } } } } }
Verifying Your Mapping
To confirm your normalizer or sub-field is set up correctly, use the mapping API:
GET /your_index/_mapping
Check the response for the field's normalizer or fields configuration to ensure everything matches your intended setup.
内容的提问来源于stack exchange,提问作者Utsav Dusad

