Elasticsearch全类型字符串适用分析器选型及精确匹配搜索无结果问题咨询
First off, let's break down why your search for 4929-3813-3266-4295 is returning no results even with the whitespace analyzer enabled.
Why Your Current Query Fails
Chances are, your text field is still using Elasticsearch's default standard analyzer (even if you intended to use whitespace). The standard analyzer splits strings on punctuation like hyphens, so your credit card number gets broken into separate tokens: 4929, 3813, 3266, 4295. When you use match_phrase to search for the full number, there’s no single token that matches the entire string—hence the empty response.
If you did correctly set the whitespace analyzer, it should split text only on spaces (keeping the full credit card number as one token). But even then, match_phrase relies on precise token positioning, so any indexing/querying discrepancies could still cause failures.
How to Enable Exact Matching for Any Content
The most flexible solution is to use a keyword sub-field alongside your text field. This lets you handle both full-text search (on the text field) and exact matches (on the keyword sub-field). Here's how to set it up:
- Update your index mapping (you’ll need to reindex documents if the index already exists):
PUT /full_text { "mappings": { "properties": { "text": { "type": "text", "analyzer": "whitespace", // Keep your preferred analyzer for full-text search "fields": { "keyword": { "type": "keyword", "ignore_above": 256 // Adjust based on your longest expected string } } } } } }
- Modify your query to use a
termsearch on the keyword sub-field for exact matches:
curl -X GET "http://username:password@localhost:9200/full_text/_search?pretty" -H 'Content-Type: application/json' -d' { "_source": { "includes": [ "filename", "filepath", "upload_time", "file_size", "file_access_date", "file_create_date", "file_modified_date", "highlight" ] }, "query": { "bool": { "must": [ { "term": { "text.keyword": "4929-3813-3266-4295" } } ], "filter": [ { "exists": { "field": "text" } } ] } }, "highlight": { "pre_tags": "<b>", "post_tags": "</b>", "fields": { "text": { "fragment_size": 100, "number_of_fragments": 1, "fragmenter": "span" } } } }'
This will find the exact string you’re searching for, since the keyword field stores the original, unanalyzed text.
Which Analyzer to Use for All String Types?
There’s no universal "one-size-fits-all" analyzer—it depends on your use case:
- For exact matches (IDs, phone numbers, credit cards, emails): Use the
keywordtype (no analysis) or thewhitespaceanalyzer (splits only on spaces, preserves punctuation). - For full-text search (natural language, sentences): Stick with the
standardanalyzer—it handles most languages well by splitting on whitespace/punctuation and lowercasing terms. - For custom needs: Build a custom analyzer (using tokenizers and filters) if you have specific formatting rules (e.g., preserving hyphens in certain strings but splitting others).
The whitespace analyzer is a solid middle ground if you want to split text into words but keep punctuation intact, but for true exact matching, the keyword sub-field approach is unbeatable because it retains the original string exactly as it was indexed.
内容的提问来源于stack exchange,提问作者sheel

