如何在Solr中搜索含数字与特殊字符的词汇?splitOnNumerics设置后仍有异常
Great question—this is a common gotcha with tokenization and special characters in search engines. Let’s break down why this is happening and how to fix it.
Root Cause
Setting splitOnNumerics: 0 tells the tokenizer not to split numbers from adjacent letters (hence why "90s" and "omega30" work), but it doesn’t change how non-word characters like " or ) are handled. Most standard tokenizers (like Elasticsearch’s standard tokenizer) split tokens at these non-word characters by default.
For example, when you index "80\"", the tokenizer will split it into just "80" (discarding the quote), so even if you escape the quote in your search query, there’s no matching token to find.
Solutions
Here are three actionable fixes depending on your use case:
1. Create a Custom Tokenizer to Include Special Characters
If you want terms like "80\"" or "40)" to be treated as single searchable tokens, define a custom analyzer that includes your desired special characters in the allowed token set.
For example, in Elasticsearch, you’d set up a custom pattern tokenizer that splits on anything not in your allowed character set (letters, numbers, ", and )):
{ "settings": { "analysis": { "analyzer": { "mixed_special_analyzer": { "tokenizer": "mixed_special_tokenizer" } }, "tokenizer": { "mixed_special_tokenizer": { "type": "pattern", "pattern": "[^a-zA-Z0-9\"()]+", // Split on non-allowed characters "group": 0 } } } } }
Apply this analyzer to your target field, and when you index "80\"" or "40)", they’ll be stored as full tokens that your escaped queries can match.
2. Use a Keyword Sub-Field for Exact Matches
If you don’t want to overhaul your main analyzer, add a keyword sub-field to your mapping. This field stores the exact, un-tokenized string, perfect for precise searches:
{ "mappings": { "properties": { "your_search_field": { "type": "text", "analyzer": "standard", "split_on_numerics": 0, "fields": { "keyword": { "type": "keyword", "ignore_above": 256 } } } } } }
Then query the keyword field directly with your escaped term:
{ "query": { "term": { "your_search_field.keyword": "80\"" } } }
Just make sure to escape characters correctly (e.g., in JSON, you’ll need to write "80\\\"" to pass the escaped quote).
3. Verify Tokenization and Adjust Queries
If you can’t modify your index, first check how your terms are being tokenized using your search engine’s analyze API. For Elasticsearch:
GET _analyze { "analyzer": "standard", "split_on_numerics": 0, "text": "80\"" }
If the result is just the token "80", you can use a wildcard query to match partial terms (note: this is less efficient than token-based searches):
{ "query": { "wildcard": { "your_search_field": "80*" } } }
Or combine a term query for "80" with a filter for the presence of the quote character if your engine supports it.
Final Notes
Always double-check that your escaped characters are being passed correctly to the search engine—different contexts (like JSON strings) may require double-escaping (e.g., \\ instead of \).
内容的提问来源于stack exchange,提问作者metylbk

