Azure Search分析器配置建议:兼容多格式模糊查询需求
Let's break down how to adjust your analyzer so queries like lowerMachineTag:/.*it.*/ return documents with values like "it's me".
Current Analyzer Limitation
Your existing analyzer uses the keyword_v2 tokenizer, which indexes the entire lowerMachineTag value as a single token (e.g., "it's me" is stored as one complete token). While a regex like /.*it\'s.*/ matches this full token, if you're not seeing results with /.*it.*/, it might be due to how Lucene regex interacts with single-token fields (though theoretically it should match). To make the field more flexible for both partial regex matches and full-text searches, we can adjust the analyzer to split tokens at apostrophes and whitespace.
Modified Analyzer Configuration
We'll swap the keyword_v2 tokenizer for standard_v2, which splits text on punctuation (like apostrophes) and whitespace. Combined with the lowercase filter, this will split "it's me" into individual tokens: "it", "s", "me". This way, any query targeting "it"—whether regex or full-text—will find the document.
"analyzers": [ { "@odata.type": "#Microsoft.Azure.Search.CustomAnalyzer", "name": "SWMLuceneAlongWithCustomHyphenAnalyser", "tokenizer": "standard_v2", "tokenFilters": [ "lowercase" ], "charFilters": [] } ]
If you need to preserve hyphenated words as single tokens (while still splitting apostrophes), the standard_v2 tokenizer already handles this by default—no extra filters needed.
Testing the Modified Analyzer
After updating the analyzer and reindexing your data, run your target query:
search=lowerMachineTag:/.*it.*/&$filter=(systemID%20ne%20null)%20and%20(ownerSalesforceRecordID%20eq%20'a0h5B000000gJKfQAM')&$count=true&$top=100&$skip=0
This should now return the 4 documents you expect, since the token "it" is explicitly present in the indexed values for lowerMachineTag.
Alternative: Keep Keyword Tokenizer but Verify Regex
If you must retain the keyword_v2 tokenizer (to keep the entire string as a single token), double-check that your regex is properly formatted. The regex /.*it.*/ should match "it's me" because the dot matches any character (including apostrophes). If it's still not working, confirm:
- The query is correctly URL-encoded (no special characters are breaking the regex)
- The
lowerMachineTagfield is assigned to your custom analyzer - There are no typos in the query or field name
That said, switching to the standard_v2 tokenizer makes the field far more versatile for different search scenarios, which is the better long-term solution.
内容的提问来源于stack exchange,提问作者Santosh

