Azure Search的PathHierarchyTokenizerV2疑似损坏,功能失效求助排查
Let’s walk through the most common reasons your PathHierarchyTokenizerV2 analyzer might have stopped splitting paths like l1/l2/l3 into l1, l1/l2, l1/l2/l3, and how to fix them:
Double-check your analyzer definition
It’s easy to accidentally tweak parameters that break path splitting. Verify your tokenizer and analyzer setup matches what worked before. For example, your custom analyzer should look something like this:{ "analyzers": [ { "name": "path_analyzer", "tokenizer": "path_tokenizer_v2" } ], "tokenizers": [ { "name": "path_tokenizer_v2", "@odata.type": "#Microsoft.Azure.Search.PathHierarchyTokenizerV2", "delimiter": "/", "reverse": false, "skip": 0 } ] }Pay close attention to:
delimiter: Ensure it’s still set to/(changing this to another character will break splitting).reverse: If switched totrue, you’ll get reversed tokens likel3,l2/l3,l1/l2/l3instead.skip: A value greater than 0 will skip leading path segments, omittingl1from results.
Verify your field mapping
Even if your analyzer is correct, if your field isn’t using it, you won’t get the expected tokens. Check that your field explicitly references the custom analyzer:{ "fields": [ { "name": "your_path_field", "type": "Edm.String", "analyzer": "path_analyzer", "searchable": true } ] }Avoid mixing up
analyzer(used for indexing) withsearchAnalyzer(used for queries)—if you only setsearchAnalyzer, the indexing process will fall back to the default analyzer.Re-test with the Analyzer Test API (carefully)
Make sure you’re using the API correctly when testing:- Send a POST request to
https://<your-search-service>.search.windows.net/indexes/<your-index>/analyze?api-version=2023-11-01 - Use this request body:
{ "text": "l1/l2/l3", "analyzer": "path_analyzer" }
Double-check that you’re selecting the right index and analyzer name—typos here are a common culprit. If the API returns incorrect tokens, the issue lies with the analyzer definition. If it returns correct tokens but your indexed data misbehaves, the problem is in your field mapping or indexing pipeline.
- Send a POST request to
Isolate the issue with a test index
Since you mentioned you can reproduce the issue by creating a specific field, build a minimal test index with just the custom analyzer and a single path field. This eliminates interference from other index settings, fields, or analyzers. Use this sample index definition to start:{ "name": "test-path-index", "fields": [ { "name": "id", "type": "Edm.String", "key": true }, { "name": "test_path", "type": "Edm.String", "analyzer": "path_analyzer", "searchable": true } ], "analyzers": [ { "name": "path_analyzer", "tokenizer": "path_tokenizer_v2" } ], "tokenizers": [ { "name": "path_tokenizer_v2", "@odata.type": "#Microsoft.Azure.Search.PathHierarchyTokenizerV2", "delimiter": "/", "reverse": false, "skip": 0 } ] }Test this index with the Analyzer API—if the problem persists here, it’s likely either a parameter mistake or a temporary service-side issue. If it works, compare this index to your original one to spot differing settings.
Check for recent index changes
Think back to when the issue started—did you update the index, analyzer, or field definitions around that time? Even a small change (like adding a new analyzer that overwrote the old one) could break things. Review your index update history if you have it.
内容的提问来源于stack exchange,提问作者Yoad Snapir

