使用NEST创建Pattern Analyzer遇pattern_syntax_exception异常求助
问题排查与解决
异常信息
使用NEST 7.17.4配合Elasticsearch 8.4创建Pattern Analyzer时触发以下异常:
Elasticsearch.Net.ElasticsearchClientException: 'Request failed to execute. Call: Status code 400 from: PUT /temp-index-for-integration-tests?pretty=true&error_trace=true. ServerError: Type: pattern_syntax_exception Reason: "Illegal repetition near index 80
([\p{L}\d]+)|(?<=\D)(?=\d)|(?<=\d)(?=\D)|(?<=[\p{L}&&[\p{Lu}]])(?=\p{Lu})|(?<=\p{Lu})(?=\p{Lu}[\p{L}&&[^\p{Lu}]])
^"'
场景对比
该正则表达式在Elasticsearch Dev Console中可正常创建分析器:
PUT test-index-3 { "settings": { "analysis": { "analyzer": { "camel": { "type": "pattern", "pattern": "([^\\p{L}\\d]+)|(?<=\\D)(?=\\d)|(?<=\\d)(?=\\D)|(?<=[\\p{L}&&[^\\p{Lu}]])(?=\\p{Lu})|(?<=\\p{Lu})(?=\\p{Lu}[\\p{L}&&[^\\p{Lu}]])" } } } } }
但调用NEST的Indices.CreateAsync方法时抛出上述异常,原.NET代码如下:
public async Task IndexCreateAsync(IEnumerable<IFieldDefinition> fieldDefinitions, string indexName) { Dictionary<PropertyName, IProperty> indexFields = fieldDefinitions .Select(f => _elasticsearchMapper.Map(f)) .ToDictionary(p => p.Name, p => p); PutMappingRequest mappings = new PutMappingRequest(indexName) { Properties = new Properties(indexFields) }; var patternAnalyzer = new PatternAnalyzer { Pattern = @"([^\\p{L}\\d]+)|(?<=\\D)(?=\\d)|(?<=\\d)(?=\\D)|(?<=[\\p{L}&&[^\\p{Lu}]])(?=\\p{Lu})|(?<=\\p{Lu})(?=\\p{Lu}[\\p{L}&&[^\\p{Lu}]])", Lowercase = true }; IndexState indexSettings = new IndexState { Mappings = mappings, Settings = new IndexSettings { Analysis = new Analysis { Analyzers = new Analyzers { { ElasticsearchConstants.TextAnalysis.CustomAnalyzers.CustomPatternRegexCasesAndSpecialChars, patternAnalyzer } } } } }; await _elasticClient.Indices.CreateAsync(indexName, s => s.InitializeUsing(indexSettings)); }
问题原因
核心问题出在字符串转义逻辑:
- Elasticsearch Dev Console的JSON中,必须用双反斜杠
\\表示正则里的单个反斜杠,因为JSON本身会转义反斜杠。 - .NET的逐字字符串(以
@开头)中,反斜杠无需额外转义,直接写单个\即可。原代码错误保留了JSON格式的双反斜杠,导致传到Elasticsearch的正则出现多余转义字符,触发语法错误。
修复方案
修改PatternAnalyzer的Pattern属性,去掉多余的反斜杠:
var patternAnalyzer = new PatternAnalyzer { Pattern = @"([^\p{L}\d]+)|(?<=\D)(?=\d)|(?<=\d)(?=\D)|(?<=[\p{L}&&[^\p{Lu}]])(?=\p{Lu})|(?<=\p{Lu})(?=\p{Lu}[\p{L}&&[^\p{Lu}]])", Lowercase = true };
修改后NEST会自动处理JSON序列化时的转义,将正确的正则表达式发送给Elasticsearch,即可成功创建分析器。
内容的提问来源于stack exchange,提问作者cah1r
相关产品推荐
相关产品推荐

