Oak Lucene索引无法按西班牙语字母序排序问题求助
问题
我有一批带有name属性的dam:Asset类型资产,需要按该属性的字母序升序查询并获取指定数量的结果。当前使用的查询语句如下:
type=dam:Asset path=/content/dam/en/foobar/contacts/ orderby=@jcr:content/data/master/@name orderby.sort=asc p.limit=3
该查询在普通英文排序中正常工作,但在西班牙语场景下,Á应与A视为同一字母。当前查询中,当资产名称包含Álvaro时,排序结果未将其与Abel归为同一序列,导致Álvaro被排除在前三结果之外,预期结果应为Abel, Álvaro, Eduardo。
为解决此问题,我创建了自定义Oak Lucene索引,配置了同义词映射(将á映射为a,Á映射为A等),也尝试过字符映射过滤器,且已确认查询使用该自定义索引,但重新索引后问题仍未解决,该如何修复?
自定义索引配置如下:
<?xml version="1.0" encoding="UTF-8"?> <jcr:root xmlns:oak="http://jackrabbit.apache.org/oak/ns/1.0" xmlns:jcr="http://www.jcp.org/jcr/1.0" xmlns:nt="http://www.jcp.org/jcr/nt/1.0" xmlns:rep="internal" jcr:mixinTypes="[rep:AccessControllable]" jcr:primaryType="nt:unstructured"> <socialLucene/> <workflowDataLucene/> <slingeventJob/> <jcrLanguage/> <versionStoreIndex/> <repMembers/> <cqReportsLucene/> <commerceLucene/> <counter/> <authorizables/> <enablementResourceName/> <externalPrincipalNames/> <cmLucene/> <foobarCFIndexFilter jcr:primaryType="oak:QueryIndexDefinition" async="[async,nrt]" evaluatePathRestrictions="{Boolean}true" includedPaths="[/content/dam/es/foobar,/content/dam/en/foobar]" queryPaths="[/content/dam/es/foobar,/content/dam/en/foobar]" reindex="{Boolean}false" reindexCount="{Long}24" seed="{Long}3850652403740003290" type="lucene"> <analyzers jcr:primaryType="nt:unstructured"> <default jcr:primaryType="nt:unstructured"> <filters jcr:primaryType="nt:unstructured"> <Synonym jcr:primaryType="nt:unstructured" format="solr" synonyms="synonyms.txt"> <synonyms.txt/> </Synonym> </filters> <tokenizer jcr:primaryType="nt:unstructured" name="Classic"/> </default> </analyzers> <indexRules jcr:primaryType="nt:unstructured"> <nt:base jcr:primaryType="nt:unstructured"> <properties jcr:primaryType="nt:unstructured"> <title jcr:primaryType="nt:unstructured" analyzed="{Boolean}true" isRegexp="{Boolean}false" name="jcr:content/data/master/title" nodeScopeIndex="{Boolean}true" ordered="{Boolean}true" propertyIndex="{Boolean}true" type="String"/> <date jcr:primaryType="nt:unstructured" name="jcr:content/data/master/date" ordered="{Boolean}true" propertyIndex="{Boolean}true"/> <sectors jcr:primaryType="nt:unstructured" name="jcr:content/data/master/sectors" propertyIndex="{Boolean}true"/> <contentFragment jcr:primaryType="nt:unstructured" name="jcr:content/contentFragment" propertyIndex="{Boolean}true"/> <model jcr:primaryType="nt:unstructured" name="cq:model" propertyIndex="{Boolean}true"/> <name jcr:primaryType="nt:unstructured" analyzed="{Boolean}true" isRegexp="{Boolean}false" name="jcr:content/data/master/name" nodeScopeIndex="{Boolean}true" ordered="{Boolean}true" propertyIndex="{Boolean}true" type="String"/> </properties> </nt:base> </indexRules> </foobarCFIndexFilter> <cqProjectLucene/> <ntFolderDamLucene/> <acPrincipalName/> <uuid/> <damAssetLucene/> <rep:policy/> <cqPayloadPath/> <nodetypeLucene/> <nodetype/> <ntBaseLucene/> <reference/> <principalName/> <cqTagLucene/> <lucene/> <repTokenIndex/> <externalId/> <authorizableId/> <cqPageLucene/> </jcr:root>
同义词文件synonyms.txt内容为:
á, a Á, A # 其他重音字符映射省略
解决方案
- 替换同义词过滤器为ICU折叠过滤器:同义词映射适用于语义同义词,不适合重音字符的归一化处理。改用
ICUFoldingFilter可直接将带重音的字符转换为对应无重音字符(如Á→A、á→a),完美匹配西班牙语排序需求。同时将分词器从ClassicTokenizer改为KeywordTokenizer,确保整个name字段作为完整字符串处理,避免分词破坏名称结构。修改后的分析器配置如下:
<analyzers jcr:primaryType="nt:unstructured"> <default jcr:primaryType="nt:unstructured"> <filters jcr:primaryType="nt:unstructured"> <ICUFoldingFilter jcr:primaryType="nt:unstructured"/> </filters> <tokenizer jcr:primaryType="nt:unstructured" name="KeywordTokenizer"/> </default> </analyzers>
强制触发全量重新索引:修改索引配置后,将
reindex="{Boolean}true",触发全量重新索引,确保所有资产的name字段都用新的分析器规则处理。待索引完成后,再将该值改回false。验证索引命中情况:执行查询时,通过查询解释器确认自定义Lucene索引是否被用于排序,避免JCR原生排序逻辑覆盖索引排序结果。
校验排序结果:重新索引完成后,执行原查询,检查
Álvaro是否与Abel归为同一序列。若仍有问题,可临时移除p.limit参数,查看完整排序结果,确认重音字符是否被正确归一化。
内容的提问来源于stack exchange,提问作者Plácid Masvidal
相关产品推荐
相关产品推荐

