如何按自定义规则排序字符串?ICU排序不符合需求的解决方法
a/á Problem Context
I'm working on string sorting using ICU, with the following target strings: a, Aaa, áa, Aa, A, á.
My desired sorted order is:
a、A、Aa、Aaa、á、áa
But the ICU Demo returns this unexpected order:
a、A、á、Aa、áa、Aaa
I need to adjust the collation rules to group all non-accented a/A strings first (sorted normally), followed by accented á strings (also sorted in their natural order).
Why ICU's Default Behavior Does This
By default, ICU uses the Unicode Collation Algorithm (UCA), which treats accents as secondary sorting weights. Since a and á share the same base character, ICU evaluates the accent difference before considering string length or case nuances. That's why á gets inserted right after A—it prioritizes resolving the accent difference over extending the string length.
Solution: Custom Collation Rules
To override this and get your desired order, you need to redefine the primary weight of á so it's treated as a separate base character that comes after all non-accented a variants. Here's how to implement this:
Step 1: Define the Custom Rule
Create a collation rule string that explicitly tells ICU to push á to the end of the a group:
&a << á
The << operator increases the primary weight of the following character, ensuring all strings with the base a (regardless of case or length) are sorted first, before any strings with á.
Step 2: Implement the Rule in Your Code
Depending on the ICU API you're using, here's how to apply this rule:
Java Example
import java.text.RuleBasedCollator; import java.util.Arrays; import java.util.Collections; import java.util.List; public class CustomSort { public static void main(String[] args) throws Exception { String customRules = "&a << á"; RuleBasedCollator collator = new RuleBasedCollator(customRules); collator.setStrength(RuleBasedCollator.TERTIARY); // Preserve case and length sensitivity List<String> strings = Arrays.asList("a", "Aaa", "áa", "Aa", "A", "á"); Collections.sort(strings, collator); System.out.println(strings); // Output: [a, A, Aa, Aaa, á, áa] } }
JavaScript (Node.js with ICU Support)
const { RuleBasedCollator } = require('icu4c'); // Use appropriate ICU bindings const customRules = "&a << á"; const collator = new RuleBasedCollator(customRules); const strings = ["a", "Aaa", "áa", "Aa", "A", "á"]; strings.sort(collator.compare); console.log(strings); // Output: ["a", "A", "Aa", "Aaa", "á", "áa"]
Extending to Other Characters
If you need this behavior for more accented character pairs (like e/é, i/í), simply extend the rule string:
&a << á &e << é &i << í
Verification
After applying this rule, ICU will first sort all non-accented a/A strings using normal case and length logic, then move on to sorting the á/áa group. This exactly matches your desired output.
内容的提问来源于stack exchange,提问作者김진필

