You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何按自定义规则排序字符串?ICU排序不符合需求的解决方法

How to Adjust ICU Collation to Group Non-Accented Strings Before Accented Ones for a/á

Problem Context

I'm working on string sorting using ICU, with the following target strings: a, Aaa, áa, Aa, A, á.

My desired sorted order is:

a、A、Aa、Aaa、á、áa

But the ICU Demo returns this unexpected order:

a、A、á、Aa、áa、Aaa

I need to adjust the collation rules to group all non-accented a/A strings first (sorted normally), followed by accented á strings (also sorted in their natural order).


Why ICU's Default Behavior Does This

By default, ICU uses the Unicode Collation Algorithm (UCA), which treats accents as secondary sorting weights. Since a and á share the same base character, ICU evaluates the accent difference before considering string length or case nuances. That's why á gets inserted right after A—it prioritizes resolving the accent difference over extending the string length.


Solution: Custom Collation Rules

To override this and get your desired order, you need to redefine the primary weight of á so it's treated as a separate base character that comes after all non-accented a variants. Here's how to implement this:

Step 1: Define the Custom Rule

Create a collation rule string that explicitly tells ICU to push á to the end of the a group:

&a << á

The << operator increases the primary weight of the following character, ensuring all strings with the base a (regardless of case or length) are sorted first, before any strings with á.

Step 2: Implement the Rule in Your Code

Depending on the ICU API you're using, here's how to apply this rule:

Java Example
import java.text.RuleBasedCollator;
import java.util.Arrays;
import java.util.Collections;
import java.util.List;

public class CustomSort {
    public static void main(String[] args) throws Exception {
        String customRules = "&a << á";
        RuleBasedCollator collator = new RuleBasedCollator(customRules);
        collator.setStrength(RuleBasedCollator.TERTIARY); // Preserve case and length sensitivity

        List<String> strings = Arrays.asList("a", "Aaa", "áa", "Aa", "A", "á");
        Collections.sort(strings, collator);

        System.out.println(strings); // Output: [a, A, Aa, Aaa, á, áa]
    }
}
JavaScript (Node.js with ICU Support)
const { RuleBasedCollator } = require('icu4c'); // Use appropriate ICU bindings
const customRules = "&a << á";
const collator = new RuleBasedCollator(customRules);

const strings = ["a", "Aaa", "áa", "Aa", "A", "á"];
strings.sort(collator.compare);

console.log(strings); // Output: ["a", "A", "Aa", "Aaa", "á", "áa"]

Extending to Other Characters

If you need this behavior for more accented character pairs (like e/é, i/í), simply extend the rule string:

&a << á &e << é &i << í

Verification

After applying this rule, ICU will first sort all non-accented a/A strings using normal case and length logic, then move on to sorting the á/áa group. This exactly matches your desired output.

内容的提问来源于stack exchange,提问作者김진필

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 11:00:56