You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

C#中含Emoji的字符串出现次数统计问题求助

解决Emoji场景下的字符串出现次数统计问题

原代码的核心问题在于:C#字符串采用UTF-16编码,许多Emoji(如😀、👨🏾)由多个char组成(代理对或字符组合序列),按单个char遍历会将一个视觉上的完整Emoji拆分成多个独立元素统计,同时Replace操作会破坏Emoji的编码结构,导致统计完全失效。

正确实现方案

使用.NET提供的StringInfo.GetTextElementEnumerator遍历字符串中的视觉字符(Grapheme Cluster)——这是用户视觉感知到的单个字符单元,包含所有类型的Unicode字符(包括单码点字符、Emoji、带修饰的组合Emoji等),以此为单位统计出现次数:

完整代码示例

// 假设CountClass的定义如下
public class CountClass
{
    public string Category { get; set; }
    public int Count { get; set; }
}

private List<CountClass> CountCharacterOccurences(string theText)
{
    var graphemeCounts = new Dictionary<string, int>();
    var enumerator = StringInfo.GetTextElementEnumerator(theText);

    // 遍历所有视觉字符
    while (enumerator.MoveNext())
    {
        string currentGrapheme = enumerator.GetTextElement();
        
        if (graphemeCounts.ContainsKey(currentGrapheme))
        {
            graphemeCounts[currentGrapheme]++;
        }
        else
        {
            graphemeCounts[currentGrapheme] = 1;
        }
    }

    // 转换成需求的List<CountClass>格式
    return graphemeCounts
        .Select(kvp => new CountClass { Category = kvp.Key, Count = kvp.Value })
        .ToList();
}

方案优势

  1. 准确性:完整识别所有视觉字符,不会拆分Emoji或其他Unicode组合序列。
  2. 效率:避免原代码中反复调用Replace产生大量临时字符串的性能损耗。
  3. 通用性:支持所有Unicode标准定义的字符类型,包括特殊符号、多语言字符、Emoji组合序列等。

内容的提问来源于stack exchange,提问作者Beb

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 12:18:12