You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

寻求可处理Unicode代理对的C# string.IndexOf安全替代实现

处理Unicode代理对的C#字符串查找实现

嘿,我懂你想要的是什么——默认的string.IndexOf是基于UTF-16码元工作的,这会把像Emoji这类由代理对组成的Unicode字符拆成两个单独的码元,导致查找子字符串时出错。你当前的代码已经迈出了第一步:用StringInfo.GetTextElementEnumerator把字符串拆成完整的文本元素(也就是真正的Unicode字符,包括代理对、组合字符序列),但只能处理单个元素的查找。下面是扩展后的实现,能像string.IndexOf一样查找子字符串:

完整实现代码

using System;
using System.Collections.Generic;
using System.Globalization;
using System.Linq;

public static class UnicodeStringExtensions
{
    // 等效于string.IndexOf,但正确处理Unicode文本元素(代理对、组合字符等)
    public static int UnicodeIndexOf(this string input, string find)
    {
        // 处理边界情况
        if (string.IsNullOrEmpty(input)) return -1;
        if (string.IsNullOrEmpty(find)) return 0;

        // 将输入和查找字符串都转换为文本元素列表
        var inputElements = input.ToTextElements().ToList();
        var findElements = find.ToTextElements().ToList();

        // 如果查找的元素数量比输入多,直接返回-1
        if (findElements.Count > inputElements.Count) return -1;

        // 滑动窗口遍历输入元素列表,查找匹配的子序列
        for (int i = 0; i <= inputElements.Count - findElements.Count; i++)
        {
            bool match = true;
            for (int j = 0; j < findElements.Count; j++)
            {
                if (inputElements[i + j] != findElements[j])
                {
                    match = false;
                    break;
                }
            }

            if (match)
            {
                // 将文本元素索引转换回原字符串的UTF-16码元索引
                return StringInfo.ParseCombiningCharacters(input)[i];
            }
        }

        // 未找到匹配
        return -1;
    }

    // 辅助方法:将字符串转换为Unicode文本元素序列
    public static IEnumerable<string> ToTextElements(this string input)
    {
        if (string.IsNullOrEmpty(input)) yield break;

        var enumerator = StringInfo.GetTextElementEnumerator(input);
        while (enumerator.MoveNext())
        {
            yield return enumerator.GetTextElement();
        }
    }
}

关键细节说明

  • 文本元素转换:ToTextElements方法用StringInfo.GetTextElementEnumerator遍历字符串,把每个完整的Unicode字符(不管是单个码元还是代理对/组合序列)作为一个元素返回。
  • 滑动窗口匹配:在文本元素列表上用嵌套循环做滑动窗口匹配,确保子序列的每个元素都完全匹配。
  • 索引转换:找到匹配的文本元素起始索引后,用StringInfo.ParseCombiningCharacters把这个元素索引转换回原字符串的UTF-16码元索引,和string.IndexOf返回的索引格式保持一致。

测试示例

var testString = "👨‍👩‍👧‍👦 Hello 👋 World";
var findString = "👨‍👩‍👧‍👦 Hello";

// 调用扩展方法
int index = testString.UnicodeIndexOf(findString);
Console.WriteLine(index); // 输出0,正确匹配开头的Emoji短语

// 对比默认IndexOf的问题
int defaultIndex = testString.IndexOf("👨‍👩‍👧‍👦");
// 注意:默认IndexOf返回的是第一个代理对的码元索引,但如果查找的是多元素子字符串,可能会出错

内容的提问来源于stack exchange,提问作者Ibrennan208

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 11:11:53