寻求可处理Unicode代理对的C# string.IndexOf安全替代实现
处理Unicode代理对的C#字符串查找实现
嘿,我懂你想要的是什么——默认的string.IndexOf是基于UTF-16码元工作的,这会把像Emoji这类由代理对组成的Unicode字符拆成两个单独的码元,导致查找子字符串时出错。你当前的代码已经迈出了第一步:用StringInfo.GetTextElementEnumerator把字符串拆成完整的文本元素(也就是真正的Unicode字符,包括代理对、组合字符序列),但只能处理单个元素的查找。下面是扩展后的实现,能像string.IndexOf一样查找子字符串:
完整实现代码
using System; using System.Collections.Generic; using System.Globalization; using System.Linq; public static class UnicodeStringExtensions { // 等效于string.IndexOf,但正确处理Unicode文本元素(代理对、组合字符等) public static int UnicodeIndexOf(this string input, string find) { // 处理边界情况 if (string.IsNullOrEmpty(input)) return -1; if (string.IsNullOrEmpty(find)) return 0; // 将输入和查找字符串都转换为文本元素列表 var inputElements = input.ToTextElements().ToList(); var findElements = find.ToTextElements().ToList(); // 如果查找的元素数量比输入多,直接返回-1 if (findElements.Count > inputElements.Count) return -1; // 滑动窗口遍历输入元素列表,查找匹配的子序列 for (int i = 0; i <= inputElements.Count - findElements.Count; i++) { bool match = true; for (int j = 0; j < findElements.Count; j++) { if (inputElements[i + j] != findElements[j]) { match = false; break; } } if (match) { // 将文本元素索引转换回原字符串的UTF-16码元索引 return StringInfo.ParseCombiningCharacters(input)[i]; } } // 未找到匹配 return -1; } // 辅助方法:将字符串转换为Unicode文本元素序列 public static IEnumerable<string> ToTextElements(this string input) { if (string.IsNullOrEmpty(input)) yield break; var enumerator = StringInfo.GetTextElementEnumerator(input); while (enumerator.MoveNext()) { yield return enumerator.GetTextElement(); } } }
关键细节说明
- 文本元素转换:
ToTextElements方法用StringInfo.GetTextElementEnumerator遍历字符串,把每个完整的Unicode字符(不管是单个码元还是代理对/组合序列)作为一个元素返回。 - 滑动窗口匹配:在文本元素列表上用嵌套循环做滑动窗口匹配,确保子序列的每个元素都完全匹配。
- 索引转换:找到匹配的文本元素起始索引后,用
StringInfo.ParseCombiningCharacters把这个元素索引转换回原字符串的UTF-16码元索引,和string.IndexOf返回的索引格式保持一致。
测试示例
var testString = "👨👩👧👦 Hello 👋 World"; var findString = "👨👩👧👦 Hello"; // 调用扩展方法 int index = testString.UnicodeIndexOf(findString); Console.WriteLine(index); // 输出0,正确匹配开头的Emoji短语 // 对比默认IndexOf的问题 int defaultIndex = testString.IndexOf("👨👩👧👦"); // 注意:默认IndexOf返回的是第一个代理对的码元索引,但如果查找的是多元素子字符串,可能会出错
内容的提问来源于stack exchange,提问作者Ibrennan208
相关产品推荐
相关产品推荐

