You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何拆解韩文单词为字母组件?C#谚文兼容字符转换方法

问题解决:将Hangul Jamo转换为Hangul Compatibility Jamo

你当前用NormalizationForm.FormD得到的是Hangul Jamo(Unicode U+1100-U+11FF范围),这是用于构建完整谚文字符的底层组件;而你期望的是Hangul Compatibility Jamo(Unicode U+3130-U+318F范围),这类字符是独立显示的谚文单元。

要实现转换,可根据两类字符的Unicode编码偏移量做映射:

  • 初声辅音(U+1100-U+1112):对应兼容块U+3131-U+3143,偏移量为0x3131 - 0x1100 = 0x2031
  • 中声元音(U+1161-U+1175):对应兼容块U+314F-U+3163,偏移量为0x314F - 0x1161 = 0x1FE8
  • 终声辅音(U+11A8-U+11C2):对应兼容块U+3131-U+314A,偏移量为0x3131 - 0x11A8 = 0x1F89

下面是修改后的C#代码,实现自动转换:

var text = "루돌프사슴코";
var normalized = text.Normalize(NormalizationForm.FormD);

foreach (var c in normalized)
{
    char converted = c;
    // 处理初声辅音
    if (c >= '\u1100' && c <= '\u1112')
    {
        converted = (char)(c + 0x2031);
    }
    // 处理中声元音
    else if (c >= '\u1161' && c <= '\u1175')
    {
        converted = (char)(c + 0x1FE8);
    }
    // 处理终声辅音
    else if (c >= '\u11A8' && c <= '\u11C2')
    {
        converted = (char)(c + 0x1F89);
    }
    Console.Write(converted + " ");
}

执行这段代码后,输出会和你期望的一致:

ㄹ ㅜ ㄷ ㅜ ㄹ ㅍ ㅡ ㅅ ㅏ ㅅ ㅡ ㅁ ㅋ ㅗ 

补充说明

  • 谚文拆解得到的初/中/终声组件都在可映射范围内,这个方法完全适用。
  • 若需要批量处理,可封装成扩展方法复用:
public static char ToHangulCompatibilityJamo(this char jamo)
{
    if (jamo >= '\u1100' && jamo <= '\u1112')
        return (char)(jamo + 0x2031);
    if (jamo >= '\u1161' && jamo <= '\u1175')
        return (char)(jamo + 0x1FE8);
    if (jamo >= '\u11A8' && jamo <= '\u11C2')
        return (char)(jamo + 0x1F89);
    return jamo;
}

使用时直接调用:

Console.Write(c.ToHangulCompatibilityJamo() + " ");

内容的提问来源于stack exchange,提问作者user3434046

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 01:04:52