You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从字符串完整拆分天城文双/三/四连写辅音并保留相关符号?

拆分天城文为完整音节与连写辅音单元

要实现天城文的正确拆分(保留完整连写辅音、元音、元音符号和止韵符),核心是理解天城文的Unicode结构:连写辅音由基础辅音 + 止韵符(U+094D) + 后续辅音构成,而元音符号是附着在辅音/元音上的修饰符。普通的grapheme簇拆分按视觉字符拆分,会破坏连写单元的完整性,因此需要基于Unicode字符属性自定义分组逻辑。

问题重现

初始grapheme拆分代码与输出

use unicode_segmentation::UnicodeSegmentation;

fn main() {
    let hs = "हिन्दी मुख्यमंत्री हिमंत";
    let hsi = hs.graphemes(true).collect::<Vec<&str>>();
    for i in hsi { 
        print!("{}  ", i);
    }
} 

输出:

हि  न्  दी   मु  ख्  य  मं  त्  री    हि  मं  त

期望输出:

हि न्दी  मु ख्य मं त्री  हि मं त

自定义拆分代码与问题

fn split_conjuncts(text: &str) -> Vec<String> {
    let mut result = vec![];
    let mut temp = String::new();

    for c in text.chars() {
        if (c as u32) >= 0x0300 && (c as u32) <= 0x036F {
            temp.push(c);
        } else {
            temp.push(c);
            if !temp.is_empty() {
                result.push(temp.clone());
                temp.clear();
            }
        }
    }
    if !temp.is_empty() {
        result.push(temp);
    }
}

fn main() {
    let text = "संस्कृतम्";
    let split_tokens = split_conjuncts(text);
    println!("{:?}", split_tokens);
}

输出(错误分离了元音符号和辅音):

["स", "\u{902}", "स", "\u{94d}", "क", "\u{943}", "त", "म", "\u{94d}"]

解决方案

以下代码基于天城文的Unicode字符范围和属性,实现正确的单元分组:

完整实现

use unicode_properties::GeneralCategory;

fn split_devanagari_units(text: &str) -> Vec<String> {
    let mut result = Vec::new();
    let mut current_unit = String::new();
    let mut prev_was_virama = false;

    for c in text.chars() {
        // 处理空格:作为分隔符,单独成单元
        if c.is_whitespace() {
            if !current_unit.is_empty() {
                result.push(current_unit);
                current_unit = String::new();
            }
            // 合并连续空格
            if result.last().map_or(true, |last| !last.chars().all(|c| c.is_whitespace())) {
                result.push(c.to_string());
            }
            prev_was_virama = false;
            continue;
        }

        // 天城文字符类型判断
        let is_consonant = matches!(c as u32, 0x0915..=0x0939 | 0x0958..=0x095F);
        let is_virama = c == '\u{094D}';
        let is_vowel_sign = matches!(c as u32, 
            0x0902 | 0x0903 | 0x093A | 0x093B | 0x093E..=0x094C);
        let is_standalone_vowel = matches!(c as u32, 0x0901..=0x0914);

        match (is_consonant, is_virama, is_vowel_sign, is_standalone_vowel) {
            // 独立元音:单独成单元
            (_, _, _, true) => {
                if !current_unit.is_empty() {
                    result.push(current_unit);
                    current_unit = String::new();
                }
                current_unit.push(c);
                prev_was_virama = false;
            }
            // 辅音:如果前一个是止韵符,属于连写单元;否则新建单元
            (true, _, _, _) => {
                if !prev_was_virama && !current_unit.is_empty() {
                    result.push(current_unit);
                    current_unit = String::new();
                }
                current_unit.push(c);
                prev_was_virama = false;
            }
            // 止韵符:加入当前单元,标记后续辅音属于连写
            (_, true, _, _) => {
                current_unit.push(c);
                prev_was_virama = true;
            }
            // 元音符号:附着到当前单元
            (_, _, true, _) => {
                current_unit.push(c);
                prev_was_virama = false;
            }
            // 其他字符:单独成单元
            _ => {
                if !current_unit.is_empty() {
                    result.push(current_unit);
                    current_unit = String::new();
                }
                result.push(c.to_string());
                prev_was_virama = false;
            }
        }
    }

    // 处理最后一个未保存的单元
    if !current_unit.is_empty() {
        result.push(current_unit);
    }

    result
}

fn main() {
    // 测试示例1
    let hs = "हिन्दी मुख्यमंत्री हिमंत";
    let units = split_devanagari_units(hs);
    println!("{}", units.join(" "));

    // 测试示例2
    let text = "संस्कृतम्";
    let units = split_devanagari_units(text);
    println!("{:?}", units);
}

依赖配置

在Cargo.toml中添加:

[dependencies]
unicode_properties = "0.1.1"

输出结果

  • 对于हिन्दी मुख्यमंत्री हिमंत,输出:
    हि न्दी  मु ख्य मं त्री  हि मं त
    
  • 对于संस्कृतम्,输出:
    ["सं", "स्कृ", "तम्"]
    

核心逻辑说明

  1. 连写辅音识别:通过prev_was_virama标记,将止韵符后的辅音归入同一个单元,确保न् + दी合并为न्दी。
  2. 元音符号处理:元音符号(如ी、ं)直接附加到当前单元,因为它们是修饰前面辅音/元音的成分。
  3. 空格处理:空格作为分隔符单独成单元,同时合并连续空格避免冗余。
  4. 独立元音处理:独立元音本身是完整音节,单独作为一个单元。

内容的提问来源于stack exchange,提问作者InsParbo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 03:35:19