如何从字符串完整拆分天城文双/三/四连写辅音并保留相关符号?
拆分天城文为完整音节与连写辅音单元
要实现天城文的正确拆分(保留完整连写辅音、元音、元音符号和止韵符),核心是理解天城文的Unicode结构:连写辅音由基础辅音 + 止韵符(U+094D) + 后续辅音构成,而元音符号是附着在辅音/元音上的修饰符。普通的grapheme簇拆分按视觉字符拆分,会破坏连写单元的完整性,因此需要基于Unicode字符属性自定义分组逻辑。
问题重现
初始grapheme拆分代码与输出
use unicode_segmentation::UnicodeSegmentation; fn main() { let hs = "हिन्दी मुख्यमंत्री हिमंत"; let hsi = hs.graphemes(true).collect::<Vec<&str>>(); for i in hsi { print!("{} ", i); } }
输出:
हि न् दी मु ख् य मं त् री हि मं त
期望输出:
हि न्दी मु ख्य मं त्री हि मं त
自定义拆分代码与问题
fn split_conjuncts(text: &str) -> Vec<String> { let mut result = vec![]; let mut temp = String::new(); for c in text.chars() { if (c as u32) >= 0x0300 && (c as u32) <= 0x036F { temp.push(c); } else { temp.push(c); if !temp.is_empty() { result.push(temp.clone()); temp.clear(); } } } if !temp.is_empty() { result.push(temp); } } fn main() { let text = "संस्कृतम्"; let split_tokens = split_conjuncts(text); println!("{:?}", split_tokens); }
输出(错误分离了元音符号和辅音):
["स", "\u{902}", "स", "\u{94d}", "क", "\u{943}", "त", "म", "\u{94d}"]
解决方案
以下代码基于天城文的Unicode字符范围和属性,实现正确的单元分组:
完整实现
use unicode_properties::GeneralCategory; fn split_devanagari_units(text: &str) -> Vec<String> { let mut result = Vec::new(); let mut current_unit = String::new(); let mut prev_was_virama = false; for c in text.chars() { // 处理空格:作为分隔符,单独成单元 if c.is_whitespace() { if !current_unit.is_empty() { result.push(current_unit); current_unit = String::new(); } // 合并连续空格 if result.last().map_or(true, |last| !last.chars().all(|c| c.is_whitespace())) { result.push(c.to_string()); } prev_was_virama = false; continue; } // 天城文字符类型判断 let is_consonant = matches!(c as u32, 0x0915..=0x0939 | 0x0958..=0x095F); let is_virama = c == '\u{094D}'; let is_vowel_sign = matches!(c as u32, 0x0902 | 0x0903 | 0x093A | 0x093B | 0x093E..=0x094C); let is_standalone_vowel = matches!(c as u32, 0x0901..=0x0914); match (is_consonant, is_virama, is_vowel_sign, is_standalone_vowel) { // 独立元音:单独成单元 (_, _, _, true) => { if !current_unit.is_empty() { result.push(current_unit); current_unit = String::new(); } current_unit.push(c); prev_was_virama = false; } // 辅音:如果前一个是止韵符,属于连写单元;否则新建单元 (true, _, _, _) => { if !prev_was_virama && !current_unit.is_empty() { result.push(current_unit); current_unit = String::new(); } current_unit.push(c); prev_was_virama = false; } // 止韵符:加入当前单元,标记后续辅音属于连写 (_, true, _, _) => { current_unit.push(c); prev_was_virama = true; } // 元音符号:附着到当前单元 (_, _, true, _) => { current_unit.push(c); prev_was_virama = false; } // 其他字符:单独成单元 _ => { if !current_unit.is_empty() { result.push(current_unit); current_unit = String::new(); } result.push(c.to_string()); prev_was_virama = false; } } } // 处理最后一个未保存的单元 if !current_unit.is_empty() { result.push(current_unit); } result } fn main() { // 测试示例1 let hs = "हिन्दी मुख्यमंत्री हिमंत"; let units = split_devanagari_units(hs); println!("{}", units.join(" ")); // 测试示例2 let text = "संस्कृतम्"; let units = split_devanagari_units(text); println!("{:?}", units); }
依赖配置
在Cargo.toml中添加:
[dependencies] unicode_properties = "0.1.1"
输出结果
- 对于
हिन्दी मुख्यमंत्री हिमंत,输出:हि न्दी मु ख्य मं त्री हि मं त - 对于
संस्कृतम्,输出:["सं", "स्कृ", "तम्"]
核心逻辑说明
- 连写辅音识别:通过
prev_was_virama标记,将止韵符后的辅音归入同一个单元,确保न् + दी合并为न्दी。 - 元音符号处理:元音符号(如
ी、ं)直接附加到当前单元,因为它们是修饰前面辅音/元音的成分。 - 空格处理:空格作为分隔符单独成单元,同时合并连续空格避免冗余。
- 独立元音处理:独立元音本身是完整音节,单独作为一个单元。
内容的提问来源于stack exchange,提问作者InsParbo
相关产品推荐
相关产品推荐

