如何获取含组合字符的文本中的唯一字符列表?
如何正确提取包含组合变音符号文本中的唯一字符
当文本包含s̈、b̃这类由基础字符+变音符号组成的组合字符时,直接用set()拆分字符串会将它们拆分为独立的基础字符和变音符号,导致结果不符合预期。以下是两种解决方法:
方法一:使用第三方regex库(简便推荐)
regex库支持匹配Unicode grapheme簇(视觉上的单个字符),能直接将组合字符作为整体提取:
- 先安装依赖:
pip install regex
- 示例代码:
import unicodedata import regex sentence = "nejon ámas̈hó T̃iqu c̈ab̃op" # 标准化字符为NFC格式,确保组合形式统一 normalized_text = unicodedata.normalize('NFC', sentence) # 提取所有视觉上的单个字符(grapheme簇) grapheme_list = regex.findall(r'\X', normalized_text) # 去重、过滤空格、排序(排序可选,按需调整) unique_chars = sorted([char for char in set(grapheme_list) if char != ' ']) print(unique_chars) # 输出:['a', 'á', 'b̃', 'c̈', 'e', 'h', 'i', 'j', 'm', 'n', 'o', 'ó', 'p', 'q', 's̈', 'T̃', 'u']
方法二:手动遍历处理(无需第三方库)
通过判断Unicode字符类别,将组合变音符号与基础字符合并:
import unicodedata sentence = "nejon ámas̈hó T̃iqu c̈ab̃op" normalized_text = unicodedata.normalize('NFC', sentence) grapheme_list = [] current_cluster = [] for char in normalized_text: # 组合变音符号的Unicode类别以'M'(Mark)开头 if unicodedata.category(char).startswith('M'): if current_cluster: current_cluster.append(char) else: if current_cluster: grapheme_list.append(''.join(current_cluster)) current_cluster = [char] # 处理最后一个字符簇 if current_cluster: grapheme_list.append(''.join(current_cluster)) # 去重、过滤空格、排序 unique_chars = sorted([char for char in set(grapheme_list) if char != ' ']) print(unique_chars) # 输出与方法一一致
关键说明
- NFC标准化:
unicodedata.normalize('NFC', text)会将字符转换为预组合形式(若存在对应预组合字符),确保变音符号与基础字符的组合格式统一,避免不必要的字符拆分。 - Grapheme簇:视觉上的单个字符可能由多个Unicode码点组成(如
s̈是s+¨),两种方法都是将这类组合视为一个整体,从而正确提取唯一字符。
内容的提问来源于stack exchange,提问作者lisa
相关产品推荐
相关产品推荐

