You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何获取含组合字符的文本中的唯一字符列表?

如何正确提取包含组合变音符号文本中的唯一字符

当文本包含s̈、b̃这类由基础字符+变音符号组成的组合字符时,直接用set()拆分字符串会将它们拆分为独立的基础字符和变音符号,导致结果不符合预期。以下是两种解决方法:

方法一:使用第三方regex库(简便推荐)

regex库支持匹配Unicode grapheme簇(视觉上的单个字符),能直接将组合字符作为整体提取:

  1. 先安装依赖:
pip install regex
  1. 示例代码:
import unicodedata
import regex

sentence = "nejon ámas̈hó T̃iqu c̈ab̃op"

# 标准化字符为NFC格式,确保组合形式统一
normalized_text = unicodedata.normalize('NFC', sentence)

# 提取所有视觉上的单个字符(grapheme簇)
grapheme_list = regex.findall(r'\X', normalized_text)

# 去重、过滤空格、排序(排序可选,按需调整)
unique_chars = sorted([char for char in set(grapheme_list) if char != ' '])

print(unique_chars)
# 输出:['a', 'á', 'b̃', 'c̈', 'e', 'h', 'i', 'j', 'm', 'n', 'o', 'ó', 'p', 'q', 's̈', 'T̃', 'u']

方法二:手动遍历处理(无需第三方库)

通过判断Unicode字符类别,将组合变音符号与基础字符合并:

import unicodedata

sentence = "nejon ámas̈hó T̃iqu c̈ab̃op"
normalized_text = unicodedata.normalize('NFC', sentence)

grapheme_list = []
current_cluster = []

for char in normalized_text:
    # 组合变音符号的Unicode类别以'M'(Mark)开头
    if unicodedata.category(char).startswith('M'):
        if current_cluster:
            current_cluster.append(char)
    else:
        if current_cluster:
            grapheme_list.append(''.join(current_cluster))
        current_cluster = [char]
# 处理最后一个字符簇
if current_cluster:
    grapheme_list.append(''.join(current_cluster))

# 去重、过滤空格、排序
unique_chars = sorted([char for char in set(grapheme_list) if char != ' '])

print(unique_chars)
# 输出与方法一一致

关键说明

  • NFC标准化:unicodedata.normalize('NFC', text)会将字符转换为预组合形式(若存在对应预组合字符),确保变音符号与基础字符的组合格式统一,避免不必要的字符拆分。
  • Grapheme簇:视觉上的单个字符可能由多个Unicode码点组成(如s̈是s+¨),两种方法都是将这类组合视为一个整体,从而正确提取唯一字符。

内容的提问来源于stack exchange,提问作者lisa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.12 18:46:04