You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中使用unidecode去重音时如何保留ç字符?

去除字符串重音但保留特定字符(如ç)的Python实现

问题描述

使用unidecode库去除字符串重音时,该库会将所有带重音的字符转为无重音形式,包括ç会被转为c:

import unidecode

print(unidecode.unidecode('ááíãôç'))

输出结果:

aaioac

但需求是保留ç,期望得到:

aaiaoç

无需硬编码的解决方案

方法1:基于Unicode规范化与字符属性处理

利用Python内置的unicodedata模块,通过Unicode规范化(NFD分解)识别并保留带软音符(CEDILLA)的字符组合,其他重音标记则移除:

import unicodedata

def remove_accents_preserve_cedilla(input_str):
    # 将字符串分解为基础字符+组合标记的形式(NFD)
    normalized = unicodedata.normalize('NFD', input_str)
    output = []
    for char in normalized:
        # 组合标记(Mn类别)通常是重音符号
        if unicodedata.category(char) == 'Mn':
            # 仅保留与'c'搭配的软音符(U+0327)
            if output and output[-1] == 'c' and char == '\u0327':
                output.append(char)
            continue
        output.append(char)
    # 将分解后的字符重新组合为标准形式(NFC),恢复ç
    return unicodedata.normalize('NFC', ''.join(output))

# 测试
print(remove_accents_preserve_cedilla('ááíãôç'))  # 输出: aaiaoç

这种方法没有硬编码具体字符,而是基于Unicode字符的类别和属性判断,同时支持保留大写形式的Ç。

方法2:结合unidecode与Unicode属性校验

如果仍想使用unidecode库,可通过Unicode字符名称判断是否为带软音符的字符,将其从处理结果中还原:

import unidecode
import unicodedata

def remove_accents_with_unidecode_preserve_cedilla(input_str):
    processed = unidecode.unidecode(input_str)
    result = list(processed)
    for idx, char in enumerate(input_str):
        # 通过Unicode名称识别带软音符的字符
        if 'CEDILLA' in unicodedata.name(char, ''):
            result[idx] = char
    return ''.join(result)

# 测试
print(remove_accents_with_unidecode_preserve_cedilla('ááíãôç'))  # 输出: aaiaoç

该方法通过字符的Unicode元数据判断,无需硬编码具体字符值,扩展性更强。


内容的提问来源于stack exchange,提问作者Vitor Boldrin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 00:01:41