Python中使用unidecode去重音时如何保留ç字符?
去除字符串重音但保留特定字符(如ç)的Python实现
问题描述
使用unidecode库去除字符串重音时,该库会将所有带重音的字符转为无重音形式,包括ç会被转为c:
import unidecode print(unidecode.unidecode('ááíãôç'))
输出结果:
aaioac
但需求是保留ç,期望得到:
aaiaoç
无需硬编码的解决方案
方法1:基于Unicode规范化与字符属性处理
利用Python内置的unicodedata模块,通过Unicode规范化(NFD分解)识别并保留带软音符(CEDILLA)的字符组合,其他重音标记则移除:
import unicodedata def remove_accents_preserve_cedilla(input_str): # 将字符串分解为基础字符+组合标记的形式(NFD) normalized = unicodedata.normalize('NFD', input_str) output = [] for char in normalized: # 组合标记(Mn类别)通常是重音符号 if unicodedata.category(char) == 'Mn': # 仅保留与'c'搭配的软音符(U+0327) if output and output[-1] == 'c' and char == '\u0327': output.append(char) continue output.append(char) # 将分解后的字符重新组合为标准形式(NFC),恢复ç return unicodedata.normalize('NFC', ''.join(output)) # 测试 print(remove_accents_preserve_cedilla('ááíãôç')) # 输出: aaiaoç
这种方法没有硬编码具体字符,而是基于Unicode字符的类别和属性判断,同时支持保留大写形式的Ç。
方法2:结合unidecode与Unicode属性校验
如果仍想使用unidecode库,可通过Unicode字符名称判断是否为带软音符的字符,将其从处理结果中还原:
import unidecode import unicodedata def remove_accents_with_unidecode_preserve_cedilla(input_str): processed = unidecode.unidecode(input_str) result = list(processed) for idx, char in enumerate(input_str): # 通过Unicode名称识别带软音符的字符 if 'CEDILLA' in unicodedata.name(char, ''): result[idx] = char return ''.join(result) # 测试 print(remove_accents_with_unidecode_preserve_cedilla('ááíãôç')) # 输出: aaiaoç
该方法通过字符的Unicode元数据判断,无需硬编码具体字符值,扩展性更强。
内容的提问来源于stack exchange,提问作者Vitor Boldrin
相关产品推荐
相关产品推荐

