如何使用Python去除字符中的变音符号
用Python移除文本文件中的变音符号
这个需求我之前处理多语言文本时经常碰到,Python自带的unicodedata模块就能完美解决,完全不用装额外依赖,上手超简单!
核心转换逻辑
首先咱们得先写个函数,把带变音符号的字符转成对应的普通字符。原理是先把Unicode字符标准化,然后过滤掉变音符号这类非字母的标记:
import unicodedata def remove_accents(text): # 先把字符标准化为NFD形式(分解字符和变音标记) normalized_text = unicodedata.normalize('NFD', text) # 过滤掉所有属于"非间距标记"的字符(也就是变音符号) cleaned_text = ''.join(c for c in normalized_text if unicodedata.category(c) != 'Mn') return cleaned_text
举个测试例子:
print(remove_accents("Cafè à la carte")) # 输出: Cafe a la carte print(remove_accents("Pão de queijo")) # 输出: Pao de queijo
处理单个文本文件
有了转换函数,接下来就是读取文件、转换内容、再写回文件。注意一定要用utf-8编码读写,避免乱码:
def process_single_file(input_path, output_path=None): # 如果没指定输出路径,就覆盖原文件(建议先备份!) if output_path is None: output_path = input_path with open(input_path, 'r', encoding='utf-8') as f: content = f.read() cleaned_content = remove_accents(content) with open(output_path, 'w', encoding='utf-8') as f: f.write(cleaned_content) # 调用示例:处理当前目录下的example.txt,覆盖原文件 # process_single_file("example.txt") # 或者指定输出到新文件 # process_single_file("example.txt", "example_cleaned.txt")
批量处理多个文件
如果要处理目录下所有.txt文件,可以结合os模块遍历文件:
import os def process_directory(dir_path, output_dir=None): # 如果没指定输出目录,就在原目录下创建一个cleaned子目录 if output_dir is None: output_dir = os.path.join(dir_path, 'cleaned') os.makedirs(output_dir, exist_ok=True) for filename in os.listdir(dir_path): if filename.endswith('.txt'): input_path = os.path.join(dir_path, filename) output_path = os.path.join(output_dir, filename) process_single_file(input_path, output_path) print(f"已处理:{filename}") # 调用示例:处理当前目录下的所有txt文件,输出到cleaned子目录 # process_directory(".")
特殊字符补充处理
有个特殊情况要注意:德语的ß用上面的函数会直接保留,因为它本身不是带变音的字符。如果需要把它转成ss,可以在函数里加个额外替换:
def remove_accents(text): # 先处理特殊字符ß text = text.replace('ß', 'ss') normalized_text = unicodedata.normalize('NFD', text) cleaned_text = ''.join(c for c in normalized_text if unicodedata.category(c) != 'Mn') return cleaned_text
这样处理Straße就会变成Strasse啦~
内容的提问来源于stack exchange,提问作者sunyata
相关产品推荐
相关产品推荐

