You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python去除字符中的变音符号

用Python移除文本文件中的变音符号

这个需求我之前处理多语言文本时经常碰到,Python自带的unicodedata模块就能完美解决,完全不用装额外依赖,上手超简单!

核心转换逻辑

首先咱们得先写个函数,把带变音符号的字符转成对应的普通字符。原理是先把Unicode字符标准化,然后过滤掉变音符号这类非字母的标记:

import unicodedata

def remove_accents(text):
    # 先把字符标准化为NFD形式(分解字符和变音标记)
    normalized_text = unicodedata.normalize('NFD', text)
    # 过滤掉所有属于"非间距标记"的字符(也就是变音符号)
    cleaned_text = ''.join(c for c in normalized_text if unicodedata.category(c) != 'Mn')
    return cleaned_text

举个测试例子:

print(remove_accents("Cafè à la carte"))  # 输出: Cafe a la carte
print(remove_accents("Pão de queijo"))    # 输出: Pao de queijo

处理单个文本文件

有了转换函数,接下来就是读取文件、转换内容、再写回文件。注意一定要用utf-8编码读写,避免乱码:

def process_single_file(input_path, output_path=None):
    # 如果没指定输出路径,就覆盖原文件(建议先备份!)
    if output_path is None:
        output_path = input_path
    
    with open(input_path, 'r', encoding='utf-8') as f:
        content = f.read()
    
    cleaned_content = remove_accents(content)
    
    with open(output_path, 'w', encoding='utf-8') as f:
        f.write(cleaned_content)

# 调用示例:处理当前目录下的example.txt,覆盖原文件
# process_single_file("example.txt")
# 或者指定输出到新文件
# process_single_file("example.txt", "example_cleaned.txt")

批量处理多个文件

如果要处理目录下所有.txt文件,可以结合os模块遍历文件:

import os

def process_directory(dir_path, output_dir=None):
    # 如果没指定输出目录,就在原目录下创建一个cleaned子目录
    if output_dir is None:
        output_dir = os.path.join(dir_path, 'cleaned')
        os.makedirs(output_dir, exist_ok=True)
    
    for filename in os.listdir(dir_path):
        if filename.endswith('.txt'):
            input_path = os.path.join(dir_path, filename)
            output_path = os.path.join(output_dir, filename)
            process_single_file(input_path, output_path)
            print(f"已处理:{filename}")

# 调用示例:处理当前目录下的所有txt文件,输出到cleaned子目录
# process_directory(".")

特殊字符补充处理

有个特殊情况要注意:德语的ß用上面的函数会直接保留,因为它本身不是带变音的字符。如果需要把它转成ss,可以在函数里加个额外替换:

def remove_accents(text):
    # 先处理特殊字符ß
    text = text.replace('ß', 'ss')
    normalized_text = unicodedata.normalize('NFD', text)
    cleaned_text = ''.join(c for c in normalized_text if unicodedata.category(c) != 'Mn')
    return cleaned_text

这样处理Straße就会变成Strasse啦~

内容的提问来源于stack exchange,提问作者sunyata

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 04:08:20