You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python:如何将UTF-8字符尽可能转换为指定拉丁类字符?

Python中替换UTF-8特殊字符为指定范围内近似字符的方案

核心思路

先通过字符归一化工具将特殊UTF-8字符转换为近似的拉丁/希腊字符或组合,再过滤掉不在指定范围内的字符。

步骤1:使用unidecode库做近似转换

unidecode库专门用于将Unicode字符转换为最接近的拉丁字符,能处理š→s、Ó→O、œ→oe这类需求。

首先安装库:

pip install unidecode

基础转换示例:

from unidecode import unidecode

print(unidecode("š"))  # 输出: s
print(unidecode("Ó"))  # 输出: O
print(unidecode("œ"))  # 输出: oe
print(unidecode("†"))  # 输出: +(后续过滤会移除)

步骤2:过滤保留指定范围内的字符

定义允许的字符集合,遍历转换后的文本,只保留符合要求的字符:

import re

allowed_chars = r'[a-zA-Z0-9.,:?!@$€α-ωΑ-Ω]'

def filter_allowed(text):
    return re.sub(f'[^{allowed_chars}]', '', text)

完整组合方案

将转换和过滤整合为一个函数:

from unidecode import unidecode
import re

allowed_chars = r'[a-zA-Z0-9.,:?!@$€α-ωΑ-Ω]'

def normalize_special_chars(text):
    # 第一步:转换为近似字符
    converted = unidecode(text)
    # 第二步:过滤保留允许的字符
    filtered = re.sub(f'[^{allowed_chars}]', '', converted)
    return filtered

# 测试示例
test_text = "ŠÓœ†Hello! 123 €αβγ"
result = normalize_special_chars(test_text)
print(result)  # 输出: SOoeHello!123€αβγ

说明

  • unidecode的转换基于字符的发音和视觉近似,基本能覆盖大部分拉丁衍生字符的需求
  • 无法转换为目标范围内的字符(如表情、特殊符号)会被过滤函数直接移除
  • 若需调整近似规则,可自定义部分字符的映射表,优先替换再调用unidecode

内容的提问来源于stack exchange,提问作者Uwe.Schneider

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 02:05:05