You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将PyMuPDF提取的Unicode字符转换为对应ASCII等效字符?

解决PyMuPDF提取文本时Unicode标点转ASCII的问题

有几种可行的转换方法,比单纯正则替换更高效全面:

方法一:自定义映射替换

直接创建Unicode标点到ASCII的映射表,批量替换,逻辑直观且易扩展:

import fitz

def replace_smart_punctuation(text):
    # 定义Unicode标点与ASCII的对应关系
    char_mapping = {
        '\u201C': '"',  # 左双引号
        '\u201D': '"',  # 右双引号
        '\u2018': "'",  # 左单引号
        '\u2019': "'",  # 右单引号
        '\u2013': '-',  # 短破折号
        '\u2014': '--', # 长破折号
        '\u2026': '...' # 省略号
    }
    # 批量替换文本中的目标字符
    for unicode_char, ascii_char in char_mapping.items():
        text = text.replace(unicode_char, ascii_char)
    return text

# 实际使用示例
doc = fitz.open("target_file.pdf")
raw_text = ""
for page in doc:
    raw_text += page.get_text()
processed_text = replace_smart_punctuation(raw_text)

方法二:正则表达式批量替换

如果偏好正则实现,可以一次性匹配同类Unicode字符并替换:

import re

def clean_special_chars(text):
    # 替换双引号类字符
    text = re.sub(r'[\u201C\u201D]', '"', text)
    # 替换单引号类字符
    text = re.sub(r'[\u2018\u2019]', "'", text)
    # 可按需扩展替换其他符号
    text = re.sub(r'[\u2013\u2014]', '-', text)
    text = re.sub(r'\u2026', '...', text)
    return text

方法三:使用unidecode库处理全场景转换

如果需要处理更多非标点类Unicode字符转ASCII,推荐使用unidecode库,它能覆盖更广泛的Unicode到ASCII的转换需求:

pip install unidecode
from unidecode import unidecode

# 直接转换提取的文本
processed_text = unidecode(raw_text)

注意:unidecode会将所有非ASCII字符转换为相近的ASCII形式(比如重音字母也会被处理),需根据实际需求选择。

内容的提问来源于stack exchange,提问作者mhay10

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 04:55:42