如何将PyMuPDF提取的Unicode字符转换为对应ASCII等效字符?
解决PyMuPDF提取文本时Unicode标点转ASCII的问题
有几种可行的转换方法,比单纯正则替换更高效全面:
方法一:自定义映射替换
直接创建Unicode标点到ASCII的映射表,批量替换,逻辑直观且易扩展:
import fitz def replace_smart_punctuation(text): # 定义Unicode标点与ASCII的对应关系 char_mapping = { '\u201C': '"', # 左双引号 '\u201D': '"', # 右双引号 '\u2018': "'", # 左单引号 '\u2019': "'", # 右单引号 '\u2013': '-', # 短破折号 '\u2014': '--', # 长破折号 '\u2026': '...' # 省略号 } # 批量替换文本中的目标字符 for unicode_char, ascii_char in char_mapping.items(): text = text.replace(unicode_char, ascii_char) return text # 实际使用示例 doc = fitz.open("target_file.pdf") raw_text = "" for page in doc: raw_text += page.get_text() processed_text = replace_smart_punctuation(raw_text)
方法二:正则表达式批量替换
如果偏好正则实现,可以一次性匹配同类Unicode字符并替换:
import re def clean_special_chars(text): # 替换双引号类字符 text = re.sub(r'[\u201C\u201D]', '"', text) # 替换单引号类字符 text = re.sub(r'[\u2018\u2019]', "'", text) # 可按需扩展替换其他符号 text = re.sub(r'[\u2013\u2014]', '-', text) text = re.sub(r'\u2026', '...', text) return text
方法三:使用unidecode库处理全场景转换
如果需要处理更多非标点类Unicode字符转ASCII,推荐使用unidecode库,它能覆盖更广泛的Unicode到ASCII的转换需求:
pip install unidecode
from unidecode import unidecode # 直接转换提取的文本 processed_text = unidecode(raw_text)
注意:unidecode会将所有非ASCII字符转换为相近的ASCII形式(比如重音字母也会被处理),需根据实际需求选择。
内容的提问来源于stack exchange,提问作者mhay10
相关产品推荐
相关产品推荐

