Python读取东方语言PDF出现文本乱码问题求助
解决PDF东方语言(孟加拉语)文本提取乱码问题
问题描述
开发REST API时,需从固定格式PDF提取表格转换为JSON数据,使用tabula-py、PyPDF2、pdfminer等库均出现孟加拉语文本乱码(如提取结果为"িনবেনর তািরখ",正确内容应为নিবন্ধনের তারিক)。因单次请求需处理400-500页PDF,无法采用速度过慢的OCR方案。
解决方案
1. 强制JVM环境使用UTF-8编码
tabula-py依赖tabula-java运行,需先配置JVM参数确保字符处理采用UTF-8:
import os os.environ["JAVA_TOOL_OPTIONS"] = "-Dfile.encoding=UTF-8"
2. 修正tabula的提取参数
将原代码中encoding='sujata'替换为encoding='utf-8',同时添加lattice=True(针对固定格式表格提升提取精度):
tables = tabula.read_pdf(pdf_path, pages='all', multiple_tables=True, encoding='utf-8', lattice=True)
3. 修复Unicode字符分离问题
乱码多因孟加拉语的变音符号与主字符拆分导致,使用unicodedata的NFC规范化合并字符:
import unicodedata def normalize_unicode(text): if isinstance(text, str): return unicodedata.normalize('NFC', text).replace("\r", " ").strip() return text
在数据清理环节调用该函数处理键值对:
cleaned_key = normalize_unicode(key) cleaned_value = normalize_unicode(value)
4. 尝试替代库camelot-py
若tabula仍存在问题,可尝试同属文本提取类的camelot-py(适配大文件处理速度):
import camelot tables = camelot.read_pdf(pdf_path, pages='all', flavor='lattice') json_data = [] for table in tables: json_data.extend(table.df.to_dict(orient="records"))
修改后的完整代码
import os import tabula import json import unicodedata # 设置JVM编码为UTF-8 os.environ["JAVA_TOOL_OPTIONS"] = "-Dfile.encoding=UTF-8" # PDF文件路径 pdf_path = "small.pdf" # 使用Tabula提取表格,指定正确编码和表格识别模式 tables = tabula.read_pdf(pdf_path, pages='all', multiple_tables=True, encoding='utf-8', lattice=True) # 存储JSON数据的列表 json_data = [] # 处理每个表格 for table in tables: json_data.extend(table.to_dict(orient="records")) # Unicode字符规范化函数 def normalize_unicode(text): if isinstance(text, str): return unicodedata.normalize('NFC', text).replace("\r", " ").strip() return text # 清理字典条目 cleaned_data = [] for entry in json_data: cleaned_entry = {} for key, value in entry.items(): cleaned_key = normalize_unicode(key) cleaned_value = normalize_unicode(value) cleaned_entry[cleaned_key] = cleaned_value cleaned_data.append(cleaned_entry) # 移除无名键 def remove_unnamed_keys(data): cleaned_data = [] for entry in data: cleaned_entry = {k: v for k, v in entry.items() if not k.startswith("Unnamed") and v is not None} cleaned_data.append(cleaned_entry) return cleaned_data cleaned_data = remove_unnamed_keys(cleaned_data) # 写入JSON文件 with open("output.json", "w", encoding="utf-8") as json_file: json.dump(cleaned_data, json_file, ensure_ascii=False, indent=4)
内容的提问来源于stack exchange,提问作者Muhammad Hossain
相关产品推荐
相关产品推荐

