You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python读取东方语言PDF出现文本乱码问题求助

解决PDF东方语言(孟加拉语)文本提取乱码问题

问题描述

开发REST API时,需从固定格式PDF提取表格转换为JSON数据,使用tabula-py、PyPDF2、pdfminer等库均出现孟加拉语文本乱码(如提取结果为"িনবেনর তািরখ",正确内容应为নিবন্ধনের তারিক)。因单次请求需处理400-500页PDF,无法采用速度过慢的OCR方案。

解决方案

1. 强制JVM环境使用UTF-8编码

tabula-py依赖tabula-java运行,需先配置JVM参数确保字符处理采用UTF-8:

import os
os.environ["JAVA_TOOL_OPTIONS"] = "-Dfile.encoding=UTF-8"

2. 修正tabula的提取参数

将原代码中encoding='sujata'替换为encoding='utf-8',同时添加lattice=True(针对固定格式表格提升提取精度):

tables = tabula.read_pdf(pdf_path, pages='all', multiple_tables=True, encoding='utf-8', lattice=True)

3. 修复Unicode字符分离问题

乱码多因孟加拉语的变音符号与主字符拆分导致,使用unicodedata的NFC规范化合并字符:

import unicodedata

def normalize_unicode(text):
    if isinstance(text, str):
        return unicodedata.normalize('NFC', text).replace("\r", " ").strip()
    return text

在数据清理环节调用该函数处理键值对:

cleaned_key = normalize_unicode(key)
cleaned_value = normalize_unicode(value)

4. 尝试替代库camelot-py

若tabula仍存在问题,可尝试同属文本提取类的camelot-py(适配大文件处理速度):

import camelot

tables = camelot.read_pdf(pdf_path, pages='all', flavor='lattice')
json_data = []
for table in tables:
    json_data.extend(table.df.to_dict(orient="records"))

修改后的完整代码

import os
import tabula
import json
import unicodedata

# 设置JVM编码为UTF-8
os.environ["JAVA_TOOL_OPTIONS"] = "-Dfile.encoding=UTF-8"

# PDF文件路径
pdf_path = "small.pdf"

# 使用Tabula提取表格,指定正确编码和表格识别模式
tables = tabula.read_pdf(pdf_path, pages='all', multiple_tables=True, encoding='utf-8', lattice=True)

# 存储JSON数据的列表
json_data = []

# 处理每个表格
for table in tables:
    json_data.extend(table.to_dict(orient="records"))

# Unicode字符规范化函数
def normalize_unicode(text):
    if isinstance(text, str):
        return unicodedata.normalize('NFC', text).replace("\r", " ").strip()
    return text

# 清理字典条目
cleaned_data = []
for entry in json_data:
    cleaned_entry = {}
    for key, value in entry.items():
        cleaned_key = normalize_unicode(key)
        cleaned_value = normalize_unicode(value)
        cleaned_entry[cleaned_key] = cleaned_value
    cleaned_data.append(cleaned_entry)

# 移除无名键
def remove_unnamed_keys(data):
    cleaned_data = []
    for entry in data:
        cleaned_entry = {k: v for k, v in entry.items() if not k.startswith("Unnamed") and v is not None}
        cleaned_data.append(cleaned_entry)
    return cleaned_data

cleaned_data = remove_unnamed_keys(cleaned_data)

# 写入JSON文件
with open("output.json", "w", encoding="utf-8") as json_file:
    json.dump(cleaned_data, json_file, ensure_ascii=False, indent=4)

内容的提问来源于stack exchange,提问作者Muhammad Hossain

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 11:23:32