You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python的tabula读取PDF时遭遇UnicodeDecodeError求助

使用tabula读取PDF时遇到UnicodeDecodeError及JAI依赖问题

问题情况

  • 用Python的tabula包读取PDF文件,触发UnicodeDecodeError
  • 用chardet检测文件编码,结果返回None

原代码

from tabula import read_pdf
from tabulate import tabulate

df = read_pdf(open(r"C:\Users\rohit\Downloads\Capstone Data\\" + "CITY OF ROCHESTER.pdf",'rb'),pages="all") #PDF文件路径
print(tabulate(df))

报错信息

1. JAI依赖警告

Oct 05, 2022 2:30:21 PM org.apache.pdfbox.contentstream.PDFStreamEngine operatorException
SEVERE: Cannot read JPEG2000 image: Java Advanced Imaging (JAI) Image I/O Tools are not installed

2. UnicodeDecodeError栈信息

---------------------------------------------------------------------------
UnicodeDecodeError                        Traceback (most recent call last)
Cell In [2], line 5
----> 5 df = read_pdf(open(r"C:\Users\rohit\Downloads\Capstone Data\\" + "CITY OF ROCHESTER, MINNESOTA - HEALTH CARE FACILITIES REVENUE BONDS, (MAYO CLINIC) SERIES 2022.pdf",'rb'),pages="all") #PDF文件路径

File ~\AppData\Local\Programs\Python\Python310\lib\site-packages\tabula\io.py:434, in read_pdf(input_path, output_format, encoding, java_options, pandas_options, multiple_tables, user_agent, use_raw_url, pages, guess, area, relative_area, lattice, stream, password, silent, columns, format, batch, output_path, options)
    432 fmt = tabula_options.format
    433 if fmt == "JSON":
--> 434     raw_json: List[Any] = json.loads(output.decode(encoding))
    435     if multiple_tables:
    436         return _extract_from(raw_json, pandas_options)

UnicodeDecodeError: 'utf-8' codec can't decode byte 0x86 in position 1962: invalid start byte

解决方法

处理JAI依赖警告

PDF中包含JPEG2000格式图片,tabula依赖的PDFBox需要JAI Image I/O Tools才能解码。如果不需要提取图片可以忽略该警告,若要彻底解决:

  • 下载JAI Image I/O Tools的jar包(如jai-imageio-core-1.1.jar)
  • 将jar包放到Python环境下的Lib\site-packages\tabula\jar目录中

解决UnicodeDecodeError

  1. 避免直接传文件对象:直接给read_pdf传文件路径字符串,tabula会自行处理文件读取,减少编码问题
  2. 指定正确编码:0x86字节对应cp1252编码中的特殊符号,尝试指定encoding='cp1252'或encoding='latin-1'(latin-1能兼容所有字节,不会触发解码错误)

修改后的代码:

from tabula import read_pdf
from tabulate import tabulate

# 直接传入文件路径,指定编码为cp1252
df = read_pdf(r"C:\Users\rohit\Downloads\Capstone Data\CITY OF ROCHESTER.pdf", pages="all", encoding='cp1252')
print(tabulate(df))

注:chardet对二进制PDF的编码检测无效,直接尝试Windows常见编码(cp1252)或通用编码(latin-1)更可靠

内容的提问来源于stack exchange,提问作者Rohit Bhargav Peesa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 11:42:03