使用Python的tabula读取PDF时遭遇UnicodeDecodeError求助
使用tabula读取PDF时遇到UnicodeDecodeError及JAI依赖问题
问题情况
- 用Python的tabula包读取PDF文件,触发
UnicodeDecodeError - 用chardet检测文件编码,结果返回
None
原代码
from tabula import read_pdf from tabulate import tabulate df = read_pdf(open(r"C:\Users\rohit\Downloads\Capstone Data\\" + "CITY OF ROCHESTER.pdf",'rb'),pages="all") #PDF文件路径 print(tabulate(df))
报错信息
1. JAI依赖警告
Oct 05, 2022 2:30:21 PM org.apache.pdfbox.contentstream.PDFStreamEngine operatorException
SEVERE: Cannot read JPEG2000 image: Java Advanced Imaging (JAI) Image I/O Tools are not installed
2. UnicodeDecodeError栈信息
--------------------------------------------------------------------------- UnicodeDecodeError Traceback (most recent call last) Cell In [2], line 5 ----> 5 df = read_pdf(open(r"C:\Users\rohit\Downloads\Capstone Data\\" + "CITY OF ROCHESTER, MINNESOTA - HEALTH CARE FACILITIES REVENUE BONDS, (MAYO CLINIC) SERIES 2022.pdf",'rb'),pages="all") #PDF文件路径 File ~\AppData\Local\Programs\Python\Python310\lib\site-packages\tabula\io.py:434, in read_pdf(input_path, output_format, encoding, java_options, pandas_options, multiple_tables, user_agent, use_raw_url, pages, guess, area, relative_area, lattice, stream, password, silent, columns, format, batch, output_path, options) 432 fmt = tabula_options.format 433 if fmt == "JSON": --> 434 raw_json: List[Any] = json.loads(output.decode(encoding)) 435 if multiple_tables: 436 return _extract_from(raw_json, pandas_options) UnicodeDecodeError: 'utf-8' codec can't decode byte 0x86 in position 1962: invalid start byte
解决方法
处理JAI依赖警告
PDF中包含JPEG2000格式图片,tabula依赖的PDFBox需要JAI Image I/O Tools才能解码。如果不需要提取图片可以忽略该警告,若要彻底解决:
- 下载JAI Image I/O Tools的jar包(如
jai-imageio-core-1.1.jar) - 将jar包放到Python环境下的
Lib\site-packages\tabula\jar目录中
解决UnicodeDecodeError
- 避免直接传文件对象:直接给
read_pdf传文件路径字符串,tabula会自行处理文件读取,减少编码问题 - 指定正确编码:0x86字节对应cp1252编码中的特殊符号,尝试指定
encoding='cp1252'或encoding='latin-1'(latin-1能兼容所有字节,不会触发解码错误)
修改后的代码:
from tabula import read_pdf from tabulate import tabulate # 直接传入文件路径,指定编码为cp1252 df = read_pdf(r"C:\Users\rohit\Downloads\Capstone Data\CITY OF ROCHESTER.pdf", pages="all", encoding='cp1252') print(tabulate(df))
注:chardet对二进制PDF的编码检测无效,直接尝试Windows常见编码(cp1252)或通用编码(latin-1)更可靠
内容的提问来源于stack exchange,提问作者Rohit Bhargav Peesa
相关产品推荐
相关产品推荐

