使用PyPDF2提取PDF时数字识别错误的修复咨询
问题:PyPDF2提取PDF表格时数字拼接错误
我用PyPDF2编写了PDF文本提取函数,目标PDF包含大量表格,但PyPDF2无法正确识别其中的数字——当Value列是单个数字时,会把Value列的数字和%列的数字拼接在一起。
错误提取结果:
sugar? Value Label Unweighted Frequency% 1- 40.1 % 2- 10 0.3 % -24 -Value Label Unweighted Frequency% 3- 60.2 % 4- 40.1 %
正确的提取结果应该是:
sugar? Value Label Unweighted Frequency% 1- 4 0.1 % 2- 10 0.3 % -24 -Value Label Unweighted Frequency% 3- 6 0.2 % 4- 4 0.1 %
我的提取函数代码如下:
def read_pdfs(dictionary): # dfs_names_text_dic = {} for pdf_file, name in dictionary.items(): pdfFileObj = open(pdf_file, 'rb') pdfReader = PyPDF2.PdfReader(pdfFileObj) output = [] for i in range(3,2000): pageObj = pdfReader.pages[i] output.append(pageObj.extract_text()) # dfs_names_text_dic[name] = output pdfFileObj.close() return output
解决方案
PyPDF2的文本提取逻辑基于文本块位置排序,不具备表格结构识别能力,处理表格时极易出现文本拼接错位问题。更可靠的解决方式是使用专门支持表格提取的库,或对提取后的文本做针对性修正:
方案1:使用pdfplumber提取表格
pdfplumber可精确识别PDF表格结构,按单元格提取内容,从根源避免数字拼接错误。
先安装依赖:
pip install pdfplumber
修改后的函数示例:
import pdfplumber def read_pdfs(dictionary): output = [] for pdf_file, name in dictionary.items(): with pdfplumber.open(pdf_file) as pdf: # 遍历指定页码范围(3到1999页),注意pdfplumber页码从1开始 for page_num in range(3, 2000): if page_num > len(pdf.pages): break page = pdf.pages[page_num - 1] # 提取页面所有表格并转换为文本格式 tables = page.extract_tables() for table in tables: table_text = "\n".join(["\t".join(str(cell) if cell is not None else "" for cell in row) for row in table]) output.append(table_text) return output
方案2:使用tabula-py提取表格
tabula-py基于Java的Tabula库,擅长提取结构化表格,支持直接导出为DataFrame格式。
先安装依赖:
pip install tabula-py
修改后的函数示例:
import tabula def read_pdfs(dictionary): output = [] for pdf_file, name in dictionary.items(): # 提取指定页码范围的表格 dfs = tabula.read_pdf(pdf_file, pages="3-1999", multiple_tables=True) for df in dfs: # 将DataFrame转换为文本格式 output.append(df.to_string(index=False)) return output
方案3:PyPDF2文本后修正(临时方案)
如果必须使用PyPDF2,可通过正则表达式对提取后的文本做针对性修正,仅适用于当前发现的特定格式问题:
import re import PyPDF2 def fix_extracted_text(text): # 匹配单个数字+0.x的错误拼接,替换为分开的格式 corrected_text = re.sub(r'(\d)(0\.\d)', r'\1 \2', text) return corrected_text def read_pdfs(dictionary): output = [] for pdf_file, name in dictionary.items(): with open(pdf_file, 'rb') as pdfFileObj: pdfReader = PyPDF2.PdfReader(pdfFileObj) for i in range(3,2000): if i >= len(pdfReader.pages): break pageObj = pdfReader.pages[i] raw_text = pageObj.extract_text() corrected_text = fix_extracted_text(raw_text) output.append(corrected_text) return output
注意:正则修正局限性强,若PDF存在其他类型的表格错位,无法覆盖所有情况。
内容的提问来源于stack exchange,提问作者ella
相关产品推荐
相关产品推荐

