You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用PyPDF2提取PDF时数字识别错误的修复咨询

问题:PyPDF2提取PDF表格时数字拼接错误

我用PyPDF2编写了PDF文本提取函数,目标PDF包含大量表格,但PyPDF2无法正确识别其中的数字——当Value列是单个数字时,会把Value列的数字和%列的数字拼接在一起。

错误提取结果:

sugar?
Value Label Unweighted
Frequency%
1- 40.1 %
2- 10 0.3 %
-24 -Value Label Unweighted
Frequency%
3- 60.2 %
4- 40.1 %

正确的提取结果应该是:

sugar?
Value Label Unweighted
Frequency%
1- 4 0.1 %
2- 10 0.3 %
-24 -Value Label Unweighted
Frequency%
3- 6 0.2 %
4- 4 0.1 %

我的提取函数代码如下:

def read_pdfs(dictionary):
#     dfs_names_text_dic = {}
    for pdf_file, name in dictionary.items():
        pdfFileObj = open(pdf_file, 'rb')
        pdfReader = PyPDF2.PdfReader(pdfFileObj)
        output = []
        for i in range(3,2000):
            pageObj = pdfReader.pages[i]
            output.append(pageObj.extract_text())
#         dfs_names_text_dic[name] = output
        pdfFileObj.close()
    return output

解决方案

PyPDF2的文本提取逻辑基于文本块位置排序,不具备表格结构识别能力,处理表格时极易出现文本拼接错位问题。更可靠的解决方式是使用专门支持表格提取的库,或对提取后的文本做针对性修正:

方案1:使用pdfplumber提取表格

pdfplumber可精确识别PDF表格结构,按单元格提取内容,从根源避免数字拼接错误。

先安装依赖:

pip install pdfplumber

修改后的函数示例:

import pdfplumber

def read_pdfs(dictionary):
    output = []
    for pdf_file, name in dictionary.items():
        with pdfplumber.open(pdf_file) as pdf:
            # 遍历指定页码范围(3到1999页),注意pdfplumber页码从1开始
            for page_num in range(3, 2000):
                if page_num > len(pdf.pages):
                    break
                page = pdf.pages[page_num - 1]
                # 提取页面所有表格并转换为文本格式
                tables = page.extract_tables()
                for table in tables:
                    table_text = "\n".join(["\t".join(str(cell) if cell is not None else "" for cell in row) for row in table])
                    output.append(table_text)
    return output

方案2:使用tabula-py提取表格

tabula-py基于Java的Tabula库,擅长提取结构化表格,支持直接导出为DataFrame格式。

先安装依赖:

pip install tabula-py

修改后的函数示例:

import tabula

def read_pdfs(dictionary):
    output = []
    for pdf_file, name in dictionary.items():
        # 提取指定页码范围的表格
        dfs = tabula.read_pdf(pdf_file, pages="3-1999", multiple_tables=True)
        for df in dfs:
            # 将DataFrame转换为文本格式
            output.append(df.to_string(index=False))
    return output

方案3:PyPDF2文本后修正(临时方案)

如果必须使用PyPDF2,可通过正则表达式对提取后的文本做针对性修正,仅适用于当前发现的特定格式问题:

import re
import PyPDF2

def fix_extracted_text(text):
    # 匹配单个数字+0.x的错误拼接,替换为分开的格式
    corrected_text = re.sub(r'(\d)(0\.\d)', r'\1 \2', text)
    return corrected_text

def read_pdfs(dictionary):
    output = []
    for pdf_file, name in dictionary.items():
        with open(pdf_file, 'rb') as pdfFileObj:
            pdfReader = PyPDF2.PdfReader(pdfFileObj)
            for i in range(3,2000):
                if i >= len(pdfReader.pages):
                    break
                pageObj = pdfReader.pages[i]
                raw_text = pageObj.extract_text()
                corrected_text = fix_extracted_text(raw_text)
                output.append(corrected_text)
    return output

注意:正则修正局限性强,若PDF存在其他类型的表格错位,无法覆盖所有情况。


内容的提问来源于stack exchange,提问作者ella

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.22 20:22:46