You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python提取PDF表格时如何正确去除脚注引用

使用Python提取PDF表格时如何正确去除脚注引用

我太懂这种烦恼了——PDF表格里的脚注上标数字总跟主数字粘在一起,比如752^5直接变成7525,把数据搞得乱七八糟。你用pdfplumber过滤字符的思路其实是对的,但全局的平均阈值太泛用,容易要么漏删要么误删。下面给你几个更靠谱的解决办法:

方案一:优化PDFPlumber的上标字符过滤逻辑

原代码用全局平均的字体大小和top位置来判断上标,很容易因为页面里有其他小字体内容(比如表头的小标注)而误判。我们改成按行分组判断,同一行的字符放在一起分析,精准度会高很多:

import pdfplumber
import pandas as pd
from IPython.display import display

pdf_path = "reports_lcc_2024_LTA BPLRT LCC Report 2024 - signed 1.pdf"
page_number = 6  # Page 7 in actual PDF
table_number = 3  # Table to extract (1-based index)

with pdfplumber.open(pdf_path) as pdf:
    page = pdf.pages[page_number]
    
    # 把字符按行分组:用top值取整作为行的标识,同一行的字符top值接近
    chars_by_row = {}
    for c in page.chars:
        row_key = round(float(c['top']))
        if row_key not in chars_by_row:
            chars_by_row[row_key] = []
        chars_by_row[row_key].append(c)
    
    filtered_chars = []
    for row_chars in chars_by_row.values():
        # 计算当前行的平均字体大小,上标字体通常是主字体的70%以下
        row_sizes = [float(c['size']) for c in row_chars]
        avg_row_size = sum(row_sizes) / len(row_sizes)
        size_thresh = avg_row_size * 0.7
        
        # 过滤当前行里的上标数字:小字体、是数字,且属于脚注的上标
        for c in row_chars:
            char_size = float(c['size'])
            if not (char_size < size_thresh and c['text'].isdigit()):
                filtered_chars.append(c)
    
    # 替换页面的字符列表,用pdfplumber的内部API修改
    page._objs["char"] = filtered_chars
    
    # 提取并处理表格
    tables = page.extract_tables()
    if tables and len(tables) >= table_number:
        table = tables[table_number - 1]
        df = pd.DataFrame(table[1:], columns=table[0])
        display(df)
    else:
        print(f"No table #{table_number} found.")

这个思路的核心是针对每行文本单独判断,不会因为页面里其他区域的小字体内容干扰表格的字符过滤,能更精准地识别并移除脚注上标。

方案二:提取表格后用正则批量清理(双保险)

有时候pdfplumber的字符过滤可能还是会有漏网之鱼,比如某些上标字体大小和主文本差距不大。这时候我们可以在提取出DataFrame之后,用正则表达式做批量清理:

import re

# 定义清理函数:移除数字末尾的脚注上标数字(比如把7525变回752)
def clean_footnotes(text):
    # 正则规则:匹配紧跟在数字后面的1-2位数字(假设脚注是短数字)
    return re.sub(r'(?<=\d)\d{1,2}$', '', str(text))

# 对DataFrame的所有单元格应用清理
df = df.applymap(clean_footnotes)
display(df)

这个方法属于事后补救,和方案一配合使用,基本能覆盖大部分脚注问题。

方案三:换用Camelot工具提取表格(备选)

如果pdfplumber还是搞不定,试试Camelot——它对带格式的表格支持更友好,尤其是流式表格:

import camelot
import re

# 用Camelot提取表格,flavor='stream'适合非严格线框的表格
tables = camelot.read_pdf(pdf_path, pages=str(page_number+1), flavor='stream')
if len(tables) >= table_number:
    df = tables[table_number-1].df
    # 同样用正则清理脚注
    df = df.applymap(clean_footnotes)
    display(df)
else:
    print(f"No table #{table_number} found.")

最后总结

  1. 优先用方案一优化pdfplumber的字符过滤,按行判断上标更精准;
  2. 配合方案二的正则清理,双管齐下解决漏网之鱼;
  3. 如果以上都不行,换方案三的Camelot工具试试,说不定会有惊喜。

内容来源于stack exchange

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.08 08:52:58