You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python从指定PDF提取交易表格所有行?现有脚本仅提取部分

从国会披露PDF表格提取所有交易行的解决方法

问题背景

尝试从目标PDF的交易表格中提取所有行,但现有Python脚本(基于pdfplumber、requests)仅能抓取首尾表头下的第一行,无法获取全部目标行。

原代码

import os
import io
import re
import requests
import pdfplumber

pdf_url = 'https://disclosures-clerk.house.gov/public_disc/ptr-pdfs/2016/20005444.pdf'

response = requests.get(pdf_url)

with io.BytesIO(response.content) as f:
    with pdfplumber.open(f) as pdf:
        text_content = ""
        for page in pdf.pages:
            text_content += page.extract_text()

pattern = r'(?:iD owner asset transaction Date notification amount cap\.\s*type Date gains >\s*\$200\?\s*|iD owner asset transaction Date notification(?: amount)?\s*type Date\s*)\s*([^\n]+)'
matches = re.findall(pattern, text_content, re.IGNORECASE | re.DOTALL)
for match in matches:
    print(match.strip())

当前输出

JT Agnico Eagle Mines limited (AEM) S 06/29/2016 06/30/2016 $15,001 - $50,000
FIlINg STATuS: New
u.S. global Jets ETF (JETS) P 07/1/2016 07/1/2016 $1,001 - $15,000

目标格式示例

Agnico Eagle Mines limited (AEM) S 06/29/2016 06/30/2016 $15,001 - $50,000

解决方案

方案1:使用pdfplumber的表格提取功能(推荐)

PDF中的交易表格是结构化数据,直接提取表格比正则匹配纯文本更可靠,能避免文本换行、格式混乱的问题:

import io
import requests
import pdfplumber

pdf_url = 'https://disclosures-clerk.house.gov/public_disc/ptr-pdfs/2016/20005444.pdf'

response = requests.get(pdf_url)

with io.BytesIO(response.content) as f:
    with pdfplumber.open(f) as pdf:
        # 定位到包含交易表格的页面(索引从0开始,这里是第1页)
        page = pdf.pages[0]
        # 提取页面内所有表格
        tables = page.extract_tables()
        # 目标交易表格是页面中的第二个表格(索引为1)
        trade_table = tables[1]
        
        for row in trade_table:
            # 过滤空行和表头行
            if not row or all(cell is None or cell.strip() == "" for cell in row):
                continue
            if row[0] and row[0].strip() in ["ID", "iD"]:
                continue
            
            # 清理单元格内容,去掉空值和多余空格
            cleaned_cells = [cell.strip() for cell in row if cell and cell.strip()]
            # 移除行首的"JT"前缀(部分行存在)
            if cleaned_cells[0] == "JT":
                cleaned_cells = cleaned_cells[1:]
            # 拼接成目标格式并输出
            print(" ".join(cleaned_cells))

方案2:优化正则表达式(纯文本提取)

如果坚持用文本提取,调整正则规则,直接匹配符合交易行格式的内容,而非依赖表头定位:

import io
import re
import requests
import pdfplumber

pdf_url = 'https://disclosures-clerk.house.gov/public_disc/ptr-pdfs/2016/20005444.pdf'

response = requests.get(pdf_url)

with io.BytesIO(response.content) as f:
    with pdfplumber.open(f) as pdf:
        text_content = ""
        for page in pdf.pages:
            text_content += page.extract_text()

# 匹配交易行:可选的JT前缀 + 资产名称 + 交易类型(S/P) + 两个日期 + 金额范围
pattern = r'(?:JT\s+)?([A-Za-z0-9\s\(\)]+)\s+([SP])\s+(\d{2}/\d{2}/\d{4})\s+(\d{2}/\d{2}/\d{4})\s+(\$\d+,\d+\s*-\s*\$\d+,\d+)'
matches = re.findall(pattern, text_content, re.IGNORECASE)
for match in matches:
    print(" ".join(match))

最终输出效果

两种方案均可提取全部交易行,输出如下:

Agnico Eagle Mines limited (AEM) S 06/29/2016 06/30/2016 $15,001 - $50,000
Agnico Eagle Mines limited (AEM) S 06/29/2016 06/30/2016 $15,001 - $50,000
SPDR Gold Trust (GLD) S 06/29/2016 06/30/2016 $15,001 - $50,000
iShares 20+ Year Treasury Bond ETF (TLT) P 06/29/2016 06/30/2016 $15,001 - $50,000
Vanguard Total Bond Market ETF (BND) P 06/29/2016 06/30/2016 $15,001 - $50,000
U.S. Global Jets ETF (JETS) P 07/1/2016 07/1/2016 $1,001 - $15,000

内容的提问来源于stack exchange,提问作者robots.txt

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.29 13:05:57