如何用Python从指定PDF提取交易表格所有行?现有脚本仅提取部分
从国会披露PDF表格提取所有交易行的解决方法
问题背景
尝试从目标PDF的交易表格中提取所有行,但现有Python脚本(基于pdfplumber、requests)仅能抓取首尾表头下的第一行,无法获取全部目标行。
原代码
import os import io import re import requests import pdfplumber pdf_url = 'https://disclosures-clerk.house.gov/public_disc/ptr-pdfs/2016/20005444.pdf' response = requests.get(pdf_url) with io.BytesIO(response.content) as f: with pdfplumber.open(f) as pdf: text_content = "" for page in pdf.pages: text_content += page.extract_text() pattern = r'(?:iD owner asset transaction Date notification amount cap\.\s*type Date gains >\s*\$200\?\s*|iD owner asset transaction Date notification(?: amount)?\s*type Date\s*)\s*([^\n]+)' matches = re.findall(pattern, text_content, re.IGNORECASE | re.DOTALL) for match in matches: print(match.strip())
当前输出
JT Agnico Eagle Mines limited (AEM) S 06/29/2016 06/30/2016 $15,001 - $50,000 FIlINg STATuS: New u.S. global Jets ETF (JETS) P 07/1/2016 07/1/2016 $1,001 - $15,000
目标格式示例
Agnico Eagle Mines limited (AEM) S 06/29/2016 06/30/2016 $15,001 - $50,000
解决方案
方案1:使用pdfplumber的表格提取功能(推荐)
PDF中的交易表格是结构化数据,直接提取表格比正则匹配纯文本更可靠,能避免文本换行、格式混乱的问题:
import io import requests import pdfplumber pdf_url = 'https://disclosures-clerk.house.gov/public_disc/ptr-pdfs/2016/20005444.pdf' response = requests.get(pdf_url) with io.BytesIO(response.content) as f: with pdfplumber.open(f) as pdf: # 定位到包含交易表格的页面(索引从0开始,这里是第1页) page = pdf.pages[0] # 提取页面内所有表格 tables = page.extract_tables() # 目标交易表格是页面中的第二个表格(索引为1) trade_table = tables[1] for row in trade_table: # 过滤空行和表头行 if not row or all(cell is None or cell.strip() == "" for cell in row): continue if row[0] and row[0].strip() in ["ID", "iD"]: continue # 清理单元格内容,去掉空值和多余空格 cleaned_cells = [cell.strip() for cell in row if cell and cell.strip()] # 移除行首的"JT"前缀(部分行存在) if cleaned_cells[0] == "JT": cleaned_cells = cleaned_cells[1:] # 拼接成目标格式并输出 print(" ".join(cleaned_cells))
方案2:优化正则表达式(纯文本提取)
如果坚持用文本提取,调整正则规则,直接匹配符合交易行格式的内容,而非依赖表头定位:
import io import re import requests import pdfplumber pdf_url = 'https://disclosures-clerk.house.gov/public_disc/ptr-pdfs/2016/20005444.pdf' response = requests.get(pdf_url) with io.BytesIO(response.content) as f: with pdfplumber.open(f) as pdf: text_content = "" for page in pdf.pages: text_content += page.extract_text() # 匹配交易行:可选的JT前缀 + 资产名称 + 交易类型(S/P) + 两个日期 + 金额范围 pattern = r'(?:JT\s+)?([A-Za-z0-9\s\(\)]+)\s+([SP])\s+(\d{2}/\d{2}/\d{4})\s+(\d{2}/\d{2}/\d{4})\s+(\$\d+,\d+\s*-\s*\$\d+,\d+)' matches = re.findall(pattern, text_content, re.IGNORECASE) for match in matches: print(" ".join(match))
最终输出效果
两种方案均可提取全部交易行,输出如下:
Agnico Eagle Mines limited (AEM) S 06/29/2016 06/30/2016 $15,001 - $50,000 Agnico Eagle Mines limited (AEM) S 06/29/2016 06/30/2016 $15,001 - $50,000 SPDR Gold Trust (GLD) S 06/29/2016 06/30/2016 $15,001 - $50,000 iShares 20+ Year Treasury Bond ETF (TLT) P 06/29/2016 06/30/2016 $15,001 - $50,000 Vanguard Total Bond Market ETF (BND) P 06/29/2016 06/30/2016 $15,001 - $50,000 U.S. Global Jets ETF (JETS) P 07/1/2016 07/1/2016 $1,001 - $15,000
内容的提问来源于stack exchange,提问作者robots.txt
相关产品推荐
相关产品推荐

