如何用Tabula匹配PDF关键词对应表格并生成DataFrame字典
问题描述
我已经用PyPDF2提取了PDF文本,手里有一组大小写敏感的关键词列表(示例:["id","VIS","INT"])。每个关键词在文本中出现后,后续可能紧跟表格。需要用Tabula工具提取每个关键词对应的表格,将关键词作为键、对应表格转换后的DataFrame作为值存入字典。当前代码无法正确关联关键词与对应表格,求实现该逻辑。
原文本示例
id: Res ID Res ID Based upon 3,302 valid cases out of 3,30 total . •Mean: 54361.66 •Minimum: 10005.00 •Maximum: 99992.00 •Standard Deviation: 25723.74 Location: 1-5 (width: 5; decimal: 0) Variable Type: numeric VIS: Study Swan visit Value Label Unweighted Frequency% 00- 3302 100.0 % Total 3,302 100% Based upon 3,302 valid cases out of 3,302 total cases. Location: 6-7 (width: 2; decimal: 0) Variable Type: character INT: Interview Date form Value Label Unweighted Frequency% 0- 3302 100.0 % Total 3,302 100% Based upon 3,302 valid cases out of 3,302 total cases. •Mean: 0.00 •Median: 0.00 •Mode: 0.00 •Minimum: 0.00 •Maximum: 0.00 •Standard Deviation: 0.00 Location: 8-8 (width: 1; decimal: 0) Variable Type: numeric
原代码示例
import tabula lis= ["id","VIS","INT"] results_dict = {} for word in lis: if word in text: try: table = tabula.read_pdf("df.pdf", pages='all', multiple_tables=True) results_dict[word] = table except: pass print(results_dict)
预期输出
字典格式示例
dic = {'id':"no_table",'VIS':对应表格DataFrame,'INT':对应表格DataFrame}
DataFrame格式示例
data = {'Value': ["0.0", "NaN"], 'Label': ["-", "Total"], 'Unweighted\rFrequency': ["3302", "3,302"], '%': ["100.0 %", "100%"], } df = pd.DataFrame(data)
解决方案
核心思路是先定位关键词在PDF中的页码,再针对对应页码提取表格并匹配关联,避免无差别提取所有表格。以下是可运行的实现代码:
import PyPDF2 import tabula import pandas as pd def get_keyword_pages(pdf_path, keywords): # 记录每个关键词出现的页码(Tabula用1-based页码) page_map = {kw: [] for kw in keywords} with open(pdf_path, 'rb') as f: reader = PyPDF2.PdfReader(f) for page_idx in range(len(reader.pages)): page_text = reader.pages[page_idx].extract_text() for kw in keywords: if kw in page_text: page_map[kw].append(page_idx + 1) return page_map def extract_keyword_tables(pdf_path, keyword_pages): results = {} for kw, pages in keyword_pages.items(): if not pages: results[kw] = "no_table" continue # 优先取关键词首次出现的页码 target_page = pages[0] # 提取当前页的所有表格 current_page_tables = tabula.read_pdf(pdf_path, pages=str(target_page), multiple_tables=True) matched_table = None # 匹配符合特征的表格(根据示例表头特征判断) for table in current_page_tables: if any(col in ['Value Label', 'Unweighted', 'Frequency%'] for col in table.columns): # 整理表格列名和内容 table.columns = ['Value', 'Label', 'Unweighted\rFrequency', '%'] table = table.dropna(how='all').reset_index(drop=True) matched_table = table break # 当前页没找到则检查下一页 if not matched_table: next_page_tables = tabula.read_pdf(pdf_path, pages=str(target_page + 1), multiple_tables=True) for table in next_page_tables: if any(col in ['Value Label', 'Unweighted', 'Frequency%'] for col in table.columns): table.columns = ['Value', 'Label', 'Unweighted\rFrequency', '%'] table = table.dropna(how='all').reset_index(drop=True) matched_table = table break results[kw] = matched_table if matched_table else "no_table" return results # 调用示例 pdf_path = "df.pdf" keywords = ["id", "VIS", "INT"] keyword_pages = get_keyword_pages(pdf_path, keywords) results_dict = extract_keyword_tables(pdf_path, keyword_pages) # 输出结果 for kw, content in results_dict.items(): print(f"=== 关键词: {kw} ===") if isinstance(content, pd.DataFrame): print(content) else: print(content)
代码说明
get_keyword_pages函数:遍历PDF所有页面,提取每页文本并记录每个关键词出现的页码,确保后续只在目标页码附近提取表格。extract_keyword_tables函数:- 对每个关键词,先在其出现的页码提取表格,通过表头特征(如
Value Label)判断是否为目标表格; - 若当前页未找到,自动检查下一页(处理关键词在页尾的情况);
- 找到表格后整理列名、清理空行,未找到则赋值为
"no_table"。
- 对每个关键词,先在其出现的页码提取表格,通过表头特征(如
- 解决了原代码中所有关键词对应全部表格的问题,实现了关键词与对应表格的精准关联。
内容的提问来源于stack exchange,提问作者ella
相关产品推荐
相关产品推荐

