如何为Tabula提取的PDF表格分配其顶部的对应关键词?
实现方案:为Tabula提取的表格匹配顶部关键词
核心思路
利用PyPDF2(或pdfplumber)提取PDF第6页的文本信息,结合Tabula提取表格时的位置数据,定位每个表格上方对应的关键词,最后将关键词与表格关联(新增列或字典存储)。
具体步骤与代码示例
1. 提取页面文本和表格
先分别用PyPDF2提取页面纯文本,Tabula提取第6页的所有表格:
import PyPDF2 import tabula import pandas as pd # 提取第6页文本(PyPDF2页码从0开始,对应索引5) with open("your_document.pdf", "rb") as pdf_file: pdf_reader = PyPDF2.PdfReader(pdf_file) page_text = pdf_reader.pages[5].extract_text() # 提取第6页的表格(Tabula页码从1开始,返回JSON格式以获取位置信息) tables_json = tabula.read_pdf( "your_document.pdf", pages=6, multiple_tables=True, output_format="json" ) # 转换为DataFrame列表方便后续处理 df_list = [pd.DataFrame(tbl["data"], columns=tbl["columns"]) for tbl in tables_json]
2. 定义关键词并定位其在文本中的位置
先定义目标关键词列表,再从页面文本中筛选出关键词的出现顺序(或位置):
# 你的关键词列表 target_keywords = ["id", "type", "name", "code"] # 方案A:基于文本行顺序匹配(适合文本提取顺序与页面布局一致的情况) # 拆分文本为非空行 text_lines = [line.strip() for line in page_text.split("\n") if line.strip()] # 按页面从上到下的顺序提取唯一关键词 matched_keys = [] seen_keys = set() for line in text_lines: line_lower = line.lower() for key in target_keywords: if key.lower() in line_lower and key not in seen_keys: matched_keys.append(key) seen_keys.add(key) # 匹配到的关键词数量和表格数量一致就停止 if len(matched_keys) == len(df_list): break if len(matched_keys) == len(df_list): break # 方案B:基于坐标精准匹配(适合文本提取顺序混乱的情况,需用pdfplumber) # 如果PyPDF2提取的文本顺序和页面布局不符,换用pdfplumber获取带坐标的文本 import pdfplumber with pdfplumber.open("your_document.pdf") as pdf: page = pdf.pages[5] # 提取带坐标的文本对象 text_objs = page.extract_words(x_tolerance=2, y_tolerance=2) # 收集所有关键词的顶部坐标和内容 keyword_coords = [] for obj in text_objs: obj_text = obj["text"].lower() for key in target_keywords: if obj_text == key.lower(): keyword_coords.append((obj["top"], key)) # 按页面从上到下排序(top值越小越靠上) keyword_coords.sort() # 获取每个表格的顶部坐标,同样排序 table_tops = [tbl["top"] for tbl in tables_json] table_tops.sort() # 为每个表格匹配最近的上方关键词 matched_keys = [] for table_top in table_tops: # 筛选出表格上方的所有关键词,取最后一个(最近的) valid_keys = [key for (top, key) in keyword_coords if top <= table_top] matched_keys.append(valid_keys[-1] if valid_keys else None)
3. 将关键词与表格关联
你可以选择为每个表格新增一列存储关键词,或者用字典统一管理:
# 方式1:为每个DataFrame新增"table_keyword"列 for df, key in zip(df_list, matched_keys): df["table_keyword"] = key # 方式2:用字典存储表格与关键词的对应关系 table_key_mapping = { f"table_{idx+1}": { "keyword": key, "dataframe": df } for idx, (df, key) in enumerate(zip(df_list, matched_keys)) }
注意事项
- 如果关键词存在大小写差异,匹配时统一转成小写避免遗漏。
- 若页面中有重复关键词,优先取表格正上方最近的那个(方案B的坐标匹配更精准)。
- 若未匹配到关键词,可以设置默认值(如
None)或添加异常处理。
内容的提问来源于stack exchange,提问作者ella
相关产品推荐
相关产品推荐

