You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何为Tabula提取的PDF表格分配其顶部的对应关键词?

实现方案:为Tabula提取的表格匹配顶部关键词

核心思路

利用PyPDF2(或pdfplumber)提取PDF第6页的文本信息,结合Tabula提取表格时的位置数据,定位每个表格上方对应的关键词,最后将关键词与表格关联(新增列或字典存储)。

具体步骤与代码示例

1. 提取页面文本和表格

先分别用PyPDF2提取页面纯文本,Tabula提取第6页的所有表格:

import PyPDF2
import tabula
import pandas as pd

# 提取第6页文本(PyPDF2页码从0开始,对应索引5)
with open("your_document.pdf", "rb") as pdf_file:
    pdf_reader = PyPDF2.PdfReader(pdf_file)
    page_text = pdf_reader.pages[5].extract_text()

# 提取第6页的表格(Tabula页码从1开始,返回JSON格式以获取位置信息)
tables_json = tabula.read_pdf(
    "your_document.pdf",
    pages=6,
    multiple_tables=True,
    output_format="json"
)
# 转换为DataFrame列表方便后续处理
df_list = [pd.DataFrame(tbl["data"], columns=tbl["columns"]) for tbl in tables_json]

2. 定义关键词并定位其在文本中的位置

先定义目标关键词列表,再从页面文本中筛选出关键词的出现顺序(或位置):

# 你的关键词列表
target_keywords = ["id", "type", "name", "code"]

# 方案A:基于文本行顺序匹配(适合文本提取顺序与页面布局一致的情况)
# 拆分文本为非空行
text_lines = [line.strip() for line in page_text.split("\n") if line.strip()]
# 按页面从上到下的顺序提取唯一关键词
matched_keys = []
seen_keys = set()
for line in text_lines:
    line_lower = line.lower()
    for key in target_keywords:
        if key.lower() in line_lower and key not in seen_keys:
            matched_keys.append(key)
            seen_keys.add(key)
            # 匹配到的关键词数量和表格数量一致就停止
            if len(matched_keys) == len(df_list):
                break
    if len(matched_keys) == len(df_list):
        break

# 方案B:基于坐标精准匹配(适合文本提取顺序混乱的情况,需用pdfplumber)
# 如果PyPDF2提取的文本顺序和页面布局不符,换用pdfplumber获取带坐标的文本
import pdfplumber
with pdfplumber.open("your_document.pdf") as pdf:
    page = pdf.pages[5]
    # 提取带坐标的文本对象
    text_objs = page.extract_words(x_tolerance=2, y_tolerance=2)
    # 收集所有关键词的顶部坐标和内容
    keyword_coords = []
    for obj in text_objs:
        obj_text = obj["text"].lower()
        for key in target_keywords:
            if obj_text == key.lower():
                keyword_coords.append((obj["top"], key))
    # 按页面从上到下排序(top值越小越靠上)
    keyword_coords.sort()
    # 获取每个表格的顶部坐标,同样排序
    table_tops = [tbl["top"] for tbl in tables_json]
    table_tops.sort()
    # 为每个表格匹配最近的上方关键词
    matched_keys = []
    for table_top in table_tops:
        # 筛选出表格上方的所有关键词,取最后一个(最近的)
        valid_keys = [key for (top, key) in keyword_coords if top <= table_top]
        matched_keys.append(valid_keys[-1] if valid_keys else None)

3. 将关键词与表格关联

你可以选择为每个表格新增一列存储关键词,或者用字典统一管理:

# 方式1:为每个DataFrame新增"table_keyword"列
for df, key in zip(df_list, matched_keys):
    df["table_keyword"] = key

# 方式2:用字典存储表格与关键词的对应关系
table_key_mapping = {
    f"table_{idx+1}": {
        "keyword": key,
        "dataframe": df
    }
    for idx, (df, key) in enumerate(zip(df_list, matched_keys))
}

注意事项

  • 如果关键词存在大小写差异,匹配时统一转成小写避免遗漏。
  • 若页面中有重复关键词,优先取表格正上方最近的那个(方案B的坐标匹配更精准)。
  • 若未匹配到关键词,可以设置默认值(如None)或添加异常处理。

内容的提问来源于stack exchange,提问作者ella

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 15:05:13