You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Tabula匹配PDF关键词对应表格并生成DataFrame字典

问题描述

我已经用PyPDF2提取了PDF文本,手里有一组大小写敏感的关键词列表(示例:["id","VIS","INT"])。每个关键词在文本中出现后,后续可能紧跟表格。需要用Tabula工具提取每个关键词对应的表格,将关键词作为键、对应表格转换后的DataFrame作为值存入字典。当前代码无法正确关联关键词与对应表格,求实现该逻辑。


原文本示例

id: Res ID
Res ID
Based upon  3,302 valid cases out of 3,30
total .
•Mean:
54361.66
•Minimum: 10005.00
•Maximum: 99992.00
•Standard Deviation:
25723.74
Location: 1-5 (width: 5; decimal: 0)
Variable Type:  numeric 
VIS: Study 
Swan  visit 
Value Label
Unweighted
Frequency%
00- 3302 100.0 %
 Total 3,302 100%
Based
upon  3,302 valid cases out of 3,302 total cases.
Location: 6-7
(width: 2; decimal: 0)
Variable Type:  character 
INT: Interview
Date form 
Value Label Unweighted
Frequency%
0- 3302 100.0 %
 Total 3,302 100%
Based upon 3,302 valid  cases out of 3,302 total
cases.
•Mean:
0.00
•Median: 0.00
•Mode: 0.00
•Minimum:
0.00
•Maximum: 0.00
•Standard Deviation:
0.00
Location: 8-8 (width: 1; decimal: 0)
Variable Type:  numeric

原代码示例

import tabula
lis= ["id","VIS","INT"]
results_dict = {}
for word in lis:
    if word in text:
        try:
            table = tabula.read_pdf("df.pdf", pages='all', multiple_tables=True)
            results_dict[word] = table
        except:
            pass
print(results_dict)

预期输出

字典格式示例

dic = {'id':"no_table",'VIS':对应表格DataFrame,'INT':对应表格DataFrame}

DataFrame格式示例

data = {'Value':  ["0.0", "NaN"],
        'Label': ["-", "Total"],
        'Unweighted\rFrequency':  ["3302", "3,302"],
        '%': ["100.0 %", "100%"],
        }
df = pd.DataFrame(data)

解决方案

核心思路是先定位关键词在PDF中的页码,再针对对应页码提取表格并匹配关联,避免无差别提取所有表格。以下是可运行的实现代码:

import PyPDF2
import tabula
import pandas as pd

def get_keyword_pages(pdf_path, keywords):
    # 记录每个关键词出现的页码(Tabula用1-based页码)
    page_map = {kw: [] for kw in keywords}
    with open(pdf_path, 'rb') as f:
        reader = PyPDF2.PdfReader(f)
        for page_idx in range(len(reader.pages)):
            page_text = reader.pages[page_idx].extract_text()
            for kw in keywords:
                if kw in page_text:
                    page_map[kw].append(page_idx + 1)
    return page_map

def extract_keyword_tables(pdf_path, keyword_pages):
    results = {}
    for kw, pages in keyword_pages.items():
        if not pages:
            results[kw] = "no_table"
            continue
        
        # 优先取关键词首次出现的页码
        target_page = pages[0]
        # 提取当前页的所有表格
        current_page_tables = tabula.read_pdf(pdf_path, pages=str(target_page), multiple_tables=True)
        matched_table = None
        
        # 匹配符合特征的表格(根据示例表头特征判断)
        for table in current_page_tables:
            if any(col in ['Value Label', 'Unweighted', 'Frequency%'] for col in table.columns):
                # 整理表格列名和内容
                table.columns = ['Value', 'Label', 'Unweighted\rFrequency', '%']
                table = table.dropna(how='all').reset_index(drop=True)
                matched_table = table
                break
        
        # 当前页没找到则检查下一页
        if not matched_table:
            next_page_tables = tabula.read_pdf(pdf_path, pages=str(target_page + 1), multiple_tables=True)
            for table in next_page_tables:
                if any(col in ['Value Label', 'Unweighted', 'Frequency%'] for col in table.columns):
                    table.columns = ['Value', 'Label', 'Unweighted\rFrequency', '%']
                    table = table.dropna(how='all').reset_index(drop=True)
                    matched_table = table
                    break
        
        results[kw] = matched_table if matched_table else "no_table"
    return results

# 调用示例
pdf_path = "df.pdf"
keywords = ["id", "VIS", "INT"]
keyword_pages = get_keyword_pages(pdf_path, keywords)
results_dict = extract_keyword_tables(pdf_path, keyword_pages)

# 输出结果
for kw, content in results_dict.items():
    print(f"=== 关键词: {kw} ===")
    if isinstance(content, pd.DataFrame):
        print(content)
    else:
        print(content)

代码说明

  1. get_keyword_pages函数:遍历PDF所有页面,提取每页文本并记录每个关键词出现的页码,确保后续只在目标页码附近提取表格。
  2. extract_keyword_tables函数:
    • 对每个关键词,先在其出现的页码提取表格,通过表头特征(如Value Label)判断是否为目标表格;
    • 若当前页未找到,自动检查下一页(处理关键词在页尾的情况);
    • 找到表格后整理列名、清理空行,未找到则赋值为"no_table"。
  3. 解决了原代码中所有关键词对应全部表格的问题,实现了关键词与对应表格的精准关联。

内容的提问来源于stack exchange,提问作者ella

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.22 16:57:04