You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中基于列名分类PDF混合表格的代码实现求助

解决PDF多结构表格归类问题的代码方案

问题分析

原代码存在两个核心缺陷:

  • 仅与单个DataFrame对比列结构,无法匹配已存在的其他类型表格
  • 无法动态创建新DataFrame来存储不同列结构的数据

解决方案代码

import tabula
import pandas as pd

def group_tables_by_columns(pdf_path, target_pages="all"):
    # 存储不同列结构的表格:每个元素是(列名元组, 对应DataFrame)
    table_groups = []
    
    # 读取指定页面的所有表格(假设每页一个表格,若一页多表可设multiple_tables=True)
    all_page_tables = tabula.read_pdf(pdf_path, pages=target_pages, multiple_tables=False)
    
    for page_num, new_table in enumerate(all_page_tables, start=1):
        # 清理列名(去除首尾空格,避免因格式差异导致匹配失败)
        new_table.columns = [col.strip() for col in new_table.columns]
        new_cols = tuple(new_table.columns)
        is_matched = False
        
        # 遍历已有表格组,检查列结构是否匹配
        for idx, (existing_cols, existing_table) in enumerate(table_groups):
            if new_cols == existing_cols:
                # 列结构匹配,追加数据
                updated_table = pd.concat([existing_table, new_table], axis=0, ignore_index=True)
                table_groups[idx] = (existing_cols, updated_table)
                is_matched = True
                break
        
        if not is_matched:
            # 无匹配结构,新增表格组
            table_groups.append((new_cols, new_table))
            print(f"页面 {page_num} 为新列结构表格,已新增分组")
    
    return table_groups

# 调用示例
pdf_file = "your_target.pdf"
result_groups = group_tables_by_columns(pdf_file)

# 查看结果:打印每个分组的列名和数据量
for cols, df in result_groups:
    print(f"列结构:{cols} | 数据行数:{len(df)}")

关键说明

  • 用table_groups列表存储不同列结构的表格,每个元素以列名元组作为标识,确保列结构对比准确
  • 自动清理列名的首尾空格,解决PDF表格列名带空格导致的匹配误差
  • 遍历所有已有表格组进行对比,确保新表格能匹配到正确的分组
  • 可通过修改target_pages参数指定读取特定页面(比如pages="1-10,106"),适配原代码的页面选择需求

内容的提问来源于stack exchange,提问作者deboshree roy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 21:33:17