Python中基于列名分类PDF混合表格的代码实现求助
解决PDF多结构表格归类问题的代码方案
问题分析
原代码存在两个核心缺陷:
- 仅与单个DataFrame对比列结构,无法匹配已存在的其他类型表格
- 无法动态创建新DataFrame来存储不同列结构的数据
解决方案代码
import tabula import pandas as pd def group_tables_by_columns(pdf_path, target_pages="all"): # 存储不同列结构的表格:每个元素是(列名元组, 对应DataFrame) table_groups = [] # 读取指定页面的所有表格(假设每页一个表格,若一页多表可设multiple_tables=True) all_page_tables = tabula.read_pdf(pdf_path, pages=target_pages, multiple_tables=False) for page_num, new_table in enumerate(all_page_tables, start=1): # 清理列名(去除首尾空格,避免因格式差异导致匹配失败) new_table.columns = [col.strip() for col in new_table.columns] new_cols = tuple(new_table.columns) is_matched = False # 遍历已有表格组,检查列结构是否匹配 for idx, (existing_cols, existing_table) in enumerate(table_groups): if new_cols == existing_cols: # 列结构匹配,追加数据 updated_table = pd.concat([existing_table, new_table], axis=0, ignore_index=True) table_groups[idx] = (existing_cols, updated_table) is_matched = True break if not is_matched: # 无匹配结构,新增表格组 table_groups.append((new_cols, new_table)) print(f"页面 {page_num} 为新列结构表格,已新增分组") return table_groups # 调用示例 pdf_file = "your_target.pdf" result_groups = group_tables_by_columns(pdf_file) # 查看结果:打印每个分组的列名和数据量 for cols, df in result_groups: print(f"列结构:{cols} | 数据行数:{len(df)}")
关键说明
- 用
table_groups列表存储不同列结构的表格,每个元素以列名元组作为标识,确保列结构对比准确 - 自动清理列名的首尾空格,解决PDF表格列名带空格导致的匹配误差
- 遍历所有已有表格组进行对比,确保新表格能匹配到正确的分组
- 可通过修改
target_pages参数指定读取特定页面(比如pages="1-10,106"),适配原代码的页面选择需求
内容的提问来源于stack exchange,提问作者deboshree roy
相关产品推荐
相关产品推荐

