使用Tabula-Py抽取PDF数据时列值不匹配问题求助
解决PDF数据抽取列值不匹配问题(针对BOE车辆价格PDF)
问题背景
处理两份BOE发布的车辆价格PDF时,原代码仅适配其中一份文件,因两份PDF格式存在细微差异(表头重复位置、冗余行数量、分页表格偏移),导致抽取后列值错位,无法生成正确的DataFrame和CSV。
解决方案:差异化适配的通用抽取函数
以下代码通过参数化配置,可分别适配两个PDF的格式特征,解决列值不匹配问题:
1. 导入依赖库
import pandas as pd import tabula
2. 通用PDF抽取清洗函数
def extract_boe_pdf(pdf_path, start_page, header_row_idx, skip_rows_after_header, redundant_row_markers=None): # 读取PDF表格,lattice模式适配带边框的表格 dfs_list = tabula.read_pdf( pdf_path, pages=f"{start_page}-end", lattice=True, pandas_options={'header': None}, multiple_tables=True ) df_combined = pd.DataFrame() current_page = start_page for df in dfs_list: # 移除全空列,避免列数混乱 df = df.dropna(axis=1, how='all') if df.empty: current_page += 1 continue # 过滤重复表头或冗余标识行 if redundant_row_markers: for marker in redundant_row_markers: df = df[~df.apply(lambda row: row.eq(marker).any(), axis=1)] # 添加页码标记,方便排查异常页 df['page'] = f"Page: {current_page}" # 统一临时列名,避免concat时列不匹配报错 df.columns = range(1, len(df.columns) + 1) df_combined = pd.concat([df_combined, df], ignore_index=True) current_page += 1 # 设置正式表头 header = df_combined.iloc[header_row_idx].astype(str).replace('nan', '') df_combined.columns = header.tolist() # 跳过表头及前置冗余行 df_combined = df_combined.iloc[header_row_idx + skip_rows_after_header:].reset_index(drop=True) # 清理全空行 df_combined = df_combined.dropna(how='all') return df_combined
3. 分别处理两个PDF
处理2014年PDF
# 适配特征:表头在第2行(索引1),表头后需跳过3行冗余内容,重复表头标记为"MARCA" df_2014 = extract_boe_pdf( pdf_path="BOE-A-2014-13181.pdf", start_page=4, header_row_idx=1, skip_rows_after_header=3, redundant_row_markers=["MARCA"] ) # 导出为CSV df_2014.to_csv("boe_cars_2014.csv", index=False, encoding='utf-8-sig')
处理2016年PDF
# 适配特征:表头在第2行(索引1),表头后需跳过4行冗余内容,重复表头标记为"MARCA"、"VALOR" df_2016 = extract_boe_pdf( pdf_path="eli-es-o-2016-12-14-hfp1895-dof-spa.pdf", start_page=4, header_row_idx=1, skip_rows_after_header=4, redundant_row_markers=["MARCA", "VALOR"] ) # 导出为CSV df_2016.to_csv("boe_cars_2016.csv", index=False, encoding='utf-8-sig')
关键调整说明
- 冗余行过滤:通过
redundant_row_markers参数批量移除跨页重复的表头行(如"MARCA""VALOR") - 表头适配:用
header_row_idx和skip_rows_after_header灵活对应两个PDF的表头位置和前置冗余行数 - 列统一处理:临时用数字序列命名列,解决不同页表格列名不一致导致的拼接失败问题
- 数据清洗:最后清理全空行,保证输出数据的整洁性
内容的提问来源于stack exchange,提问作者Michael Picazo
相关产品推荐
相关产品推荐

