You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Tabula-Py抽取PDF数据时列值不匹配问题求助

解决PDF数据抽取列值不匹配问题(针对BOE车辆价格PDF)

问题背景

处理两份BOE发布的车辆价格PDF时,原代码仅适配其中一份文件,因两份PDF格式存在细微差异(表头重复位置、冗余行数量、分页表格偏移),导致抽取后列值错位,无法生成正确的DataFrame和CSV。

解决方案:差异化适配的通用抽取函数

以下代码通过参数化配置,可分别适配两个PDF的格式特征,解决列值不匹配问题:

1. 导入依赖库

import pandas as pd
import tabula

2. 通用PDF抽取清洗函数

def extract_boe_pdf(pdf_path, start_page, header_row_idx, skip_rows_after_header, redundant_row_markers=None):
    # 读取PDF表格,lattice模式适配带边框的表格
    dfs_list = tabula.read_pdf(
        pdf_path,
        pages=f"{start_page}-end",
        lattice=True,
        pandas_options={'header': None},
        multiple_tables=True
    )
    
    df_combined = pd.DataFrame()
    current_page = start_page
    
    for df in dfs_list:
        # 移除全空列,避免列数混乱
        df = df.dropna(axis=1, how='all')
        if df.empty:
            current_page += 1
            continue
        
        # 过滤重复表头或冗余标识行
        if redundant_row_markers:
            for marker in redundant_row_markers:
                df = df[~df.apply(lambda row: row.eq(marker).any(), axis=1)]
        
        # 添加页码标记,方便排查异常页
        df['page'] = f"Page: {current_page}"
        
        # 统一临时列名,避免concat时列不匹配报错
        df.columns = range(1, len(df.columns) + 1)
        df_combined = pd.concat([df_combined, df], ignore_index=True)
        
        current_page += 1
    
    # 设置正式表头
    header = df_combined.iloc[header_row_idx].astype(str).replace('nan', '')
    df_combined.columns = header.tolist()
    
    # 跳过表头及前置冗余行
    df_combined = df_combined.iloc[header_row_idx + skip_rows_after_header:].reset_index(drop=True)
    
    # 清理全空行
    df_combined = df_combined.dropna(how='all')
    
    return df_combined

3. 分别处理两个PDF

处理2014年PDF

# 适配特征:表头在第2行(索引1),表头后需跳过3行冗余内容,重复表头标记为"MARCA"
df_2014 = extract_boe_pdf(
    pdf_path="BOE-A-2014-13181.pdf",
    start_page=4,
    header_row_idx=1,
    skip_rows_after_header=3,
    redundant_row_markers=["MARCA"]
)

# 导出为CSV
df_2014.to_csv("boe_cars_2014.csv", index=False, encoding='utf-8-sig')

处理2016年PDF

# 适配特征:表头在第2行(索引1),表头后需跳过4行冗余内容,重复表头标记为"MARCA"、"VALOR"
df_2016 = extract_boe_pdf(
    pdf_path="eli-es-o-2016-12-14-hfp1895-dof-spa.pdf",
    start_page=4,
    header_row_idx=1,
    skip_rows_after_header=4,
    redundant_row_markers=["MARCA", "VALOR"]
)

# 导出为CSV
df_2016.to_csv("boe_cars_2016.csv", index=False, encoding='utf-8-sig')

关键调整说明

  • 冗余行过滤:通过redundant_row_markers参数批量移除跨页重复的表头行(如"MARCA""VALOR")
  • 表头适配:用header_row_idx和skip_rows_after_header灵活对应两个PDF的表头位置和前置冗余行数
  • 列统一处理:临时用数字序列命名列,解决不同页表格列名不一致导致的拼接失败问题
  • 数据清洗:最后清理全空行,保证输出数据的整洁性

内容的提问来源于stack exchange,提问作者Michael Picazo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.05 02:42:49