You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将指定PDF转换为DataFrame?解决tabula-py返回NaN问题

解决tabula-py转换PDF到DataFrame时空白行生成NaN的方法

方法1:读取阶段跳过空白行

调用read_pdf时启用skip_blank_lines参数,直接跳过PDF中的空白行,避免生成全NaN的行:

from tabula import read_pdf

# 读取指定PDF文件
df_2018 = read_pdf("2018.pdf", pages="all", skip_blank_lines=True)
df_2021 = read_pdf("2021.pdf", pages="all", skip_blank_lines=True)

方法2:读取后清理全NaN行

如果读取后仍残留全NaN行,用pandas的dropna方法删除所有值均为NaN的行:

# 清理单个DataFrame并重置索引
df_2018_cleaned = df_2018.dropna(how="all").reset_index(drop=True)
df_2021_cleaned = df_2021.dropna(how="all").reset_index(drop=True)

方法3:指定表格读取区域

若PDF中空白区域被误识别为表格行,可通过area参数限定读取的坐标范围(坐标可通过tabula.interactive()启动GUI工具获取):

# 示例:替换为实际PDF的表格坐标(格式:[顶部, 左侧, 底部, 右侧])
df_2018 = read_pdf(
    "2018.pdf",
    pages="all",
    area=[50, 20, 700, 800],
    skip_blank_lines=True
)

方法4:精准指定列分隔

关闭自动表格识别(guess=False),配合columns参数手动指定列分隔的横坐标,减少空白行误判:

df_2018 = read_pdf(
    "2018.pdf",
    pages="all",
    guess=False,
    columns=[100, 250, 400, 550],  # 替换为实际列分隔的横坐标
    skip_blank_lines=True
)

内容的提问来源于stack exchange,提问作者emiley mille

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 14:14:54