You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用tabula-py解析PDF时如何忽略单元格内换行符避免内容截断

解决方法

1. 使用内置参数(适配tabula-py 2.3.0及以上版本)

tabula-py 后续版本新增了remove_newline_in_cells参数,专门用于移除单元格内部的换行符,避免内容被拆分为多行,你可以直接在原有调用中新增该参数即可,调整后代码如下:

tabula.read_pdf(
    file, 
    pandas_options={"header":None}, 
    pages='all', 
    stream=True, 
    lattice=True, 
    multiple_tables=True, 
    guess=False, 
    password=password,
    remove_newline_in_cells=True
)

该参数会自动把单元格内的\r、\n类换行符替换为空格,原本换行的内容会合并为连续文本,不会出现内容遗漏。

2. 低版本手动后处理

如果你使用的tabula-py版本低于2.3.0,不支持上述参数,可以解析完成后对返回的DataFrame做批量处理,手动替换换行符:

# 按原有逻辑读取表格
dfs = tabula.read_pdf(
    file, 
    pandas_options={"header":None}, 
    pages='all', 
    stream=True, 
    lattice=True, 
    multiple_tables=True, 
    guess=False, 
    password=password
)

# 批量清理所有表格的单元格换行
processed_dfs = []
for df in dfs:
    df = df.apply(lambda col: col.astype(str).str.replace(r'[\r\n]+', ' ', regex=True).str.strip())
    processed_dfs.append(df)

额外优化提示

你当前同时开启了stream和lattice两种解析模式,前者适配无框线表格,后者适配有明确框线的表格,同时开启可能出现解析冲突,你可以根据自己PDF的表格格式,只保留对应模式的参数即可。


内容的提问来源于stack exchange,提问作者shekwo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.27 11:36:06