You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用camelot-py lattice模式提取PDF表格时split_text参数无效问题

Camelot lattice模式split_text参数失效问题排查

问题现象

使用camelot提取带边框的PDF表格时,表格线可以被正确识别,但相邻两列的文本会被合并为同一个单元格内容。使用lattice模式并设置split_text = True后参数仍未生效。

故障代码示例

# -*- coding: utf-8 -*-

from pdfminer.layout import LAParams
from pdfminer.high_level import extract_text
import camelot

file = "test.pdf"

        
laparams = LAParams(
                line_overlap=0.5,
                char_margin=0.5,        # 尝试调小该参数,默认值为5
                word_margin=0.1,
                line_margin=0.0,
                boxes_flow=0.5,
                detect_vertical=False,
                all_texts=False
            )

# 提取表格
tables = camelot.read_pdf(
             file, 
             flavor='lattice', 
             pages="1", 
             process_background=False, 
             line_tol=2,
             joint_tol=2,
             line_scale=30,           # 从默认15调大以识别更细的表格线
             layout_params = laparams,
             split_text = True                                         
        )

# 网格识别结果正常
camelot.plot(tables[0], kind='grid').show()
# 文本没有按照网格拆分,特定文本如'Requirement/Function/Configuration'和'GxP'被合并
camelot.plot(tables[0], kind='text').show()


# 直接使用pdfminer时char_margin参数生效,但传入camelot.read_pdf后没有效果
texts = extract_text(file, page_numbers=[0], maxpages=1, laparams=laparams) 
texts = texts.split('\n')
print(texts)

故障代码输出的可视化结果如下,可见表格列边界识别完全正常,但文本跨列展示未被拆分:

  • 表格网格可视化图
    表格网格可视化图
  • 表格文本可视化图
    表格文本可视化图

修复后可正常运行的代码

仅调整了参数的传递格式,split_text就可以正常生效:

import camelot

file = "test.pdf"
    
laparams = {
        'line_overlap': 0.5,
        'char_margin': 0.5,
        'word_margin': 0.1,
        'line_margin': 0.0,
        'boxes_flow': 0.5,
        'detect_vertical': False,
        'all_texts': False
    }


camelotArgs = {
            'flavor': 'lattice', 
            'process_background': False, 
            'line_tol': 2,
            'joint_tol': 2,
            'line_scale': 30,           # 从默认15调大以识别更细的表格线
            'split_text': True,
            'layout_kwargs': laparams
        }

# 提取表格
tables = camelot.read_pdf(
         file, 
         pages="1", 
         **camelotArgs
    )

# 查看结果
camelot.plot(tables[0], kind='grid').show()
camelot.plot(tables[0], kind='text').show()

版本信息

  • Python版本:3.7.11
  • camelot-py版本:0.10.1
  • pdfminer.six版本:20211012

问题原因

核心原因是camelot的read_pdf方法不支持直接传入LAParams实例作为参数,自定义的pdfminer布局参数必须通过layout_kwargs参数以字典形式传入,camelot内部会自动根据该字典生成对应的LAParams实例。

故障代码中传入的layout_params不是camelot官方定义的入参字段,属于无效参数会被直接忽略,程序实际运行时使用了默认的char_margin=5的布局参数,字符间距判断阈值过大,导致相邻列的文本被识别为同一个文本块,即便开启split_text也无法按照网格拆分跨列的文本。

修复后的代码使用了官方要求的layout_kwargs传参,自定义的char_margin=0.5生效,pdfminer会将相邻列的文本识别为独立块,split_text即可正常按照表格网格拆分文本。


内容的提问来源于stack exchange,提问作者Tomper

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.30 12:39:03