You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Camelot提取PDF表格时跨单元格文本复制异常求助

Camelot提取PDF跨单元格数据时的复制缺失问题

我在使用Camelot工具提取PDF数据时,遇到了跨单元格文本复制不一致的问题,具体出现在某数据手册的第3页。

问题详情

表格的跨单元格已被正确识别,但提取时出现数据缺失:

  • 第3列内容仅复制到两个跨单元格中的一个
  • 第4列内容仅复制到三个跨单元格中的两个
    每列均缺失一个单元格数据。

相关截图

  • 存在问题的原表格:
    存在问题的表格

  • 已正确识别的表格网格:
    表格网格

  • 提取后缺失数据的结果:
    提取到的表格数据

使用的测试代码

table_areas=['86, 697, 529, 95'] # 忽略页面边框
tables = camelot.read_pdf(single_source, pages='all', 
                          flavor = 'lattice', 
                          copy_text=['v'], 
                          line_scale = 110, 
                          table_regions=table_areas, 
                          flag_size = False, 
                          process_background=False)

Colab环境配置与提取代码

依赖安装

!pip install "camelot-py[cv]" -q
!pip install PyPDF2==2.12.1
!apt-get install ghostscript

导入模块与数据提取

import camelot
import pandas as pd
from tabulate import tabulate
import re
import fitz

single_source = '/content/FDB9406_F085-D.PDF'
print("Extracting ", single_source, "...")

table_areas=['86, 697, 529, 95']
tables = camelot.read_pdf(single_source, pages='all', flavor = 'lattice', copy_text=['v'], line_scale = 110, table_regions=table_areas, flag_size = False, process_background=False)

print("Extracting ", single_source, "is finished!")

表格可视化代码

for table in accurate_tables:
  print(table.parsing_report, table.shape, table._bbox)
  print(tabulate(table.df, headers='keys', tablefmt='psql'))
  camelot.plot(table, kind='grid').show()

print("Extracting ", single_source, "is finished!")

内容的提问来源于stack exchange,提问作者Said Akyuz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 20:50:29