使用Camelot提取PDF表格时跨单元格文本复制异常求助
Camelot提取PDF跨单元格数据时的复制缺失问题
我在使用Camelot工具提取PDF数据时,遇到了跨单元格文本复制不一致的问题,具体出现在某数据手册的第3页。
问题详情
表格的跨单元格已被正确识别,但提取时出现数据缺失:
- 第3列内容仅复制到两个跨单元格中的一个
- 第4列内容仅复制到三个跨单元格中的两个
每列均缺失一个单元格数据。
相关截图
存在问题的原表格:

已正确识别的表格网格:

提取后缺失数据的结果:

使用的测试代码
table_areas=['86, 697, 529, 95'] # 忽略页面边框 tables = camelot.read_pdf(single_source, pages='all', flavor = 'lattice', copy_text=['v'], line_scale = 110, table_regions=table_areas, flag_size = False, process_background=False)
Colab环境配置与提取代码
依赖安装
!pip install "camelot-py[cv]" -q !pip install PyPDF2==2.12.1 !apt-get install ghostscript
导入模块与数据提取
import camelot import pandas as pd from tabulate import tabulate import re import fitz single_source = '/content/FDB9406_F085-D.PDF' print("Extracting ", single_source, "...") table_areas=['86, 697, 529, 95'] tables = camelot.read_pdf(single_source, pages='all', flavor = 'lattice', copy_text=['v'], line_scale = 110, table_regions=table_areas, flag_size = False, process_background=False) print("Extracting ", single_source, "is finished!")
表格可视化代码
for table in accurate_tables: print(table.parsing_report, table.shape, table._bbox) print(tabulate(table.df, headers='keys', tablefmt='psql')) camelot.plot(table, kind='grid').show() print("Extracting ", single_source, "is finished!")
内容的提问来源于stack exchange,提问作者Said Akyuz
相关产品推荐
相关产品推荐

