如何让Camelot仅提取PDF表格,PyPDF2提取表格外文本?
解决Camelot误识别文本为表格及PyPDF2提取非表格文本的问题
一、限制Camelot仅提取真实表格
Camelot误将单行文本识别为单列表格,核心是优化其表格检测逻辑,可通过以下方式处理:
1. 精准指定表格区域
若已知PDF中表格的大致位置,直接用table_areas参数限定提取范围,避免扫描整个页面。坐标格式为[x1, y1, x2, y2](x从左到右,y从下到上,单位为点):
import camelot # 限定表格在页面x1=50、y1=100、x2=750、y2=500的区域内 tables = camelot.read_pdf("your_file.pdf", table_areas=["50,100,750,500"], flavor="lattice")
2. 过滤小表格剔除误判
单列表格通常行数极少,可手动筛选出行数≥3的表格,或调整边缘检测参数降低敏感度:
# 提取后过滤行数不足3的表格 tables = camelot.read_pdf("your_file.pdf", flavor="stream") valid_tables = [table for table in tables if table.shape[0] >= 3] # 调整edge_tol参数增强边框检测精度 tables = camelot.read_pdf("your_file.pdf", flavor="lattice", edge_tol=50)
3. 选择匹配的提取模式
- 优先用
lattice模式提取有清晰边框的表格,能自动排除无框的普通文本; - 若用
stream模式,可通过columns参数指定列分隔线的x坐标,避免将单行文本识别为单列:
# 指定表格列分隔在x=200和x=400的位置 tables = camelot.read_pdf("your_file.pdf", flavor="stream", columns=["200,400"])
二、让PyPDF2仅提取表格外文本
需先通过Camelot获取表格的页面坐标,再从PyPDF2提取的文本中剔除表格区域内的内容:
1. 获取表格的坐标范围
Camelot的每个表格对象自带bbox属性,对应表格的[x1, y1, x2, y2]坐标:
valid_tables = camelot.read_pdf("your_file.pdf", ...) # 收集当前页面所有有效表格的坐标(多页需循环处理) table_bboxes = [table.bbox for table in valid_tables]
2. 提取并过滤表格区域内的文本
利用PyPDF2的get_text_words()方法获取每个文本块的坐标,判断其是否在表格区域外:
from PyPDF2 import PdfReader reader = PdfReader("your_file.pdf") page = reader.pages[0] text_words = page.get_text_words() non_table_text = [] for word in text_words: x0, y0, x1, y1, text = word[:5] inside_table = False # 检查文本块是否与任一表格区域重叠 for bbox in table_bboxes: tb_x1, tb_y1, tb_x2, tb_y2 = bbox if x0 >= tb_x1 and x1 <= tb_x2 and y0 >= tb_y1 and y1 <= tb_y2: inside_table = True break if not inside_table: non_table_text.append(text) # 拼接成最终的非表格文本 final_text = " ".join(non_table_text)
注:Camelot与PyPDF2的坐标系统一致,均以页面左下角为原点(x向右、y向上),无需额外转换。
内容的提问来源于stack exchange,提问作者Varad Kulkarni
相关产品推荐
相关产品推荐

