使用borb检测PDF表格时触发AttributeError问题求助
问题分析
你遇到的AttributeError是因为TableDetectionByLines.get_table_bounding_boxes()返回的是字典类型——键是页码(int类型),值是对应页面的表格矩形列表。你直接遍历这个字典时,拿到的是页码(int),而非表格的Rectangle对象,所以调用grow方法会报错。
解决方案
修正循环逻辑,先根据目标页码(比如第0页)取出该页的所有表格矩形,再遍历这些矩形对象:
from decimal import Decimal from borb.pdf import Document from borb.pdf.page.page import Page from borb.pdf.pdf import PDF from borb.toolkit.table.table_detection_by_lines import TableDetectionByLines as TDBL from borb.pdf.canvas.layout.annotation.square_annotation import SquareAnnotation from borb.pdf.canvas.color.color import HexColor, X11Color def main(infile, outfile): # 获取文档并检测表格 t = TDBL() doc = None with open(infile, "rb") as pdf_file_handle: doc = PDF.loads(pdf_file_handle, [t]) assert doc is not None # 获取第0页 p = doc.get_page(0) # 获取第0页的所有表格矩形 page_tables = t.get_table_bounding_boxes().get(0, []) for r in page_tables: r = r.grow(Decimal(5)) p.add_annotation(SquareAnnotation(r, stroke_color=X11Color("Green"))) # 保存修改后的文档 with open(outfile, "wb") as pdf_file_handle: PDF.dumps(pdf_file_handle, doc) return if __name__ == "__main__": infile = "Parameter.pdf" outfile = "ParameterOut.pdf" main(infile, outfile)
关键说明
get_table_bounding_boxes()返回Dict[int, List[Rectangle]],key为页码,value为该页所有表格的Rectangle对象列表- 使用
.get(0, [])可以安全获取第0页的表格列表,即使该页没有表格也不会报错 - 遍历表格列表时,每个
r都是Rectangle对象,具备grow方法,可以正常调用
内容的提问来源于stack exchange,提问作者Matt J Kenney
相关产品推荐
相关产品推荐

