从Amazon Textract提取表格文本与Bounding Box时遇错求助
错误原因分析
你遇到的TypeError是因为:
- 你将
table_blocks处理成了仅包含TABLE类型块的列表,但后续代码里用ROW/CELL的Id(字符串)去索引这个列表——列表只能用整数索引,自然触发报错。 - 更关键的是,ROW、CELL类型的块根本不在这个过滤后的
table_blocks列表里,你需要从Textract返回的所有Blocks中通过Id查找这些块。
修正后的代码
import boto3 # Initialize the Textract client client = boto3.client('textract') with open('table_document.pdf', 'rb') as file: # Call Amazon Textract to analyze the document response = client.analyze_document(Document={'Bytes': file.read()}, FeatureTypes=['TABLES']) # 把所有Blocks转换成字典,用Id作为键,方便快速查找任意块 blocks_dict = {block['Id']: block for block in response['Blocks']} # 过滤出所有TABLE类型的块 table_blocks = [b for b in response['Blocks'] if b['BlockType'] == 'TABLE'] # Iterate over each table block for table_block in table_blocks: # 获取当前表格关联的所有ROW块Id row_blocks_ids = table_block['Relationships'][0]['Ids'] # 从字典中取出ROW块,并按Top坐标从上到下排序 row_blocks = [blocks_dict[row_id] for row_id in row_blocks_ids] row_blocks.sort(key=lambda row: row['Geometry']['BoundingBox']['Top']) # Iterate over each row block for row_block in row_blocks: # 获取当前行关联的所有CELL块Id cell_blocks_ids = row_block['Relationships'][0]['Ids'] # 从字典中取出CELL块,并按Left坐标从左到右排序 cell_blocks = [blocks_dict[cell_id] for cell_id in cell_blocks_ids] cell_blocks.sort(key=lambda cell: cell['Geometry']['BoundingBox']['Left']) # Iterate over each cell block for cell_block in cell_blocks: # 获取单元格文本和 bounding box cell_text = cell_block['Text'] box = cell_block['Geometry']['BoundingBox'] # 打印结果 print(f'{cell_text}: {box}')
关键修改点
- 新增
blocks_dict:将所有返回的Block按Id映射,彻底解决通过Id查找块的问题。 - 排序ROW/CELL时,直接从字典中取出对应的块对象,避免用字符串索引列表的错误。
- 调整ROW和CELL的处理逻辑,先取出完整的块对象再排序,代码更直观易读。
内容的提问来源于stack exchange,提问作者InterestingPenguin80
相关产品推荐
相关产品推荐

