如何用Python的boto3调用AWS Textract提取三栏格式文档文本?
AWS Textract提取三栏文档文本的实现方案
完全可以用AWS Textract提取三栏(或更多栏)格式的文档文本,你现有的两栏处理代码只需要做少量调整就能适配多栏场景,以下是具体实现方法:
核心思路
原代码通过LINE块的边界框(BoundingBox)识别列的逻辑是通用的,支持自动识别任意数量的列。问题在于原代码仅按列索引排序,同一列内的行没有按垂直位置排序,导致输出顺序混乱。我们只需要补充垂直位置的排序逻辑即可。
修改后的适配代码
import boto3 # 文档路径 documentName = "three-column-image.jpg" # 初始化Textract客户端 textract = boto3.client('textract') # 调用Textract检测文档文本 with open(documentName, "rb") as document: response = textract.detect_document_text( Document={ 'Bytes': document.read(), } ) # 识别列并整理行数据 columns = [] lines = [] for item in response["Blocks"]: if item["BlockType"] == "LINE": column_found = False bbox = item["Geometry"]["BoundingBox"] bbox_left = bbox["Left"] bbox_right = bbox["Left"] + bbox["Width"] bbox_centre = bbox_left + bbox["Width"] / 2 bbox_top = bbox["Top"] # 记录行的垂直位置 # 检查当前行属于已有的哪一列 for index, column in enumerate(columns): column_centre = column['left'] + (column['right'] - column['left']) / 2 # 判断行与列的重叠关系 if (bbox_centre > column['left'] and bbox_centre < column['right']) or \ (column_centre > bbox_left and column_centre < bbox_right): lines.append({ 'column_index': index, 'text': item["Text"], 'top': bbox_top }) column_found = True break # 如果不属于任何现有列,新建列 if not column_found: columns.append({ 'left': bbox_left, 'right': bbox_right }) lines.append({ 'column_index': len(columns)-1, 'text': item["Text"], 'top': bbox_top }) # 按列索引(从左到右)、垂直位置(从上到下)排序 lines.sort(key=lambda x: (x['column_index'], x['top'])) # 输出结果 for line in lines: print(line['text'])
进阶优化:使用LAYOUT特征提升准确率
如果文档布局复杂(比如存在跨栏文本、不规则列宽),推荐使用analyze_document API并指定LAYOUT特征,Textract会直接返回COLUMN类型的块,无需手动计算边界框:
import boto3 documentName = "three-column-image.jpg" textract = boto3.client('textract') with open(documentName, "rb") as document: response = textract.analyze_document( Document={'Bytes': document.read()}, FeatureTypes=["LAYOUT"] ) # 先整理COLUMN块的位置和索引 columns = [] for block in response["Blocks"]: if block["BlockType"] == "COLUMN": bbox = block["Geometry"]["BoundingBox"] columns.append({ 'id': block["Id"], 'left': bbox["Left"], 'top': bbox["Top"] }) # 列按从左到右排序 columns.sort(key=lambda x: x['left']) # 构建COLUMN ID到索引的映射 column_id_map = {col['id']: idx for idx, col in enumerate(columns)} # 整理LINE块,关联到对应的列 lines = [] for block in response["Blocks"]: if block["BlockType"] == "LINE": # 找到当前LINE所属的COLUMN父块 for rel in block["Relationships"]: if rel["Type"] == "CHILD_OF": for parent_id in rel["Ids"]: if parent_id in column_id_map: bbox = block["Geometry"]["BoundingBox"] lines.append({ 'column_index': column_id_map[parent_id], 'text': block["Text"], 'top': bbox["Top"] }) break break # 按列和垂直位置排序输出 lines.sort(key=lambda x: (x['column_index'], x['top'])) for line in lines: print(line['text'])
内容的提问来源于stack exchange,提问作者Shady
相关产品推荐
相关产品推荐

