You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python的boto3调用AWS Textract提取三栏格式文档文本?

AWS Textract提取三栏文档文本的实现方案

完全可以用AWS Textract提取三栏(或更多栏)格式的文档文本,你现有的两栏处理代码只需要做少量调整就能适配多栏场景,以下是具体实现方法:

核心思路

原代码通过LINE块的边界框(BoundingBox)识别列的逻辑是通用的,支持自动识别任意数量的列。问题在于原代码仅按列索引排序,同一列内的行没有按垂直位置排序,导致输出顺序混乱。我们只需要补充垂直位置的排序逻辑即可。

修改后的适配代码

import boto3

# 文档路径
documentName = "three-column-image.jpg"

# 初始化Textract客户端
textract = boto3.client('textract')

# 调用Textract检测文档文本
with open(documentName, "rb") as document:
    response = textract.detect_document_text(
        Document={
            'Bytes': document.read(),
        }
    )

# 识别列并整理行数据
columns = []
lines = []
for item in response["Blocks"]:
    if item["BlockType"] == "LINE":
        column_found = False
        bbox = item["Geometry"]["BoundingBox"]
        bbox_left = bbox["Left"]
        bbox_right = bbox["Left"] + bbox["Width"]
        bbox_centre = bbox_left + bbox["Width"] / 2
        bbox_top = bbox["Top"]  # 记录行的垂直位置

        # 检查当前行属于已有的哪一列
        for index, column in enumerate(columns):
            column_centre = column['left'] + (column['right'] - column['left']) / 2
            # 判断行与列的重叠关系
            if (bbox_centre > column['left'] and bbox_centre < column['right']) or \
               (column_centre > bbox_left and column_centre < bbox_right):
                lines.append({
                    'column_index': index,
                    'text': item["Text"],
                    'top': bbox_top
                })
                column_found = True
                break
        # 如果不属于任何现有列,新建列
        if not column_found:
            columns.append({
                'left': bbox_left,
                'right': bbox_right
            })
            lines.append({
                'column_index': len(columns)-1,
                'text': item["Text"],
                'top': bbox_top
            })

# 按列索引(从左到右)、垂直位置(从上到下)排序
lines.sort(key=lambda x: (x['column_index'], x['top']))

# 输出结果
for line in lines:
    print(line['text'])

进阶优化:使用LAYOUT特征提升准确率

如果文档布局复杂(比如存在跨栏文本、不规则列宽),推荐使用analyze_document API并指定LAYOUT特征,Textract会直接返回COLUMN类型的块,无需手动计算边界框:

import boto3

documentName = "three-column-image.jpg"
textract = boto3.client('textract')

with open(documentName, "rb") as document:
    response = textract.analyze_document(
        Document={'Bytes': document.read()},
        FeatureTypes=["LAYOUT"]
    )

# 先整理COLUMN块的位置和索引
columns = []
for block in response["Blocks"]:
    if block["BlockType"] == "COLUMN":
        bbox = block["Geometry"]["BoundingBox"]
        columns.append({
            'id': block["Id"],
            'left': bbox["Left"],
            'top': bbox["Top"]
        })
# 列按从左到右排序
columns.sort(key=lambda x: x['left'])

# 构建COLUMN ID到索引的映射
column_id_map = {col['id']: idx for idx, col in enumerate(columns)}

# 整理LINE块,关联到对应的列
lines = []
for block in response["Blocks"]:
    if block["BlockType"] == "LINE":
        # 找到当前LINE所属的COLUMN父块
        for rel in block["Relationships"]:
            if rel["Type"] == "CHILD_OF":
                for parent_id in rel["Ids"]:
                    if parent_id in column_id_map:
                        bbox = block["Geometry"]["BoundingBox"]
                        lines.append({
                            'column_index': column_id_map[parent_id],
                            'text': block["Text"],
                            'top': bbox["Top"]
                        })
                        break
                break

# 按列和垂直位置排序输出
lines.sort(key=lambda x: (x['column_index'], x['top']))
for line in lines:
    print(line['text'])

内容的提问来源于stack exchange,提问作者Shady

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 16:43:12