You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Amazon Textract解析多页PDF,如何获取LAYOUT类BLOCK内容?

问题

我尝试用Amazon Textract抓取多页PDF,需要将内容按章节、子章节、表格格式化为JSON。通过Textract UI演示(开启LAYOUT和Table功能)能正常显示布局标题、章节、文本等信息,下载的layout.csv和analyzeDocResponse.json也包含LAYOUT相关数据。

编写代码1打印Block完整JSON结构时能看到LAYOUT类块,但代码2打印Block类型对应文本时,仅LINES、WORDS类型有内容,LAYOUT_TITLE、LAYOUT_SECTION_HEADER等LAYOUT类块的Text字段为空。我是Textract新手,查过文档和相关问题没找到解法,寻求帮助。测试PDF为药品SmPC文件。

代码1

def start_textract_job(bucket, document):
    response = textract.start_document_analysis(
        DocumentLocation={
            'S3Object': {
                'Bucket': bucket,
                'Name': document
            }
        },
        FeatureTypes=["LAYOUT"]  # 可根据需求调整FeatureTypes
    )
    return response['JobId']


def print_blocks(job_id):
    next_token = None
    while True:
        if next_token:
            response = textract.get_document_analysis(JobId=job_id, NextToken=next_token)
        else:
            response = textract.get_document_analysis(JobId=job_id)

        for block in response.get('Blocks', []):
            print(json.dumps(block, indent=4))

        next_token = response.get('NextToken', None)
        if not next_token:
            break

代码2

# start_textract_job函数与代码1相同,启用LAYOUT

def print_blocks(job_id):
    next_token = None
    while True:
        if next_token:
            response = textract.get_document_analysis(JobId=job_id, NextToken=next_token)
        else:
            response = textract.get_document_analysis(JobId=job_id)

        for block in response.get('Blocks', []):
            print(f"{block['BlockType']}: {block.get('Text', '')}")

        next_token = response.get('NextToken', None)
        if not next_token:
            break

解决方案

LAYOUT类型的Block(如LAYOUT_TITLE、LAYOUT_SECTION_HEADER)本身不存储文本内容,它们是布局容器块,仅用于标记文档的结构区域。实际文本内容存储在它们关联的子Block(如WORDS、LINES类型)中,通过Block的Relationships字段建立关联。

要获取LAYOUT类块对应的文本,需按以下步骤处理:

  1. 先遍历所有Block,建立以BlockId为键的字典,方便快速查找关联子块
  2. 对每个LAYOUT类块,通过其Relationships中的CHILD类型提取关联的子Block ID
  3. 从字典中取出这些子Block,拼接它们的Text内容

修改后的print_blocks函数示例:

def print_blocks(job_id):
    next_token = None
    block_dict = {}
    
    # 先收集所有Block到字典
    while True:
        if next_token:
            response = textract.get_document_analysis(JobId=job_id, NextToken=next_token)
        else:
            response = textract.get_document_analysis(JobId=job_id)
        
        for block in response.get('Blocks', []):
            block_dict[block['Id']] = block
        
        next_token = response.get('NextToken', None)
        if not next_token:
            break
    
    # 遍历所有Block,处理LAYOUT类块
    for block_id, block in block_dict.items():
        block_type = block['BlockType']
        if block_type.startswith('LAYOUT_'):
            # 获取关联的子Block ID
            child_ids = []
            if 'Relationships' in block:
                for rel in block['Relationships']:
                    if rel['Type'] == 'CHILD':
                        child_ids.extend(rel['Ids'])
            
            # 拼接子Block的文本内容
            text_parts = []
            for child_id in child_ids:
                child_block = block_dict.get(child_id)
                if child_block and 'Text' in child_block:
                    text_parts.append(child_block['Text'])
            
            full_text = ' '.join(text_parts)
            print(f"{block_type}: {full_text}")
        else:
            # 保留其他类型Block的原有打印逻辑
            print(f"{block_type}: {block.get('Text', '')}")

通过这种方式就能正确获取LAYOUT类块对应的文本,后续可基于这些布局块的层级关系,组织出章节、子章节结构,最终生成符合要求的JSON格式文档。

内容的提问来源于stack exchange,提问作者Santhosh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.26 16:23:14