使用Amazon Textract解析多页PDF,如何获取LAYOUT类BLOCK内容?
问题
我尝试用Amazon Textract抓取多页PDF,需要将内容按章节、子章节、表格格式化为JSON。通过Textract UI演示(开启LAYOUT和Table功能)能正常显示布局标题、章节、文本等信息,下载的layout.csv和analyzeDocResponse.json也包含LAYOUT相关数据。
编写代码1打印Block完整JSON结构时能看到LAYOUT类块,但代码2打印Block类型对应文本时,仅LINES、WORDS类型有内容,LAYOUT_TITLE、LAYOUT_SECTION_HEADER等LAYOUT类块的Text字段为空。我是Textract新手,查过文档和相关问题没找到解法,寻求帮助。测试PDF为药品SmPC文件。
代码1
def start_textract_job(bucket, document): response = textract.start_document_analysis( DocumentLocation={ 'S3Object': { 'Bucket': bucket, 'Name': document } }, FeatureTypes=["LAYOUT"] # 可根据需求调整FeatureTypes ) return response['JobId'] def print_blocks(job_id): next_token = None while True: if next_token: response = textract.get_document_analysis(JobId=job_id, NextToken=next_token) else: response = textract.get_document_analysis(JobId=job_id) for block in response.get('Blocks', []): print(json.dumps(block, indent=4)) next_token = response.get('NextToken', None) if not next_token: break
代码2
# start_textract_job函数与代码1相同,启用LAYOUT def print_blocks(job_id): next_token = None while True: if next_token: response = textract.get_document_analysis(JobId=job_id, NextToken=next_token) else: response = textract.get_document_analysis(JobId=job_id) for block in response.get('Blocks', []): print(f"{block['BlockType']}: {block.get('Text', '')}") next_token = response.get('NextToken', None) if not next_token: break
解决方案
LAYOUT类型的Block(如LAYOUT_TITLE、LAYOUT_SECTION_HEADER)本身不存储文本内容,它们是布局容器块,仅用于标记文档的结构区域。实际文本内容存储在它们关联的子Block(如WORDS、LINES类型)中,通过Block的Relationships字段建立关联。
要获取LAYOUT类块对应的文本,需按以下步骤处理:
- 先遍历所有Block,建立以
BlockId为键的字典,方便快速查找关联子块 - 对每个LAYOUT类块,通过其
Relationships中的CHILD类型提取关联的子Block ID - 从字典中取出这些子Block,拼接它们的Text内容
修改后的print_blocks函数示例:
def print_blocks(job_id): next_token = None block_dict = {} # 先收集所有Block到字典 while True: if next_token: response = textract.get_document_analysis(JobId=job_id, NextToken=next_token) else: response = textract.get_document_analysis(JobId=job_id) for block in response.get('Blocks', []): block_dict[block['Id']] = block next_token = response.get('NextToken', None) if not next_token: break # 遍历所有Block,处理LAYOUT类块 for block_id, block in block_dict.items(): block_type = block['BlockType'] if block_type.startswith('LAYOUT_'): # 获取关联的子Block ID child_ids = [] if 'Relationships' in block: for rel in block['Relationships']: if rel['Type'] == 'CHILD': child_ids.extend(rel['Ids']) # 拼接子Block的文本内容 text_parts = [] for child_id in child_ids: child_block = block_dict.get(child_id) if child_block and 'Text' in child_block: text_parts.append(child_block['Text']) full_text = ' '.join(text_parts) print(f"{block_type}: {full_text}") else: # 保留其他类型Block的原有打印逻辑 print(f"{block_type}: {block.get('Text', '')}")
通过这种方式就能正确获取LAYOUT类块对应的文本,后续可基于这些布局块的层级关系,组织出章节、子章节结构,最终生成符合要求的JSON格式文档。
内容的提问来源于stack exchange,提问作者Santhosh
相关产品推荐
相关产品推荐

