调用start_human_loop传入S3对象至HumanLoopInput报错解决
解决Amazon A2I调用start_human_loop时的ValidationException错误
问题背景
使用Amazon Textract完成多页PDF文本提取后,编写Lambda代码调用SageMaker A2I的start_human_loop接口时触发ValidationException,错误提示「Provided InputContent is not valid」。使用的是AWS默认Worker模板,该模板依赖以下字段:
task.input.aiServiceRequest.document.s3Object.buckettask.input.aiServiceRequest.document.s3Object.nametask.input.selectedAiServiceResponse.blockstask.input.humanLoopContext.importantFormKeys
但当前代码传入的InputContent结构与模板要求不匹配。
原因分析
默认Worker模板的Liquid语法严格依赖特定层级的JSON结构,原代码中InputContent仅传入了InitialValue字段,完全不符合模板所需的字段结构,导致A2I无法解析输入内容,触发验证错误。
修正后的Lambda代码
import os import json import time import uuid from urllib.parse import unquote_plus import boto3 def lambda_handler(event, context): textract = boto3.client("textract") a2i = boto3.client("sagemaker-a2i-runtime") FLOW_ARN = os.environ["FLOW_ARN"] if event: file_obj = event["Records"][0] bucketname = str(file_obj["s3"]["bucket"]["name"]) filename = unquote_plus(str(file_obj["s3"]["object"]["key"])) # 启动文档分析任务 response = textract.start_document_analysis( DocumentLocation={ "S3Object": { "Bucket": bucketname, "Name": filename, } }, FeatureTypes=["FORMS"], ClientRequestToken=str(uuid.uuid4()), ) job_id = response["JobId"] # 轮询任务状态 while True: job_status = textract.get_document_analysis(JobId=job_id)['JobStatus'] if job_status in ['SUCCEEDED', 'FAILED']: break time.sleep(5) # 获取完整的Textract分析结果(处理分页) textract_results = [] next_token = None while True: if next_token: response = textract.get_document_analysis(JobId=job_id, NextToken=next_token) else: response = textract.get_document_analysis(JobId=job_id) textract_results.extend(response['Blocks']) next_token = response.get('NextToken') if not next_token: break # 调用A2I人工审核 a2i.start_human_loop( HumanLoopName=uuid.uuid4().hex, FlowDefinitionArn=FLOW_ARN, HumanLoopInput={ 'InputContent': json.dumps({ "aiServiceRequest": { "document": { "s3Object": { "bucket": bucketname, "name": filename } } }, "selectedAiServiceResponse": { "blocks": textract_results }, "humanLoopContext": { "importantFormKeys": [] # 可指定需要重点审核的表单键,如["姓名", "地址"] } }) }, DataAttributes={ 'ContentClassifiers': ['FreeOfAdultContent'] } ) return { "statusCode": 200, "body": json.dumps("文档处理成功,已触发人工审核!") } return {"statusCode": 500, "body": json.dumps("文件处理失败!")}
关键修正说明
- 对齐模板字段结构:严格按照模板依赖的层级构建
InputContent,包含aiServiceRequest、selectedAiServiceResponse、humanLoopContext三个核心节点 - 完整获取Textract结果:处理Textract分页返回的
Blocks,确保所有分析结果都传入A2I作为人工审核的初始值 - 可选配置重点字段:
humanLoopContext.importantFormKeys可填入需要优先审核的表单键名数组,为空则默认展示所有键值对
内容的提问来源于stack exchange,提问作者Ritesh Manglani
相关产品推荐
相关产品推荐

