You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

调用start_human_loop传入S3对象至HumanLoopInput报错解决

解决Amazon A2I调用start_human_loop时的ValidationException错误

问题背景

使用Amazon Textract完成多页PDF文本提取后,编写Lambda代码调用SageMaker A2I的start_human_loop接口时触发ValidationException,错误提示「Provided InputContent is not valid」。使用的是AWS默认Worker模板,该模板依赖以下字段:

  • task.input.aiServiceRequest.document.s3Object.bucket
  • task.input.aiServiceRequest.document.s3Object.name
  • task.input.selectedAiServiceResponse.blocks
  • task.input.humanLoopContext.importantFormKeys
    但当前代码传入的InputContent结构与模板要求不匹配。

原因分析

默认Worker模板的Liquid语法严格依赖特定层级的JSON结构,原代码中InputContent仅传入了InitialValue字段,完全不符合模板所需的字段结构,导致A2I无法解析输入内容,触发验证错误。

修正后的Lambda代码

import os
import json
import time
import uuid
from urllib.parse import unquote_plus
import boto3

def lambda_handler(event, context):
    textract = boto3.client("textract")
    a2i = boto3.client("sagemaker-a2i-runtime")
    FLOW_ARN = os.environ["FLOW_ARN"]
    if event:
        file_obj = event["Records"][0]
        bucketname = str(file_obj["s3"]["bucket"]["name"])
        filename = unquote_plus(str(file_obj["s3"]["object"]["key"]))
        
        # 启动文档分析任务
        response = textract.start_document_analysis(
            DocumentLocation={
                "S3Object": {
                    "Bucket": bucketname,
                    "Name": filename,
                }
            },
            FeatureTypes=["FORMS"],
            ClientRequestToken=str(uuid.uuid4()),
        )
        
        job_id = response["JobId"]
        
        # 轮询任务状态
        while True:
            job_status = textract.get_document_analysis(JobId=job_id)['JobStatus']
            if job_status in ['SUCCEEDED', 'FAILED']:
                break
            time.sleep(5)
        
        # 获取完整的Textract分析结果(处理分页)
        textract_results = []
        next_token = None
        while True:
            if next_token:
                response = textract.get_document_analysis(JobId=job_id, NextToken=next_token)
            else:
                response = textract.get_document_analysis(JobId=job_id)
            textract_results.extend(response['Blocks'])
            next_token = response.get('NextToken')
            if not next_token:
                break
        
        # 调用A2I人工审核
        a2i.start_human_loop(
            HumanLoopName=uuid.uuid4().hex,
            FlowDefinitionArn=FLOW_ARN,
            HumanLoopInput={
                'InputContent': json.dumps({
                    "aiServiceRequest": {
                        "document": {
                            "s3Object": {
                                "bucket": bucketname,
                                "name": filename
                            }
                        }
                    },
                    "selectedAiServiceResponse": {
                        "blocks": textract_results
                    },
                    "humanLoopContext": {
                        "importantFormKeys": []  # 可指定需要重点审核的表单键,如["姓名", "地址"]
                    }
                })
            },
            DataAttributes={
                'ContentClassifiers': ['FreeOfAdultContent']
            }
        )

        return {
            "statusCode": 200,
            "body": json.dumps("文档处理成功,已触发人工审核!")
        }

    return {"statusCode": 500, "body": json.dumps("文件处理失败!")}

关键修正说明

  1. 对齐模板字段结构:严格按照模板依赖的层级构建InputContent,包含aiServiceRequest、selectedAiServiceResponse、humanLoopContext三个核心节点
  2. 完整获取Textract结果:处理Textract分页返回的Blocks,确保所有分析结果都传入A2I作为人工审核的初始值
  3. 可选配置重点字段:humanLoopContext.importantFormKeys可填入需要优先审核的表单键名数组,为空则默认展示所有键值对

内容的提问来源于stack exchange,提问作者Ritesh Manglani

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 18:21:00