Lambda调用start_document_analysis时AdaptersConfig未被识别的问题
AWS Textract
start_document_analysis 添加 AdaptersConfig 参数报错排查 问题背景
使用Textract Queries从上传的PDF中提取特定信息,通过Lambda函数textract_async_job_creation触发异步分析任务,调用start_document_analysis时添加AdaptersConfig参数以使用自定义训练的适配器,但触发"参数未被识别"的错误,官方文档显示该参数支持与start_document_analysis配合使用。
可能原因及解决方案
1. Boto3版本过低
Textract的适配器功能属于较新特性,旧版本的Boto3客户端未包含该参数的定义,导致调用时被判定为非法参数。
解决方法:
升级Lambda环境中的Boto3和Botocore版本:
- 本地创建Lambda层:
- 创建临时目录,执行
pip install boto3 botocore -t ./ - 将目录打包为zip文件,上传至AWS Lambda作为层
- 将该层关联到
textract_async_job_creation函数
- 创建临时目录,执行
- 或者直接在函数的部署包中包含更新后的Boto3库
2. 参数拼写/格式错误
确认参数名称和内部结构是否完全符合要求:
- 注意参数名是**
AdaptersConfig**(复数形式),而非AdapterConfig - 内部嵌套的
Adapters同样需要保持复数格式,且每个适配器需包含AdapterId和Version字段
修正后的参数示例:
AdaptersConfig={ 'Adapters': [ {'AdapterId': 'xxxxxxxxxxx', 'Version': '1'} ] }
3. 区域兼容性问题
部分AWS区域可能暂未推出Textract适配器功能,需确认us-east-2区域是否支持该特性。可通过AWS官方服务区域列表确认功能覆盖范围。
修正后的完整代码
import os import json import boto3 from botocore.config import Config from urllib.parse import unquote_plus my_config = Config( region_name='us-east-2', retries={ 'max_attempts': 10, 'mode': 'adaptive' } ) textract = boto3.client('textract', config=my_config) OUTPUT_BUCKET_NAME = os.environ["OUTPUT_BUCKET_NAME"] OUTPUT_S3_PREFIX = os.environ["OUTPUT_S3_PREFIX"] SNS_TOPIC_ARN = os.environ["SNS_TOPIC_ARN"] SNS_ROLE_ARN = os.environ["SNS_ROLE_ARN"] def lambda_handler(event, context): responses = [] for record in event["Records"]: file_obj = record["s3"] bucketname = str(file_obj["bucket"]["name"]) filename = unquote_plus(str(file_obj["object"]["key"])) print(f"Bucket: {bucketname} ::: Key: {filename}") response = textract.start_document_analysis( DocumentLocation={'S3Object': {'Bucket': bucketname, 'Name': filename}}, FeatureTypes=['QUERIES'], OutputConfig={'S3Bucket': OUTPUT_BUCKET_NAME, 'S3Prefix': OUTPUT_S3_PREFIX}, NotificationChannel={'SNSTopicArn': SNS_TOPIC_ARN, 'RoleArn': SNS_ROLE_ARN}, QueriesConfig={ 'Queries': [ {'Text': 'What is the name of the claimant?', 'Pages': ['1']}, {'Text': 'What is the date on the document?', 'Pages': ['1']}, {'Text': 'What is the phone number?', 'Pages': ['1']}, {'Text': 'What is the address of the office?', 'Pages': ['1']} ] }, AdaptersConfig={ 'Adapters': [ {'AdapterId': 'xxxxxxxxxxx', 'Version': '1'} ] } ) responses.append(response) successful_responses = [resp for resp in responses if resp["ResponseMetadata"]["HTTPStatusCode"] == 200] failed_responses = [resp for resp in responses if resp["ResponseMetadata"]["HTTPStatusCode"] != 200] if successful_responses: return {"statusCode": 200, "body": json.dumps(f"Job(s) created successfully for {len(successful_responses)} file(s)!")} else: return {"statusCode": 500, "body": json.dumps(f"Job creation failed for {len(failed_responses)} file(s)!")}
内容的提问来源于stack exchange,提问作者user21067592
相关产品推荐
相关产品推荐

