Amazon Textract Python代码处理多页PDF遇UnsupportedDocumentException求助
问题:Amazon Textract调用AnalyzeDocument时报UnsupportedDocumentException错误
我有一份多页PDF文档,在Amazon Textract网页界面中可正常处理并提取键值对,但使用Python代码调用AnalyzeDocument操作时,返回以下错误:
UnsupportedDocumentException: An error occurred (UnsupportedDocumentException) when calling the AnalyzeDocument operation: Request has unsupported document format
初始代码如下:
response = textract.analyze_document( Document={ "S3Object": { "Bucket": bucketname, "Name": filename, } }, FeatureTypes=["FORMS"], HumanLoopConfig={ "HumanLoopName": uuid.uuid4().hex, "FlowDefinitionArn": FLOW_ARN, "DataAttributes": { "ContentClassifiers": [ "FreeOfPersonallyIdentifiableInformation", "FreeOfAdultContent", ] }, }, ) print(json.dumps(response)) return { "statusCode": 200, "body": json.dumps("Document processed successfully!"), } return {"statusCode": 500, "body": json.dumps("Issue processing file!")}
原以为是多页PDF导致无法读取,于是修改代码尝试逐页处理,修改后的代码如下:
# Start document text detection response = textract.start_document_text_detection( DocumentLocation={ "S3Object": { "Bucket": bucketname, "Name": filename, } }, ClientRequestToken=str(uuid.uuid4()) # Generate a unique client request token ) # Retrieve the job ID from the response job_id = response["JobId"] # Poll for the completion of the job while True: job_status = textract.get_document_text_detection(JobId=job_id)['JobStatus'] if job_status in ['SUCCEEDED', 'FAILED']: break time.sleep(5) # Wait for 5 seconds before checking again # Get the results of the detection response = textract.get_document_text_detection(JobId=job_id) # Process each page of the document for page_result in response['Blocks']: if page_result['BlockType'] == 'PAGE': page_number = page_result['Page'] response = textract.analyze_document( Document={ "S3Object": { "Bucket": bucketname, "Name": filename, } }, FeatureTypes=["FORMS"], HumanLoopConfig={ "HumanLoopName": uuid.uuid4().hex, "FlowDefinitionArn": FLOW_ARN, "DataAttributes": { "ContentClassifiers": [ "FreeOfAdultContent", ] }, }, ) print(json.dumps(response)) return { "statusCode": 200, "body": json.dumps("Document processed successfully!"), } return {"statusCode": 500, "body": json.dumps("Issue processing file!")}
但修改后仍出现相同的UnsupportedDocumentException错误,恳请提供帮助或解决思路。
解决思路与修复方案
核心修复:指定多页PDF的处理页码
Textract网页端会自动处理多页PDF,但AnalyzeDocument是同步API,默认不支持直接处理多页文档,必须显式指定Pages参数。你修改后的代码虽然获取了页码,但后续调用AnalyzeDocument时未传入该参数,导致报错。修改
AnalyzeDocument调用部分,添加Pages参数:response = textract.analyze_document( Document={ "S3Object": { "Bucket": bucketname, "Name": filename, } }, FeatureTypes=["FORMS"], HumanLoopConfig={ "HumanLoopName": uuid.uuid4().hex, "FlowDefinitionArn": FLOW_ARN, "DataAttributes": { "ContentClassifiers": [ "FreeOfAdultContent", ] }, }, Pages=[str(page_number)] # 新增:指定当前要处理的页码 )其他排查方向
- 检查文档完整性:确认PDF未损坏、未加密(Textract不支持加密文档)
- 验证S3权限:确保执行代码的IAM角色拥有目标S3对象的
s3:GetObject权限 - 更新SDK版本:升级boto3到最新版本,避免旧版本兼容性问题:
pip install --upgrade boto3 - 改用异步API:如果文档页数较多,推荐使用
StartDocumentAnalysis异步API,专门处理多页文档,无需逐页调用:response = textract.start_document_analysis( DocumentLocation={ "S3Object": { "Bucket": bucketname, "Name": filename, } }, FeatureTypes=["FORMS"], HumanLoopConfig={ "HumanLoopName": uuid.uuid4().hex, "FlowDefinitionArn": FLOW_ARN, "DataAttributes": { "ContentClassifiers": [ "FreeOfAdultContent", ] }, } ) # 后续轮询获取结果逻辑与start_document_text_detection一致
内容的提问来源于stack exchange,提问作者Ritesh Manglani
相关产品推荐
相关产品推荐

