You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Amazon Textract Python代码处理多页PDF遇UnsupportedDocumentException求助

问题:Amazon Textract调用AnalyzeDocument时报UnsupportedDocumentException错误

我有一份多页PDF文档,在Amazon Textract网页界面中可正常处理并提取键值对,但使用Python代码调用AnalyzeDocument操作时,返回以下错误:

UnsupportedDocumentException: An error occurred (UnsupportedDocumentException) when calling the AnalyzeDocument operation: Request has unsupported document format

初始代码如下:

response = textract.analyze_document(
    Document={
        "S3Object": {
            "Bucket": bucketname,
            "Name": filename,
        }
    },
    FeatureTypes=["FORMS"],
    HumanLoopConfig={
        "HumanLoopName": uuid.uuid4().hex,
        "FlowDefinitionArn": FLOW_ARN,
        "DataAttributes": {
            "ContentClassifiers": [
                "FreeOfPersonallyIdentifiableInformation",
                "FreeOfAdultContent",
            ]
        },
    },
)
print(json.dumps(response))

return {
    "statusCode": 200,
    "body": json.dumps("Document processed successfully!"),
}

return {"statusCode": 500, "body": json.dumps("Issue processing file!")}

原以为是多页PDF导致无法读取,于是修改代码尝试逐页处理,修改后的代码如下:

# Start document text detection
response = textract.start_document_text_detection(
    DocumentLocation={
        "S3Object": {
            "Bucket": bucketname,
            "Name": filename,
        }
    },
    ClientRequestToken=str(uuid.uuid4())  # Generate a unique client request token
)

# Retrieve the job ID from the response
job_id = response["JobId"]

# Poll for the completion of the job
while True:
    job_status = textract.get_document_text_detection(JobId=job_id)['JobStatus']
    if job_status in ['SUCCEEDED', 'FAILED']:
        break
    time.sleep(5)  # Wait for 5 seconds before checking again

# Get the results of the detection
response = textract.get_document_text_detection(JobId=job_id)

# Process each page of the document
for page_result in response['Blocks']:
    if page_result['BlockType'] == 'PAGE':
        page_number = page_result['Page']
        response = textract.analyze_document(
            Document={
                "S3Object": {
                    "Bucket": bucketname,
                    "Name": filename,
                }
            },
            FeatureTypes=["FORMS"],
            HumanLoopConfig={
                "HumanLoopName": uuid.uuid4().hex,
                "FlowDefinitionArn": FLOW_ARN,
                "DataAttributes": {
                    "ContentClassifiers": [
                        "FreeOfAdultContent",
                    ]
                },
            },
        )
        print(json.dumps(response))

return {
    "statusCode": 200,
    "body": json.dumps("Document processed successfully!"),
}

return {"statusCode": 500, "body": json.dumps("Issue processing file!")}

但修改后仍出现相同的UnsupportedDocumentException错误,恳请提供帮助或解决思路。


解决思路与修复方案

  • 核心修复:指定多页PDF的处理页码
    Textract网页端会自动处理多页PDF,但AnalyzeDocument是同步API,默认不支持直接处理多页文档,必须显式指定Pages参数。你修改后的代码虽然获取了页码,但后续调用AnalyzeDocument时未传入该参数,导致报错。

    修改AnalyzeDocument调用部分,添加Pages参数:

    response = textract.analyze_document(
        Document={
            "S3Object": {
                "Bucket": bucketname,
                "Name": filename,
            }
        },
        FeatureTypes=["FORMS"],
        HumanLoopConfig={
            "HumanLoopName": uuid.uuid4().hex,
            "FlowDefinitionArn": FLOW_ARN,
            "DataAttributes": {
                "ContentClassifiers": [
                    "FreeOfAdultContent",
                ]
            },
        },
        Pages=[str(page_number)]  # 新增:指定当前要处理的页码
    )
    
  • 其他排查方向

    • 检查文档完整性:确认PDF未损坏、未加密(Textract不支持加密文档)
    • 验证S3权限:确保执行代码的IAM角色拥有目标S3对象的s3:GetObject权限
    • 更新SDK版本:升级boto3到最新版本,避免旧版本兼容性问题:
      pip install --upgrade boto3
      
    • 改用异步API:如果文档页数较多,推荐使用StartDocumentAnalysis异步API,专门处理多页文档,无需逐页调用:
      response = textract.start_document_analysis(
          DocumentLocation={
              "S3Object": {
                  "Bucket": bucketname,
                  "Name": filename,
              }
          },
          FeatureTypes=["FORMS"],
          HumanLoopConfig={
              "HumanLoopName": uuid.uuid4().hex,
              "FlowDefinitionArn": FLOW_ARN,
              "DataAttributes": {
                  "ContentClassifiers": [
                      "FreeOfAdultContent",
                  ]
              },
          }
      )
      # 后续轮询获取结果逻辑与start_document_text_detection一致
      

内容的提问来源于stack exchange,提问作者Ritesh Manglani

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 12:59:51