You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Amazon Textract异步GetDocumentAnalysis分页时遇无效参数异常

Amazon Textract GetDocumentAnalysis 分页迭代时的Invalid Parameter异常问题分析

问题场景

使用Amazon Textract异步函数GetDocumentAnalysis实现分页(利用NextToken属性),采用迭代方式处理时,首次循环返回JobStatus为InProgress,后续调用API触发Invalid Parameter异常,相关代码如下:

def receive_analysis_data(job_id_value, next_token):
    """
    This method uses Textract async function GetDocumentAnalysis to extract
    the analyzed data from the document in JSON Key-value pair format and store it in a Python dictionary
    :param next_token: pagination value initialized to None and sent as argument
    :parameter job_id_value is the JobId received by Textract async function StartDocumentAnalysis
    """
    token = " " # used for next_token value
    job_status = {} # used for storing JobStatus
    retry_attempts = 5
    block_list = {} # used to store the JSON data returned by GetDocumentAnalysis
    result = {} # used to store JSON data of response and next_response
    while retry_attempts > 0: 
        if next_token is None:
            response = client.get_document_analysis(JobId=job_id_value)
            job_status = response['JobStatus']
            if job_status == "SUCCEEDED":
                result.update(response)
                token = result.get('NextToken')
            else:
                retry_attempts -= 1
                logger.info("job status is: %s ", job_status)
                logger.info("Retrying to analyze the document. retry attempts left %d:", retry_attempts)
                time.sleep(10)

        if token is not None:
            next_response = client.get_document_analysis(JobId=job_id_value, NextToken=token)
            job_status = next_response['JobStatus']
            if job_status == "SUCCEEDED":
                result.update(next_response)
                token = next_response.get('NextToken')
            else:
                retry_attempts -= 1
                logger.info("job status is: %s ", job_status)
                logger.info("Retrying to analyze the document. retry attempts left %d:", retry_attempts)
                time.sleep(10)
        else:
            logger.info("Analyzing document completed. JobStatus %s ", job_status)
            write_json_file(block_list)
            parse_json(block_list)
            break

问题根源

  1. Token初始化错误:token初始值设为空格字符串" ",而非None。首次循环中,即使任务处于InProgress状态,代码仍会进入token is not None的分支,将无效的空格作为NextToken传入API,直接触发Invalid Parameter异常。
  2. 分支逻辑重叠:两个if分支是独立判断,当next_token为None但任务未完成时,执行完第一个分支的重试逻辑后,会立即执行第二个if分支,此时token还是初始的无效值,导致错误调用。
  3. 变量逻辑错误:block_list全程未被赋值,最后写入和解析的是空字典,不符合业务预期;job_status初始化为字典{},但实际存储的是字符串状态,语义不符。

修复后的代码

import time
import logging

logger = logging.getLogger(__name__)

def receive_analysis_data(job_id_value, next_token=None):
    """
    调用Textract异步函数GetDocumentAnalysis提取文档分析数据,转为JSON键值对格式并存入Python字典
    :param job_id_value: 调用StartDocumentAnalysis后获取的JobId
    :param next_token: 分页标识,初始为None
    """
    token = next_token  # 直接使用传入的next_token初始化
    job_status = ""  # 初始化为字符串类型
    retry_attempts = 5
    block_list = []  # 用列表存储所有Blocks数据
    result = {}

    while retry_attempts > 0:
        try:
            if token is None:
                # 首次请求或无分页时调用
                response = client.get_document_analysis(JobId=job_id_value)
            else:
                # 分页请求
                response = client.get_document_analysis(JobId=job_id_value, NextToken=token)
            
            job_status = response['JobStatus']
            
            if job_status == "SUCCEEDED":
                # 合并Blocks数据
                if 'Blocks' in response:
                    block_list.extend(response['Blocks'])
                # 更新分页token
                token = response.get('NextToken')
                
                # 无更多分页数据时退出循环
                if token is None:
                    logger.info("文档分析完成,JobStatus: %s", job_status)
                    write_json_file(block_list)
                    parse_json(block_list)
                    break
            elif job_status == "IN_PROGRESS":
                # 任务进行中,重试
                retry_attempts -= 1
                logger.info("任务处理中,剩余重试次数: %d", retry_attempts)
                time.sleep(10)
            else:
                # 任务失败,直接退出
                logger.error("任务失败,JobStatus: %s", job_status)
                break
        except Exception as e:
            logger.error("调用GetDocumentAnalysis失败: %s", str(e))
            retry_attempts -= 1
            time.sleep(10)
    
    if retry_attempts == 0:
        logger.error("重试次数耗尽,任务未完成")

关键修改说明

  • 将token初始化为传入的next_token(默认None),避免无效初始值
  • 合并分支逻辑,用if-else区分首次/分页请求,避免分支重叠执行
  • 用列表block_list存储所有Blocks数据,符合分页数据合并的业务需求
  • 增加异常捕获,处理API调用中的其他潜在错误
  • 明确区分任务状态(SUCCEEDED/IN_PROGRESS/其他),逻辑更清晰

内容的提问来源于stack exchange,提问作者syed zain

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 17:53:17