使用Amazon Textract异步GetDocumentAnalysis分页时遇无效参数异常
Amazon Textract GetDocumentAnalysis 分页迭代时的Invalid Parameter异常问题分析
问题场景
使用Amazon Textract异步函数GetDocumentAnalysis实现分页(利用NextToken属性),采用迭代方式处理时,首次循环返回JobStatus为InProgress,后续调用API触发Invalid Parameter异常,相关代码如下:
def receive_analysis_data(job_id_value, next_token): """ This method uses Textract async function GetDocumentAnalysis to extract the analyzed data from the document in JSON Key-value pair format and store it in a Python dictionary :param next_token: pagination value initialized to None and sent as argument :parameter job_id_value is the JobId received by Textract async function StartDocumentAnalysis """ token = " " # used for next_token value job_status = {} # used for storing JobStatus retry_attempts = 5 block_list = {} # used to store the JSON data returned by GetDocumentAnalysis result = {} # used to store JSON data of response and next_response while retry_attempts > 0: if next_token is None: response = client.get_document_analysis(JobId=job_id_value) job_status = response['JobStatus'] if job_status == "SUCCEEDED": result.update(response) token = result.get('NextToken') else: retry_attempts -= 1 logger.info("job status is: %s ", job_status) logger.info("Retrying to analyze the document. retry attempts left %d:", retry_attempts) time.sleep(10) if token is not None: next_response = client.get_document_analysis(JobId=job_id_value, NextToken=token) job_status = next_response['JobStatus'] if job_status == "SUCCEEDED": result.update(next_response) token = next_response.get('NextToken') else: retry_attempts -= 1 logger.info("job status is: %s ", job_status) logger.info("Retrying to analyze the document. retry attempts left %d:", retry_attempts) time.sleep(10) else: logger.info("Analyzing document completed. JobStatus %s ", job_status) write_json_file(block_list) parse_json(block_list) break
问题根源
- Token初始化错误:
token初始值设为空格字符串" ",而非None。首次循环中,即使任务处于InProgress状态,代码仍会进入token is not None的分支,将无效的空格作为NextToken传入API,直接触发Invalid Parameter异常。 - 分支逻辑重叠:两个
if分支是独立判断,当next_token为None但任务未完成时,执行完第一个分支的重试逻辑后,会立即执行第二个if分支,此时token还是初始的无效值,导致错误调用。 - 变量逻辑错误:
block_list全程未被赋值,最后写入和解析的是空字典,不符合业务预期;job_status初始化为字典{},但实际存储的是字符串状态,语义不符。
修复后的代码
import time import logging logger = logging.getLogger(__name__) def receive_analysis_data(job_id_value, next_token=None): """ 调用Textract异步函数GetDocumentAnalysis提取文档分析数据,转为JSON键值对格式并存入Python字典 :param job_id_value: 调用StartDocumentAnalysis后获取的JobId :param next_token: 分页标识,初始为None """ token = next_token # 直接使用传入的next_token初始化 job_status = "" # 初始化为字符串类型 retry_attempts = 5 block_list = [] # 用列表存储所有Blocks数据 result = {} while retry_attempts > 0: try: if token is None: # 首次请求或无分页时调用 response = client.get_document_analysis(JobId=job_id_value) else: # 分页请求 response = client.get_document_analysis(JobId=job_id_value, NextToken=token) job_status = response['JobStatus'] if job_status == "SUCCEEDED": # 合并Blocks数据 if 'Blocks' in response: block_list.extend(response['Blocks']) # 更新分页token token = response.get('NextToken') # 无更多分页数据时退出循环 if token is None: logger.info("文档分析完成,JobStatus: %s", job_status) write_json_file(block_list) parse_json(block_list) break elif job_status == "IN_PROGRESS": # 任务进行中,重试 retry_attempts -= 1 logger.info("任务处理中,剩余重试次数: %d", retry_attempts) time.sleep(10) else: # 任务失败,直接退出 logger.error("任务失败,JobStatus: %s", job_status) break except Exception as e: logger.error("调用GetDocumentAnalysis失败: %s", str(e)) retry_attempts -= 1 time.sleep(10) if retry_attempts == 0: logger.error("重试次数耗尽,任务未完成")
关键修改说明
- 将
token初始化为传入的next_token(默认None),避免无效初始值 - 合并分支逻辑,用
if-else区分首次/分页请求,避免分支重叠执行 - 用列表
block_list存储所有Blocks数据,符合分页数据合并的业务需求 - 增加异常捕获,处理API调用中的其他潜在错误
- 明确区分任务状态(SUCCEEDED/IN_PROGRESS/其他),逻辑更清晰
内容的提问来源于stack exchange,提问作者syed zain
相关产品推荐
相关产品推荐

