IBM Cloud自定义镜像创建间歇超时,如何在代码层面处理该问题
IBM Cloud 自定义镜像间歇超时问题代码修复方案
从报错和代码逻辑来看,超时偶发、重试即可成功的核心原因是现有代码的容错逻辑覆盖不足,偶发的接口波动、服务端临时状态都会被误判为失败,可从以下几个代码层面优化:
- 扩展可重试的API异常范围
现有代码仅处理了502网络错误,实际IBM云接口常见的临时错误还有429(限流)、500(服务端内部错误)、503(服务不可用)、504(网关超时),这类错误都属于可重试范畴,不需要直接抛出终止流程。 - 优化镜像状态判断逻辑
不要仅把pending判定为中间状态,IBM云镜像创建过程中可能返回queued、processing等临时中间状态,现有代码遇到这类非pending非available的状态会直接抛出错误,建议增加中间状态白名单,只有遇到failed、deleted、rejected这类明确的失败终态才终止流程。 - 替换重试次数计时为绝对时间计时
现有逻辑通过重试次数乘以间隔来估算超时时间,容易因为退避逻辑的设置误差导致实际等待时长不足2小时就提前超时,改为记录查询起始时间,每次循环判断当前时间与起始时间的差值是否达到2小时,计时更准确。 - 上层增加整个创建流程的重试机制
因为偶发失败后重新运行任务即可成功,可在调用镜像创建的逻辑中增加全局重试,每次重试前先清理本次创建的无效镜像资源,避免资源残留。
修复后代码示例
状态查询函数修改
import time from datetime import datetime, timedelta # 可重试的API错误码 RETRYABLE_ERROR_CODES = {429, 500, 502, 503, 504} # 镜像创建中间状态白名单 INTERMEDIATE_STATUSES = {"pending", "queued", "processing", "preparing"} # 最大等待时长2小时 MAX_WAIT_DURATION = timedelta(hours=2) # 固定轮询间隔(可根据需求调整,建议30-60秒) POLL_INTERVAL = 60 def retrieve_image_status(ibm_client, image_id): start_time = datetime.now() while datetime.now() - start_time < MAX_WAIT_DURATION: try: time.sleep(POLL_INTERVAL) response = ibm_client.get_image(image_id) status = response.result["status"] if status in INTERMEDIATE_STATUSES: print(f"Image creation for image {response.result['name']} is {status}.....") elif status == "available": print(f"Custom image creation succeeded at {response.result['created_at']}.") return else: # 明确失败终态,直接抛出错误 raise RuntimeError( f"Image creation failed: Status {status}, reasons: {response.result.get('status_reasons', 'unknown')}" ) except ApiException as e: if e.code in RETRYABLE_ERROR_CODES: print(f"Encountered retryable error {e.code}: {e.message}, retrying...") continue # 非可重试错误直接抛出 raise e # 超时退出 raise RuntimeError(f"Image import has timed out after 2 hours")
上层调用逻辑修改(增加全局重试)
# 全局创建流程最大重试次数 MAX_CREATE_RETRIES = 2 def create_image_flow(): ssm_client = create_ssm_client() api_key = retrieve_ibm_config(ssm_client) authenticator = create_authenticator(api_key) ibm_client = create_ibm_client(authenticator) service_client = create_service_client(authenticator) resource_group_id = retrieve_resource_group_id(service_client, dc_ibm_env) image_prototype = create_image_prototype( source_bucket, source_object, resource_group_id ) for retry in range(MAX_CREATE_RETRIES + 1): try: images = ibm_client.list_images(name=source_object) if images.result["images"]: print(f"Image already exists with name {source_object}") image_id = images.result["images"][0]["id"] note_image_id(image_id) return response = ibm_client.create_image(image_prototype) image_id = response.result["id"] retrieve_image_status(ibm_client, image_id) note_image_id(image_id) return except (ApiException, RuntimeError) as e: print(f"Image creation attempt {retry + 1} failed: {str(e)}") # 清理残留的失败镜像 if 'image_id' in locals(): try: ibm_client.delete_image(image_id) print(f"Cleaned up failed image {image_id}") except Exception as clean_err: print(f"Failed to clean up image {image_id}: {str(clean_err)}") # 最后一次重试失败,抛出错误 if retry == MAX_CREATE_RETRIES: raise RuntimeError(f"All {MAX_CREATE_RETRIES + 1} image creation attempts failed") from e # 重试前等待一段时间 time.sleep(30)
内容的提问来源于stack exchange,提问作者Deepali Mittal
相关产品推荐
相关产品推荐

