如何让Kubeflow组件内部函数的异常终止VertexAI流水线?
关于Vertex AI流水线临时组件异常无法终止的问题处理
当在ephemeral_component内部调用的函数中抛出异常时,Vertex AI流水线不会停止。示例流水线的日志中已出现明确的堆栈跟踪:
2024-04-09 14:10:34.753 BST File "/tmp/tmp.c37YCfzDVS/ephemeral_component.py", line 109, in train_classifier 2024-04-09 14:10:34.753 BST train_df = preprocess_dataset_sequence( 2024-04-09 14:10:34.753 BST File "/pytorch_gpu/src/processes.py", line 530, in preprocess_dataset_sequence 2024-04-09 14:10:34.753 BST raise NotImplementedError("Debug raising exception to check if VertexAI stops pipeline.") 2024-04-09 14:10:34.753 BST NotImplementedError: Debug raising exception to check if VertexAI stops pipeline.
但该流水线始终处于Running状态。直接在train_classifier组件内抛出异常会终止流水线,而在其调用的任意子函数内抛出异常则无法触发流水线终止。请问是否存在处理Kubeflow组件异常的正确方式?
最小代码示例
@component( base_image=os.environ["PYTORCH_GPU_IMAGE_URI"], ) def train_classifier(): from src.processes import debug_throw debug_throw() def debug_throw(): raise NotImplementedError("Debug raising exception to check if VertexAI stops pipeline.")
编辑补充
经核查,组件出错后Vertex AI的成本可忽略不计(运行数日仅花费3美分),这可能只是Vertex AI流水线的状态显示问题。
处理方案
- 显式捕获并重新抛出异常:在组件顶层函数中捕获子函数抛出的异常并重新抛出,确保异常能被Vertex AI的流水线监控机制识别:
@component( base_image=os.environ["PYTORCH_GPU_IMAGE_URI"], ) def train_classifier(): from src.processes import debug_throw import logging try: debug_throw() except Exception as e: logging.error("组件执行失败,触发流水线终止", exc_info=True) raise # 重新抛出异常,让进程返回非0退出码
- 确认进程退出码:Kubeflow组件通过进程退出码判断执行状态,子函数的异常若未传递到顶层导致进程退出码异常,流水线会误判为运行中。重新抛出异常会让Python解释器返回非0退出码,触发流水线终止逻辑。
- 排查状态显示延迟:结合成本核查结果,若实际资源已释放仅状态显示异常,可通过流水线的组件详情页查看具体执行状态,或使用API查询实际状态:
from google.cloud import aiplatform # 替换为你的流水线任务ID pipeline_job = aiplatform.PipelineJob.get("projects/[PROJECT_ID]/locations/[REGION]/pipelineJobs/[JOB_ID]") print("流水线实际状态:", pipeline_job.state) print("任务详情:", pipeline_job.job_details)
内容的提问来源于stack exchange,提问作者WurmD
相关产品推荐
相关产品推荐

