在Vertex AI(Kubeflow Runner)运行自定义TFX组件遇Executor导入错误
解决GCP Vertex AI上TFX自定义组件Executor找不到的问题
针对你遇到的ImportError: Executor class couldn’t be found in main错误,结合TFX 1.15、KubeflowDagRunner和Vertex AI的环境,以下是两种组件定义方式的具体解决步骤:
一、使用@component装饰器的组件
- 必须给装饰器指定
package_path参数,明确组件代码所在的根目录(该目录会被完整打包上传),示例:@component( package_path='./my_custom_components', # 组件代码所在的根目录 base_image='tensorflow/tfx:1.15.0' ) def sample_custom_component(input_data: str) -> str: # 自定义组件逻辑 return f"{input_data}_processed" - 编译流水线时,配置
KubeflowDagRunner的代码上传器,确保本地代码被上传到GCS的指定路径,示例:from tfx.orchestration.kubeflow import KubeflowDagRunner from tfx.orchestration.kubeflow.gcp import GcsCodeUploader runner_config = KubeflowDagRunnerConfig( default_image='gcr.io/your-project-id/tfx-1.15:custom', code_uploader=GcsCodeUploader( gcs_output_uri='gs://your-bucket-name/tfx-component-code' ) ) KubeflowDagRunner(config=runner_config).run(your_pipeline)
二、自定义Executor+ComponentSpec的组件
- 严格指定
EXECUTOR_SPEC的完整模块路径:如果你的Executor类在my_components/executors/custom_executor.py中,类名为MyExecutor,则ComponentSpec的定义要写全路径:from tfx import types from tfx.dsl.component.experimental import executor_spec class MyCustomComponentSpec(types.ComponentSpec): PARAMETERS = { 'param1': types.ExecutionParameter(type=str), } INPUTS = { 'input_data': types.ChannelParameter(type=types.Artifact), } OUTPUTS = { 'output_data': types.ChannelParameter(type=types.Artifact), } EXECUTOR_SPEC = executor_spec.ExecutorClassSpec( 'my_components.executors.custom_executor.MyExecutor' # 完整模块路径+类名 ) - 打包代码时要保留完整目录结构:打包命令需包含整个组件模块目录,不能只打包单个文件,示例:
tar -czf custom_components.tar.gz my_components/ - 上传GCS后,确保组件定义时指向正确的代码包路径,同时使用与TFX 1.15兼容的base镜像(如
tensorflow/tfx:1.15.0)。
通用排查步骤
- 镜像版本一致性:确认Vertex AI使用的镜像中TFX版本为1.15,与本地开发环境完全匹配,避免版本差异导致的导入问题。
- 验证GCS代码包:下载GCS上的tar.gz文件解压,检查目录结构是否与你指定的模块路径一致,确保没有遗漏文件或目录层级错误。
- 查看Pod日志:在Vertex AI Pipeline的运行详情中,找到报错组件对应的Pod,查看完整日志,定位具体的导入路径错误(比如模块名拼写错误、代码未被正确挂载到Pod的Python路径中)。
内容的提问来源于stack exchange,提问作者crbl
相关产品推荐
相关产品推荐

