You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Vertex AI Kubeflow Pipeline中Dataset元数据无法存储传递求助

问题解决:Vertex AI Kubeflow Pipeline中Output工件元数据传递失败

核心问题原因

你使用的kfp==2.0.0b13属于beta测试版本,存在Output工件元数据无法持久化传递的已知bug,该问题在后续稳定版本中已修复。此外,直接通过data.metadata[key] = value的赋值方式在部分旧版本中无法正确触发元数据的持久化逻辑。

解决方案

1. 升级KFP到稳定版本

首先升级KFP相关依赖至稳定版本,推荐版本组合:

pip install --upgrade kfp>=2.0.1 kfp-pipeline-spec>=0.2.1 kfp-server-api>=2.0.1

2. 修改组件代码(兼容多版本)

修改make_metadata组件,使用set_custom_metadata方法设置元数据(该方法在新旧版本中兼容性更好,稳定版中直接赋值字典也可生效):

from kfp.dsl import pipeline, component
from kfp.dsl import Input, Output, Dataset
from kfp import compiler, dsl

@component(packages_to_install=["pandas"], base_image='python:3.9')
def make_metadata(
  data: Output[Dataset],
):
    import pandas as pd
    param_out_df = pd.DataFrame({"dummy_col": "dummy_row"}, index=[0])
    param_out_df.to_csv(data.path, index=False)
    
    # 使用set_custom_metadata方法设置元数据,确保持久化
    data.set_custom_metadata("data_num", 1)
    data.set_custom_metadata("data_str", "random string")    
  
@component(packages_to_install=["pandas"], base_image='python:3.9')
def use_metadata(
    data: Input[Dataset],
):
    print("data - metadata")
    # 稳定版中直接访问data.metadata即可获取所有自定义元数据
    print(data.metadata)
    
@dsl.pipeline(
   name='test-pipeline',
   description='An example pipeline that performs arithmetic calculations.', 
   pipeline_root=f'{BUCKET}/pipelines'
)
def metadata_pipeline():
    metadata_made = make_metadata()
    used_metadata = use_metadata(data=metadata_made.outputs["data"])
    
PIPELINE_NAME = "test-pipeline"    
PIPELINE_FILENAME = f"{PIPELINE_NAME}.yaml"

compiler.Compiler().compile(
  pipeline_func=metadata_pipeline, 
  package_path=PIPELINE_FILENAME
)

3. 重新编译并运行Pipeline

升级依赖后,重新编译Pipeline YAML文件,再通过Vertex AI提交运行:

import datetime as datetime
from google.cloud import aiplatform

current_time = datetime.datetime.now().strftime("%Y%m%d_%H%M%S")
test_run_name = f"{PIPELINE_NAME}_{current_time}"

aiplatform.init(project=PROJECT_ID, location=LOCATION)
job = aiplatform.pipeline_jobs.PipelineJob(
    display_name=test_run_name,
    template_path=PIPELINE_FILENAME
)
job.run(sync=False)

验证结果

  • use_metadata组件的运行日志会打印出完整元数据:{'data_num': 1, 'data_str': 'random string'}
  • 在Vertex AI元数据谱系中查看Dataset工件,metadata字段会包含你设置的所有键值对

内容的提问来源于stack exchange,提问作者Gavin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.27 17:12:49