如何在MLrun中读取已存储为工件的CSV文件?
解决MLrun中读取已存储工件的问题
直接用pd.read_csv()访问S3路径报错,是因为MLrun的存储后端需要特定的认证与配置,手动读取无法自动处理这些逻辑。推荐使用MLrun原生API来读取工件,以下是具体方法:
在MLrun函数内部读取
如果是在另一个MLrun函数中读取,利用上下文对象context的API即可自动处理存储认证:
方法1:通过工件名称+项目+来源函数名获取
from mlrun.execution import MLClientCtx def target_function(context: MLClientCtx): # 获取指定工件 artifact = context.get_artifact( name="mydf", project="test-pipeline", function="data-prep-test-data-generator" ) # 直接转为DataFrame df = artifact.as_df()
方法2:通过MLrun的store:// URI获取
MLrun提供了统一的工件引用格式store://artifacts/<项目名>/<来源函数名>/<工件名>,用这个URI更稳定,无需硬编码S3路径:
from mlrun.execution import MLClientCtx def target_function(context: MLClientCtx): artifact_uri = "store://artifacts/test-pipeline/data-prep-test-data-generator/mydf" # 获取DataItem对象 data_item = context.get_dataitem(artifact_uri) # 转为DataFrame df = data_item.as_df()
在MLrun外部(本地脚本)读取
如果是在MLrun集群外的本地环境读取,先初始化MLrun客户端,再用get_dataitem方法:
import mlrun # 设置当前项目 mlrun.set_environment(project="test-pipeline") # 通过store URI获取工件 data_item = mlrun.get_dataitem("store://artifacts/test-pipeline/data-prep-test-data-generator/mydf") # 转为DataFrame df = data_item.as_df()
关键说明
store://URI是MLrun的标准化工件引用,比直接用S3路径更可靠,不会因为存储配置变更失效。as_df()方法会自动根据工件类型(CSV/Parquet等)选择正确的读取方式,无需手动指定格式。
内容的提问来源于stack exchange,提问作者Sangeeta Prasad
相关产品推荐
相关产品推荐

