You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在AzureML中保存分区parquet文件并自定义存储路径的方案咨询

AzureML环境下自定义路径存储分区Parquet到Blob Storage的实现方法

完全可以在AzureML的许可使用范围内实现自定义路径的分区Parquet存储,不需要依赖默认生成随机UUID路径的数据集注册方法,以下是两种生产可用的实现方案:

方案1:Datastore + 本地临时写入上传(兼容性最好)

该方案依托AzureML官方原生的Datastore能力实现,适配所有AzureML计算集群/实例环境:

  • 第一步:获取目标Blob存储对应的AzureML Datastore(需提前在AzureML工作区中将目标Blob挂载为Datastore,属于AzureML标准操作)
    from azureml.core import Workspace, Datastore, Dataset
    # 加载当前AzureML工作区配置
    ws = Workspace.from_config()
    # 替换为你要写入的目标Blob对应的Datastore名称
    target_datastore = Datastore.get(ws, datastore_name="customer_segment_datastore")
    
  • 第二步:将pandas分群结果按reference_dt分区写入计算节点本地临时目录
    import pandas as pd
    # 你的分群结果DataFrame,此处替换为实际变量名
    segment_df: pd.DataFrame
    # 按reference_dt生成分区Parquet到本地临时目录
    segment_df.to_parquet(
        path="./temp_segment_result",
        partition_cols=["reference_dt"],
        engine="pyarrow",
        index=False
    )
    
  • 第三步:将本地分区文件上传到Blob的自定义路径,路径完全由你指定
    target_blob_path = "customer_segmentation/production_results/"
    target_datastore.upload(
        src_dir="./temp_segment_result",
        target_path=target_blob_path,
        overwrite=True,
        show_progress=False
    )
    
  • 可选:如果需要在AzureML中注册数据集方便后续使用,直接基于你自定义的路径注册即可,不会修改已经存储的文件路径
    # 基于自定义路径创建表格数据集,支持自动识别分区字段
    segment_dataset = Dataset.Tabular.from_parquet_files(
        path=[(target_datastore, f"{target_blob_path}**/*.parquet")],
        partition_format="/reference_dt={reference_dt:yyyy-MM-dd}/"
    )
    # 注册数据集,更新时仅生成新的逻辑版本,底层存储路径固定
    segment_dataset.register(
        workspace=ws,
        name="customer_segment_result",
        create_new_version=True
    )
    

方案2:直接内存写入Blob(大文件场景更高效)

依托Azure官方维护的adlfs库(AzureML环境默认预装),无需写入本地临时文件,直接将DataFrame写入Blob指定路径:

from azureml.core import Workspace, Datastore
import adlfs
import pandas as pd

ws = Workspace.from_config()
target_datastore = Datastore.get(ws, datastore_name="customer_segment_datastore")
# 初始化Blob文件系统实例
fs = adlfs.AzureBlobFileSystem(
    account_name=target_datastore.account_name,
    account_key=target_datastore.account_key
)
# 直接写入分区Parquet到Blob自定义路径
segment_df.to_parquet(
    path=f"abfs://{target_datastore.container_name}/customer_segmentation/production_results/",
    partition_cols=["reference_dt"],
    engine="pyarrow",
    filesystem=fs,
    index=False
)

注意事项

  • 若需要保留历史分群版本,可在自定义路径中增加版本标识字段,例如按处理日期设置路径为customer_segmentation/results/dt=20240520/即可
  • 官方register_pandas_dataframe方法本身定位为快速实验场景的便捷接口,默认由系统托管存储路径,不适用于需要固定路径的生产自动化流程
  • 上述所有操作均完全在AzureML的权限与功能体系内实现,符合仅使用AzureML开展工作的合规要求

内容的提问来源于stack exchange,提问作者Original BBQ Sauce

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.28 15:36:08