You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从Azure机器学习笔记本直接将DataFrame写入Blob存储?

解决方案:直接将Pandas DataFrame写入Azure ML关联Blob存储

以下是两种无需本地临时文件、直接将处理后的DataFrame写入Blob存储的可行方案:

方案1:使用Azure Storage Blob SDK直接写入

通过Azure Blob存储官方SDK,将DataFrame的CSV内容以字节流形式直接上传,完全跳过本地存储步骤。

from azureml.core import Workspace
from azure.storage.blob import BlobClient
import pandas as pd

# 加载Azure ML工作区
ws = Workspace.from_config()
# 获取目标Blob存储对应的datastore
target_datastore = ws.datastores.get("exampleblobstore")

# 初始化Blob客户端
blob_client = BlobClient(
    account_url=f"https://{target_datastore.account_name}.blob.core.windows.net",
    container_name=target_datastore.container_name,
    blob_name="path/to/target.csv",  # 存储路径+目标文件名
    credential=target_datastore.account_key
)

# 将DataFrame转为CSV字节流并上传
csv_content = df.to_csv(index=False).encode("utf-8")
blob_client.upload_blob(csv_content, overwrite=True)

方案2:通过Azure ML Dataset直接注册DataFrame

利用Azure ML的Dataset生态,直接将Pandas DataFrame注册为Tabular Dataset,自动将数据写入指定datastore,同时可复用该Dataset进行后续机器学习流程。

from azureml.core import Workspace, Dataset
import pandas as pd

ws = Workspace.from_config()
target_datastore = ws.datastores.get("exampleblobstore")

# 注册DataFrame到指定datastore
tabular_dataset = Dataset.Tabular.register_pandas_dataframe(
    dataframe=df,
    target=(target_datastore, "path/to/target.csv"),  # 目标datastore及存储路径
    name="processed_data_set",  # Dataset标识名称
    overwrite=True  # 覆盖已存在的同名文件/Dataset
)

关于AzureMachineLearningFileSystem的说明

当前AzureMachineLearningFileSystem类主要聚焦于读取操作与文件系统浏览,官方暂未提供完整的写入支持,因此上述两种方案是更直接的替代选择。

内容的提问来源于stack exchange,提问作者Matt_Haythornthwaite

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 09:34:55