Azure SDK v2中如何将pandas DataFrame转回Mltable?
解决方案:处理后DataFrame转回MLTable及转为URI File(Azure SDK v2)
一、将处理后的Pandas DataFrame转回MLTable
MLTable本质是数据读取定义文件,并不直接存储数据,所以必须先把处理后的数据落地到Azure存储(比如Datastore),再基于存储路径重新创建MLTable对象。具体步骤如下:
import mltable from azure.ai.ml import MLClient from azure.identity import DefaultAzureCredential import pandas as pd import os import tempfile # 假设processed_df是你处理完成后的DataFrame processed_df = df # 替换为你的实际处理后数据 # 1. 临时保存数据到本地CSV(避免内存占用过大) temp_dir = tempfile.mkdtemp() local_output = os.path.join(temp_dir, "processed_data.csv") processed_df.to_csv(local_output, index=False) # 2. 连接Azure ML工作区 ml_client = MLClient( DefaultAzureCredential(), subscription_id=subscription_id, resource_group_name=resource_group, workspace_name=workspace_name ) # 3. 上传本地文件到目标Datastore datastore_name = "my_container" target_path = "processed_data/processed_data.csv" # Datastore内的存储路径 ml_client.datastores.upload( datastore_name=datastore_name, source_file=local_output, target_path=target_path, overwrite=True ) # 4. 基于存储路径创建MLTable processed_mltable_path = { "file": f"azureml://subscriptions/{subscription_id}/resourcegroups/{resource_group}/workspaces/{workspace_name}/datastores/{datastore_name}/paths/{target_path}" } processed_tbl = mltable.from_delimited_files(paths=[processed_mltable_path])
这样生成的processed_tbl就可以直接作为管道组件的MLTable类型输入使用。
二、将Pandas DataFrame转为URI File用于管道组件
如果你的管道组件要求输入/输出为uri_file类型,只需将处理后的数据写入Datastore,然后获取对应的Azure存储URI即可:
from azure.storage.blob import BlobServiceClient # 沿用上面的ml_client、datastore_name、processed_df datastore = ml_client.datastores.get(datastore_name) target_uri_path = "processed_data/processed_data.csv" # 1. 获取Datastore的存储密钥,连接Blob服务 datastore_key = ml_client.datastores.get_secret(datastore_name) blob_service_client = BlobServiceClient( account_url=f"https://{datastore.account_name}.blob.core.windows.net", credential=datastore_key ) # 2. 将DataFrame转为CSV字节流,直接上传到Blob存储 container_client = blob_service_client.get_container_client(datastore.container_name) csv_bytes = processed_df.to_csv(index=False).encode("utf-8") blob_client = container_client.get_blob_client(target_uri_path) blob_client.upload_blob(csv_bytes, overwrite=True) # 3. 生成URI File路径 uri_file = f"azureml://subscriptions/{subscription_id}/resourcegroups/{resource_group}/workspaces/{workspace_name}/datastores/{datastore_name}/paths/{target_uri_path}"
这个uri_file字符串就是符合要求的URI File,可直接传入管道组件的对应输入参数。
注意事项
- 若管道组件接受MLTable类型,用第一部分的
processed_tbl;若接受URI File,用第二部分的uri_file,根据组件定义选择即可。 - 频繁创建临时文件时,记得清理临时目录避免磁盘占用。
内容的提问来源于stack exchange,提问作者Egorsky
相关产品推荐
相关产品推荐

