如何使用Azure ML SDK v2读取已注册容器实例中的CSV文件
在Azure ML SDK v2中读取Datastore文件的正确方法
问题分析
你之前的代码失败原因:
- 第一段
mltable.from_delimited_files的paths参数格式错误,mltable SDK v2要求路径以特定格式(URI或带pattern的字典)传入,而非直接传递Datastore对象和路径字符串。 - 第二段
ml_client.data._get_latest_version是用于获取已注册到工作区的Data Asset(数据集),并非直接读取Datastore中的原始文件,因此无法关联到你的Datastore路径。
方法1:用mltable加载为表格数据
这是替代SDK v1中Dataset.Tabular的正确方式,需按mltable要求的路径格式构造:
# 导入依赖包 from azure.ai.ml import MLClient from azure.identity import DefaultAzureCredential import mltable # 初始化MLClient(若未初始化) ml_client = MLClient( DefaultAzureCredential(), subscription_id="你的订阅ID", resource_group_name="你的资源组名称", workspace_name="你的工作区名称" ) # 获取目标Datastore target_datastore = ml_client.datastores.get("my_container") # 构造符合mltable要求的路径:使用Datastore的URI格式 path_config = [{"pattern": f"azureml://datastores/{target_datastore.name}/paths/data/my_data.csv"}] # 加载为mltable对象 tbl = mltable.from_delimited_files(paths=path_config) # 转换为Pandas DataFrame(按需使用) df = tbl.to_pandas_dataframe()
方法2:直接读取/下载文件内容
如果不需要tabular格式操作,可直接下载文件到本地或读取到内存:
方式A:通过Datastore直接下载
# 将文件下载到本地指定路径 target_datastore.download( target_path="./local_download", # 本地保存路径 prefix="data/my_data.csv" # Datastore中的文件路径 )
方式B:通过Azure Storage SDK读取到内存
from azure.storage.blob import BlobClient import pandas as pd from io import StringIO # 从Datastore获取存储账户信息 storage_account_name = target_datastore.account_name storage_account_key = target_datastore.account_key container_name = target_datastore.container_name # 初始化Blob客户端 blob_client = BlobClient( account_url=f"https://{storage_account_name}.blob.core.windows.net", container_name=container_name, blob_name="data/my_data.csv", credential=storage_account_key ) # 读取文件内容并转换为DataFrame with blob_client.download_blob() as blob: file_content = blob.readall().decode("utf-8") df = pd.read_csv(StringIO(file_content))
可选:注册为Data Asset后管理和读取
若需版本化管理数据集,可先将Datastore中的文件注册为Data Asset,再通过ml_client获取:
from azure.ai.ml.entities import Data from azure.ai.ml.constants import AssetTypes # 定义Data Asset my_data_asset = Data( path=f"azureml://datastores/my_container/paths/data/my_data.csv", type=AssetTypes.URI_FILE, name="my_data_csv", description="从my_container datastore读取的CSV数据" ) # 注册到工作区 ml_client.data.create_or_update(my_data_asset) # 后续获取最新版本的Data Asset data_asset = ml_client.data.get(name="my_data_csv", version="latest") # 可通过data_asset.path获取文件URI,再用mltable或Storage SDK读取
内容的提问来源于stack exchange,提问作者Egorsky
相关产品推荐
相关产品推荐

